Antique information identification system and method introducing self-attention mechanism
By introducing a self-attention mechanism into the antique information recognition system, the problem of low accuracy and inability to extract global features in antique image recognition is solved, and higher recognition accuracy and more comprehensive feature extraction are achieved.
Patent Information
- Application Number
- CN202510550943.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art has low accuracy in antique image recognition, and cannot effectively extract global features and deal with complex damaged areas.
Introduce a self-attention mechanism, perform hierarchical structured deployment through multiple self-attention dimensions, build a self-attention recognizer, and deploy it in an embedded manner on the antique information platform. The system performs image recognition by blurring texture attention, feature attention, relational attention, cross-modal attention and global attention, performs directional recognition under meta-concurrency self-attention, and aggregates and reorganizes the concurrent recognition results.
It improves the accuracy of antique image recognition, comprehensively extracts feature information, can effectively deal with complex damaged areas, and provides more accurate and comprehensive recognition results.
Smart Images

Figure CN120107978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an antique information recognition system and method that introduces a self-attention mechanism. Background Art
[0002] At present, antique image recognition technology faces the problems of low accuracy, inability to effectively extract global features and handle complex damaged areas. Traditional methods usually rely on local feature extraction, ignoring the association between global information and complex areas in the image, resulting in unsatisfactory recognition results when facing comprehensive recognition of antique images. Summary of the invention
[0003] The present application provides an antique information recognition system and method that introduces a self-attention mechanism, which is used to solve the technical problems of low accuracy in antique image recognition, inability to effectively extract global features and process complex damaged areas in the prior art.
[0004] In view of the above problems, the present application provides an antique information recognition system and method that introduces a self-attention mechanism.
[0005] In a first aspect of the present application, an antique information recognition system is provided that introduces a self-attention mechanism, the system comprising:
[0006] The identifier deployment module is used to introduce multiple self-attention dimensions, perform hierarchical structured deployment, build a self-attention identifier and embed it in the antique information platform; the information identification module is used to upload antique scans according to the multi-threaded port of the antique information platform, pre-process the antique scans and import them into the self-attention identifier, perform self-attention meta-triggering and decision-making based on recognition needs, and determine the information recognition results; wherein, the recognition steps based on the self-attention identifier include: by performing image recognition and block segmentation based on fuzzy textures, performing directional recognition under meta-concurrent self-attention on the block images, aggregating and reorganizing the concurrent recognition results, and determining the information recognition results, wherein the meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations.
[0007] The second aspect of the present application provides an antique information recognition method that introduces a self-attention mechanism, the method comprising:
[0008] Introduce multiple self-attention dimensions, perform hierarchical structured deployment, build a self-attention recognizer and embed it in an antique information platform; upload antique scans according to the multi-threaded port of the antique information platform, pre-process the antique scans and import them into the self-attention recognizer, perform self-attention meta-triggering and decision-making based on recognition needs, and determine information recognition results; wherein, the recognition steps based on the self-attention recognizer include: by performing fuzzy texture-based image recognition and block segmentation, performing directional recognition under meta-concurrent self-attention on the block image, aggregating and reorganizing the concurrent recognition results, and determining the information recognition results, wherein the meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] This application introduces multiple self-attention dimensions, performs hierarchical structured deployment, constructs a self-attention identifier and deploys it in an embedded manner on an antique information platform; uploads antique scans according to the multi-threaded port of the antique information platform, pre-processes the antique scans and imports the self-attention identifier, performs self-attention meta-triggering and decision-making based on recognition needs, and determines information recognition results; wherein, the recognition steps based on the self-attention identifier include: by performing image recognition and block segmentation based on fuzzy textures, performing directional recognition under meta-concurrent self-attention for block images, aggregating and reorganizing the concurrent recognition results, and determining information recognition results, wherein meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations. The present invention solves the technical problems of low accuracy, inability to effectively extract global features, and inability to process complex damaged areas in the prior art in antique image recognition, and achieves the technical effect of improving antique image recognition accuracy and comprehensively extracting feature information by introducing multiple self-attention dimensions, performing hierarchical structured deployment, and concurrent self-attention mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 A schematic diagram of the structure of an antique information recognition system that introduces a self-attention mechanism provided in an embodiment of the present application;
[0013] Figure 2 A flow chart of an antique information recognition method that introduces a self-attention mechanism provided in an embodiment of the present application.
[0014] Explanation of the reference numerals: identifier deployment module 11, information identification module 12. DETAILED DESCRIPTION
[0015] The present application provides an antique information recognition system and method that introduces a self-attention mechanism, aiming to solve the technical problems of low accuracy in antique image recognition, inability to effectively extract global features and process complex damaged areas in the prior art. By introducing multiple self-attention dimensions, hierarchical structured deployment and concurrent self-attention mechanisms, the technical effect of improving the accuracy of antique image recognition and comprehensively extracting feature information is achieved.
[0016] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0017] It should be noted that any variations of the terms "include" and "have" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or inherent to these processes, methods, products or devices.
[0018] Embodiment 1, as Figure 1 As shown, the embodiment of the present application provides an antique information recognition system that introduces a self-attention mechanism, and the system includes:
[0019] The identifier deployment module 11 is used to introduce multiple self-attention dimensions, perform hierarchical structured deployment, build a self-attention identifier and embed it in the antique information platform.
[0020] In the embodiment of the present application, the main task of the recognizer deployment module 11 is to build and deploy a self-attention recognizer with high-precision and multi-dimensional analysis capabilities to support image recognition and feature judgment tasks in the antique information platform. First, the recognizer deployment module 11 establishes an attention mechanism system covering different perception levels and information structures by introducing multiple self-attention dimensions. This mechanism covers fuzzy texture attention, feature attention, relational attention, cross-modal attention and global attention. Fuzzy texture attention is used to capture the fuzzy areas in antique images due to age, wear and aging, and enhance the perception of fuzzy features; feature attention focuses on identifying the most representative and recognizable image areas in antiques, such as specific patterns, carvings, material traces, etc.; relational attention is used to model the spatial relationship or semantic connection between local features in the image, such as positional dependence or style combination between components; cross-modal attention enables the recognizer to integrate visual information with existing text descriptions, expert labels or database knowledge in the platform, to achieve linkage understanding between image information and non-image information; and global attention integrates image information from a holistic perspective, enhancing the recognizer's ability to understand the overall style, configuration, layout and other macro features of antiques.
[0021] On the basis of the establishment of the attention mechanism, the recognizer deployment module 11 deploys various attention mechanisms in a hierarchical structure to improve the functional coordination efficiency and hierarchical perception ability. First, the first recognition layer is constructed based on the fuzzy texture attention. This layer is responsible for parsing the fuzzy area of the antique image and preliminary block processing, providing high-quality local input for subsequent layers; then the feature attention, relationship attention and cross-modal attention are deployed in parallel in the second recognition layer. In this layer, the recognizer simultaneously processes local feature extraction, regional relationship modeling and image and text information fusion to form multi-channel concurrent perception capabilities; finally, the third recognition layer is deployed according to the global attention mechanism, and the results of the first two layers are aggregated and reorganized, and the recognition decision results are output based on the overall characteristics. A unified self-attention recognizer structure is constructed between the three recognition layers through a fully connected method to ensure that different attention mechanisms are both independent and fused and collaborative in function, thereby improving the overall recognition accuracy and generalization ability.
[0022] The actual deployment of the self-attention recognizer depends on a sufficient training process and diverse data support. In the training phase, a dataset containing a large number of antique images is first prepared. These images cover antique samples of various materials, various historical periods, different shooting angles and resolutions to ensure the comprehensiveness and representativeness of the training data. In addition to image data, cross-modal data such as expert annotation information, literature, text descriptions and expert comments are also included to train the cross-modal attention mechanism and enhance the recognizer's ability to integrate multi-source information. Before training, all image data must be standardized and pre-processed, including image denoising, brightness alignment, resolution scaling, etc., so that the input data meets the model processing requirements. During the training process, the self-attention recognizer optimizes the weight parameters in the self-attention structure through multiple rounds of iterative learning, so that different attention mechanisms can learn the appropriate weight distribution and trigger path according to the training samples. Fuzzy texture attention learns how to distinguish texture details in high-noise images, feature attention learns to identify the most distinguishing image areas in different antiques, and relationship attention trains the model to understand the spatial or structural relationship between local images. Cross-modal attention builds semantic linkage between images and texts to improve the accuracy and interpretation ability of multi-dimensional recognition. The training process also includes the training of the self-attention meta-trigger mechanism for recognition needs, so that the recognizer can autonomously activate the relevant recognition layer or attention node according to the recognition task, thereby realizing the recognition process of on-demand activation and flexible response. The performance of the model is continuously evaluated and tuned throughout the training process, and cross-validation, precision evaluation, recall rate monitoring and other means are used to ensure that the recognizer has stable and reliable performance after training.
[0023] After the self-attention recognizer is trained and optimized, the recognizer deployment module 11 integrates it into the antique information platform in an embedded manner to realize an integrated hardware and software operating environment.
[0024] Furthermore, in the system provided by the embodiment of the application, the introduction of multiple self-attention dimensions also includes:
[0025] Introduce fuzzy texture attention, deploy the block mode, and set the first attention node; introduce feature attention and set the second attention node; introduce relational attention and set the third attention node; introduce cross-modal attention and set the fourth attention node; introduce global attention and set the fifth attention node.
[0026] In an embodiment of the present application, fuzzy texture attention is first introduced and processed through a block mode. Fuzzy texture attention is used to identify fuzzy areas in an image due to aging, wear or other external factors. In order to effectively extract the texture features of these fuzzy areas, the image is divided into multiple small blocks, and the blocking method is selected according to the regularity or irregularity of the texture. In the case of regular textures, non-uniform geometric blocks, such as irregular rectangles or polygons, are used to better adapt to the texture structure in the image; for irregular textures, uniform geometric blocks, such as squares or rectangles, are used to ensure that the size and shape of each block are consistent. In this way, fuzzy texture attention can focus on the key texture areas in each small block and enhance the recognition of details in complex images. This step is set as the first attention node, which is responsible for processing the fuzzy texture areas in the image.
[0027] Next, feature attention is introduced and set as the second attention node. The feature attention mechanism helps the model focus on key feature areas in the image, especially those local features that are most important for recognition. For example, carvings, patterns or unique marks in antique images are crucial to judging the authenticity and historical value of objects. Feature attention weights these key features so that the model can prioritize and accurately identify these local areas, even in complex backgrounds or in the presence of noise. This mechanism ensures that the system can efficiently extract the most important information in the image, thereby improving recognition accuracy.
[0028] Then, relational attention is introduced and set as the third attention node. The relational attention mechanism helps the model understand the spatial relationship or semantic connection between different regions in the image. By analyzing the interdependence and spatial relationship between local regions, relational attention can identify the structural relationship between multiple parts in the image. For example, in the image of ancient sculptures, there may be certain spatial and structural connections between different parts of the sculpture (such as hands, faces, bases, etc.). Relational attention can identify these dependencies and use them as the basis for overall recognition. In this way, the model can capture the association between local features, thereby enhancing the understanding of the overall structure of the image.
[0029] Next, cross-modal attention is introduced and set as the fourth attention node. Cross-modal attention enhances the recognition ability of the model by combining images with other types of information (such as text descriptions, label information, etc.). The image itself may contain some features that are difficult to recognize directly visually, and the text information related to the image (such as expert labels, item descriptions, etc.) can provide important contextual clues for recognition. The cross-modal attention mechanism combines image information with these text information, so that the model not only relies on visual information during the recognition process, but also obtains supplementary information from other modalities to improve the understanding of complex scenes.
[0030] Finally, global attention is introduced and set as the fifth attention node. The purpose of the global attention mechanism is to analyze the image as a whole and integrate the features of each local area. Global attention ensures that the model can not only focus on local features when processing images, but also understand the overall structure of the image through the integration of global information. For example, in antique images, the overall layout and the relationship between the parts may provide important clues. Through global attention, the model can integrate all local features to form a comprehensive understanding of the entire image. The global attention mechanism enables the recognizer to understand the overall picture of the image from a higher level and improve the overall analysis ability of the image.
[0031] Through these steps, the introduced attention mechanisms are set as the first to fifth attention nodes in turn, ensuring that the system can deeply analyze and understand the image from multiple dimensions such as texture, local features, regional relations, cross-modal information and global structure, thereby achieving accurate image recognition.
[0032] Furthermore, in the system provided in the embodiment of the application, performing hierarchical structured deployment also includes:
[0033] According to the first attention node, a first recognition layer is constructed; the second attention node, the third attention node and the fourth attention node are deployed in parallel to construct a second recognition layer; according to the fifth attention node, a third recognition layer is constructed; the first recognition layer, the second recognition layer and the third recognition layer are fully connected in three layers to determine a self-attention recognizer.
[0034] In an embodiment of the present application, in the process of constructing a self-attention recognizer, first, according to the first attention node, a first recognition layer is constructed. The recognition layer focuses on processing the fuzzy texture information in the image, using the fuzzy texture attention mechanism. The task of fuzzy texture attention is to enhance the features of the fuzzy area in the image due to factors such as illumination, wear, and aging. In the specific operation, the Canny edge detection algorithm is first applied to enhance the image edge to enhance the detail information in the image, especially the boundary of the fuzzy area. Then the image is divided into a plurality of small blocks, which are segmented according to the regularity or irregularity of the image texture. For images with regular textures, non-uniform geometric blocks, such as irregular rectangles or polygons, are used to adapt to the complex texture in the image; and for images with irregular textures, uniform geometric blocks, such as squares or rectangles, are used. Each image block is processed by the fuzzy texture attention mechanism, which enhances the detail information of the fuzzy area in the image through a weighted strategy, so that the fuzzy texture features become clearer. During the training process, an image data set with fuzzy area annotations is used, and through supervised learning, the model adjusts parameters according to the loss function to optimize the texture enhancement ability of the fuzzy area, thereby effectively extracting the potential texture features in the image.
[0035] Next, the second, third, and fourth attention nodes are introduced and deployed in parallel to construct the second recognition layer. This layer processes different feature dimensions of the image through three parallel sub-channels. First, the second attention node uses a convolutional neural network (CNN) to extract local features in the image. The convolution operation scans the image through a series of convolution kernels to automatically learn key local features in the image, such as carvings, patterns, or other details. During training, a dataset of antique images with local feature labels, such as images with patterns or cracks, is used. After these local features are processed by the convolution layer, feature maps are formed and passed to downstream processing. Secondly, the third attention node uses a graph convolutional network (GCN) to model the spatial relationship between image regions. In this step, each region of the image is regarded as a node of the graph, and adjacent regions are connected by the edges of the graph, and the information between regions is aggregated through the graph convolution operation. For example, the relative position relationship of different parts in the image (such as handles and bases) will be modeled by a graph convolutional network to help the recognizer understand the dependencies between these regions. Finally, the fourth attention node aligns the image and text description through the joint embedding space, matching the image features with the text description (such as "Qing Dynasty blue and white porcelain with dragon patterns and flowers") in the same semantic space. By using the multi-head attention mechanism, the image and text information can interact in this space, helping the model to obtain the semantic context of the image from the text description. During the training process, through the pairing information of the image and text description, the model learns how to establish an effective connection between the image content and the text description. In this way, the second recognition layer deployed in parallel can simultaneously extract the local features, structural relationships and semantic information of the image, providing a more comprehensive feature representation for subsequent recognition tasks.
[0036] Subsequently, the third recognition layer is constructed based on the fifth attention node. This layer uses the self-attention mechanism to model the entire image. The self-attention mechanism calculates the similarity between each region in the image and other regions, and uses these similarities as weights to integrate the information of each part of the image. Specifically, the image is divided into multiple patches of fixed size, each patch is mapped to a vector, and the correlation between each patch is calculated through the self-attention mechanism. During the training process, an image dataset containing global labels, such as the style or category label of the image, is used to supervise the model to extract important information from the global features of the image. For example, when identifying complex sculptures, the third recognition layer can integrate the information of all parts of the image (such as the head, base, decoration, etc.) through the global attention mechanism to obtain an overall understanding of the image. The purpose of this layer is to capture structural features within the entire image and improve the model's perception of global information.
[0037] Finally, the first recognition layer, the second recognition layer and the third recognition layer are fully connected to determine the self-attention recognizer. The goal of this process is to fuse the feature information output by the three layers, integrating the fuzzy texture features, local detail features, spatial relationship features and global information from different recognition layers. Through the fully connected layer, the output features of each layer are weighted summed or spliced to form a high-dimensional comprehensive feature representation. Subsequently, after processing by a multi-layer neural network, all the information in the image is integrated together to finally generate the recognition result. Through this process, the final self-attention recognizer can not only process local textures, local features, and spatial relationships, but also understand the overall structure of the image, providing efficient and accurate image recognition results.
[0038] Furthermore, the system provided in the application embodiment also includes:
[0039] The blocking mode is determined according to the texture features; if it is a regular texture, non-uniform geometric blocking is adopted, and if it is an irregular texture, uniform geometric blocking is adopted.
[0040] In the embodiment of the present application, the appropriate block mode is first determined according to the texture features. The core of this process is to select the most appropriate image segmentation strategy according to the texture type (regular or irregular) of the image, so as to better capture and process the details in the image.
[0041] When the image presents regular texture, non-uniform geometric blocking is used for segmentation. Regular textures usually have obvious, repetitive structures and patterns, such as neat grid-like textures, uniform patterns or designs, and these features are distributed more consistently throughout the image. For example, in images of antique porcelain, regular textures may appear as uniform geometric patterns or consistent line layouts. In order to process such images more effectively, non-uniform geometric blocking is used to segment image blocks with irregular shapes (such as irregular rectangles, pentagons, etc.) to accommodate the more regular details in the image. This blocking method can accurately capture the local features of the texture and avoid the loss of texture features caused by the traditional uniform blocking method.
[0042] For images with irregular textures, uniform geometric blocking is used for segmentation. Irregular textures often appear as irregular texture structures and forms, such as wear marks in damaged areas or naturally formed random patterns. These textures lack predictable regularity, and the details in the image vary greatly, so they need to be processed using uniform geometric blocking (such as squares or rectangles). This approach ensures that each block has a consistent shape and size, avoiding the effects of texture irregularities. Uniform blocking enables the model to process each area in the image and extract independent features from each area, while reducing interference caused by inconsistent textures during the blocking process.
[0043] By selecting these block segmentation methods, the texture information of the image can be effectively extracted and optimized according to the characteristics of the image (regular or irregular) after block processing.
[0044] The information recognition module 12 is used to upload the antique scan image according to the multi-threaded port of the antique information platform, pre-process the antique scan image and import it into the self-attention recognizer, and perform self-attention meta-triggering and decision-making based on the recognition needs to determine the information recognition result; wherein, the recognition step based on the self-attention recognizer includes: by performing fuzzy texture-based image recognition and block segmentation, performing meta-concurrent self-attention directional recognition on the block image, aggregating and reorganizing the concurrent recognition results, and determining the information recognition result, wherein the meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations.
[0045] In the embodiment of the present application, the information recognition module 12 is responsible for uploading the antique scans, performing preprocessing and importing the self-attention recognizer to achieve efficient and accurate image recognition. In this process, the multi-threaded port first undertakes the image uploading task. Through multi-threading technology, the antique information platform can achieve parallel uploading of multiple images, thereby effectively improving the transmission speed and processing efficiency of image data. Each thread processes an image, which enables the platform to process multiple antique scans at the same time, ensuring that the efficient operation of the system can be maintained when a large amount of data flows in.
[0046] After the upload is complete, the image enters the preprocessing stage. The goal of this stage is to denoise, standardize, resize, and other operations on the image to improve the quality and consistency of the image so that it is suitable for subsequent recognition tasks. Image processing techniques include Gaussian blur denoising, histogram equalization (enhancing contrast), and normalization to ensure that the image content is clear, standardized, and the color and brightness meet the standards. After the image is preprocessed, it is input into the previously trained self-attention recognizer. This recognizer relies on the self-attention mechanism, a technology that can capture the relationship between various regions in an image. The self-attention mechanism weights each local region in the image based on its correlation with other regions to determine the importance of different regions. In this way, the model can dynamically focus on the recognizable parts of the image and enhance its understanding of the global structure of the image. As a result, the model can not only efficiently process the details of the image, but also understand the connections and global patterns between different regions.
[0047] After entering the self-attention recognizer, recognition demand, as a key factor in the recognition process, will guide the meta-triggering and decision-making of the self-attention recognizer. Specifically, depending on the type of image, the recognizer will flexibly adjust the processing strategy according to the task requirements. For example, for antique images with complex textures, the recognizer may prioritize the details of the image based on the texture features; while for images with obvious morphological features, it may focus on recognizing the shape and structure of the image. This recognition demand-oriented mechanism enables the self-attention recognizer to adaptively optimize and adjust according to the actual characteristics of the image and the task objectives, thereby improving recognition efficiency and accuracy.
[0048] The next steps involve specific recognition steps based on the self-attention recognizer. First, the image will undergo fuzzy texture-based image recognition and block processing. By analyzing the texture features of the image, the image will first be divided into multiple small blocks, and the block mode will be selected according to the texture type of the image: if the image presents a regular texture, non-uniform geometric block, such as irregular rectangles or polygons, is used to adapt to the texture pattern of the image; if it is an irregular texture, uniform geometric block, such as squares or rectangles, is used to ensure that the scale of each block is consistent, which is convenient for processing areas with irregular textures. This block processing ensures the independence and flexibility of each area of the image during the recognition process, providing important support for subsequent directional recognition.
[0049] After the block processing is completed, the image blocks will be directional recognized under meta-concurrent self-attention. Meta-concurrent self-attention refers to the parallel calculation of the features of each image block in multiple parallel channels, and the weighting of the feature elements of each block through the self-attention mechanism. This process not only processes the block feature elements of the image (such as texture, color, shape, etc.), but also considers the inter-block relationship (that is, the spatial and semantic dependencies between different blocks). For example, in an image, different parts of the pattern (such as decoration, cracks, etc.) may be in different image blocks. The meta-concurrent self-attention mechanism helps the recognizer understand the connection between these parts by establishing the relationship between the blocks. In addition, the processing of cross-modal relationships is also carried out at this stage. The combination of images and text descriptions (such as "Ming Dynasty blue and white porcelain with dragon patterns and flowers") is processed through the self-attention mechanism, which improves the accuracy and richness of image recognition.
[0050] The recognition results of all blocks are then aggregated and reorganized. This process ensures that the feature contribution of each recognition path is fully utilized by weighted fusion of concurrent recognition results. Common aggregation methods include weighted summation, splicing, and gating mechanisms, which can combine the local features, spatial relationship features, and semantic information of the image to generate the final recognition results. Ultimately, through this process, the information recognition module 12 can accurately integrate the outputs from different recognition layers to provide users with clear and accurate antique image recognition results.
[0051] Finally, the information recognition results will be output to the platform users. These results include the classification labels and related attributes of the image, providing reliable data support for antique identification, style analysis, etc.
[0052] Furthermore, the system provided in the application embodiment also includes:
[0053] Through permission constraints, a multi-threaded port is deployed, wherein the multi-threaded port includes at least a user side and an expert side; the antique scan image is uploaded according to the multi-threaded port; wherein the uploading method includes antique scanning and uploading based on the built-in data acquisition card of the antique information platform, and image document retrieval and uploading.
[0054] In an embodiment of the present application, in the information identification platform, the access rights of different users and experts when using the platform are ensured through a permission constraint mechanism to ensure the security and rationality of the data. First, the information identification platform supports different types of operations through multi-threaded port deployment. Specifically, the multi-threaded port includes at least a user side and an expert side. On the user side, ordinary users can access and upload image information related to antiques, while on the expert side, expert users can further analyze and identify images and provide professional feedback. In this way, the platform can provide different functional modules for users and experts according to different permissions, while realizing the management and control of uploaded data, ensuring that the data flow and operation of the platform comply with the permission regulations.
[0055] Based on the multi-threaded port, the information recognition platform can efficiently support the image upload function. There are two main ways to upload. The first is to upload the antique scans through the built-in data acquisition card of the antique information platform. The data acquisition card is part of the platform's internal hardware, which is used to convert the antique scans in the physical world into digital images. The data acquisition card connects the antique scanner to the platform, collects images and uploads them to the platform in real time. During the image acquisition process, the acquisition card will use a high-resolution scanner to perform a fine scan of the antique, and upload the image data to the platform in real time through a high-speed transmission channel to ensure high fidelity of the image quality. This upload method is often used in scenarios that require high-precision image acquisition, especially for items with rich details such as antiques.
[0056] The second way to upload is through image document retrieval and uploading. This method is often used to retrieve and upload antique images from existing image libraries or databases. Users or experts can use the document retrieval function provided by the platform to select digital images related to antiques and upload them through the platform's interface. These images may come from external databases, archived files or other platforms, and are quickly imported into the platform through the image document interface. In this way, image uploading is more convenient and suitable for the integration and management of antique images that already exist in digital formats.
[0057] The two uploading methods each have their own advantages. Data acquisition card uploading can provide extremely high image quality and a fine scanning process, which is particularly suitable for original images that need to be acquired at high resolution. Image document retrieval and uploading is more convenient, and can quickly integrate existing digital image resources and improve the efficiency of platform data processing.
[0058] Through these steps, the platform provides flexible and efficient upload channels based on different needs and permission constraints, ensuring that antique image data can enter the system safely and timely for subsequent identification and analysis.
[0059] Furthermore, in the system provided in the embodiment of the application, self-attention element triggering and decision making are performed based on identification needs, and further include:
[0060] Recognition requirements are received, and the recognition requirements are converted into self-attention elements, wherein the self-attention element identifier has a focus direction; based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer.
[0061] In the embodiment of the present application, the information recognition module 12 receives the user's recognition requirements and converts them into self-attention elements, so that the self-attention recognizer can process the image in a targeted manner according to the task requirements. The recognition requirements are specific requests from users for the content of the image, such as identifying the texture features, age identification, or material analysis of antiques. After receiving the recognition requirements, the requirements are parsed using natural language processing (NLP) technology and converted into self-attention elements, each of which has a clear "focus", indicating the image feature area that the recognizer should focus on. For example, if the user requires the recognition of the texture features of the image, the "focus" of the self-attention element will instruct the model to prioritize the texture area of the image.
[0062] Based on these self-attention elements, the corresponding layers and nodes of the self-attention recognizer are triggered. First, the first recognition layer is triggered, which is mainly used to process the blurred texture in the image. Through the image recognition technology based on blurred texture, the blurred area in the image is analyzed and recognized, and the image is divided into multiple small blocks through the block technology for subsequent processing. This process uses the self-attention mechanism to focus on the key blurred areas in the image to accurately restore and enhance the texture information of these areas.
[0063] The block images are then input into the second recognition layer, where concurrent block image directional recognition based on the second, third, and fourth attention nodes is triggered. In this stage, each block of the image is processed through multiple parallel channels, focusing on the local features, spatial relationships, and cross-modal information of the image. A lateral interaction compensation mechanism is introduced to enhance the interaction between different channels and ensure that the various features of the image can be fully integrated and compensated during the processing. After the concurrent recognition results are processed, they will be passed to the next stage.
[0064] Finally, the concurrent recognition results are passed to the third recognition layer, where the recognition results are aggregated and reorganized according to the characteristics of the past and present. This process analyzes the association between historical and modern features in the image, integrates the various parts of the image into a complete recognition output, and thus determines the final information recognition result.
[0065] Furthermore, in the system provided by the embodiment of the application, based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer, and further includes:
[0066] Trigger the first recognition layer, identify the fuzzy texture of the antique scan image, and determine the block image; import the block image into the second recognition layer, trigger the concurrent block image directional recognition based on the second attention node, the third attention node and the fourth attention node, and introduce lateral interaction compensation, and output the concurrent recognition result; import the concurrent recognition result into the third recognition layer, aggregate and reorganize it according to the ancient and modern characteristics, and determine the information recognition result.
[0067] In the embodiment of the present application, in the process of the information recognition module, the image first enters the first recognition layer, the main task of which is to identify the fuzzy texture in the image and determine the block image. Since antique images usually have fuzzy areas due to factors such as age, wear or aging, these areas require special processing to extract effective features. In this layer, a convolutional neural network is used to extract fuzzy texture features in the image. Specifically, CNN uses a multi-layer convolution kernel to scan the image, so that when processing the fuzzy area, it can effectively enhance the details of the image, especially the edges and texture features of the fuzzy area. In this process, the boundaries in the image are further enhanced using techniques such as Canny edge detection, making the features of the fuzzy area more prominent. After the image processing is completed, the image is divided into multiple small blocks for subsequent analysis. The block method is adjusted according to the regularity or irregularity of the texture. If the image presents a regular texture, a non-uniform geometric block, such as an irregular rectangle or polygon, is used to better adapt to the repeated texture in the image; if the image presents an irregular texture, a uniform geometric block, such as a square or rectangle, is used to ensure that the size of each block is consistent, which is convenient for subsequent processing and feature extraction. Through this block division, each small block in the image can be analyzed independently, thus providing high-quality local input for the subsequent recognition layer.
[0068] After the block segmentation is completed, the block image is passed to the second recognition layer for further analysis. In this layer, the second, third and fourth attention nodes are triggered to process multiple feature dimensions of the image in parallel. The task of the second attention node is to extract local features in the image, such as carvings, patterns, cracks, etc. These areas are very important for identifying the authenticity, age and historical value of antiques. The third attention node focuses on spatial relationship modeling, analyzing the dependencies and relative positions between different areas in the image, such as the spatial relationship between the carved hands, faces and bases in the sculpture image. The fourth attention node is responsible for processing cross-modal information fusion, which enhances the recognition ability of the model by combining images and other types of information (such as text descriptions, expert labels, database knowledge, etc.). Through multi-channel parallel processing, the local features, spatial relationships and cross-modal information of the image are extracted simultaneously, providing multi-dimensional feature representation for subsequent recognition tasks.
[0069] A lateral interaction compensation mechanism is introduced between these parallel processing nodes to ensure that the features output by each node can be effectively integrated. The core of lateral interaction compensation is to dynamically adjust the weights according to the contribution and relevance of each node output. Specifically, the relevance of each node output to the current recognition task is calculated, and the node output is weighted according to these weights. For the recognition of texture features, the output weight of the second attention node is enhanced to ensure that texture information is processed first; for nodes with weak spatial relationships, their weights are reduced according to their relative contributions to ensure the accuracy and consistency of the final recognition results. For example, if the texture feature is very important for the current image recognition task (such as the pattern recognition of ancient pottery), the output of this node will be given a higher weight, so that it occupies a more important position in the final recognition result.
[0070] Next, after the concurrent processing of the second recognition layer, the concurrent recognition results will be passed to the third recognition layer. The task of this layer is to aggregate and reorganize the outputs of the first two layers and conduct a comprehensive analysis based on the ancient and modern characteristics. Specifically, the results are weighted and fused according to the historical features and modern features contained in the image. Historical features include textures, cracks, aging marks, etc. in antiques, which are crucial for judging the age and origin of antiques; while modern features mainly involve restoration marks in images, elements of modern aesthetics, etc., which are helpful for judging the restoration history and modern aesthetic style of objects. In this layer, the weights of these features are adjusted according to their time background and contribution to the final recognition task. Through the weighted summation method, the ancient and modern characteristics are fused to ensure the organic combination of historical and modern information. Specifically, higher weights are assigned to historical features because they are more critical for judging the age of antiques; while lower weights are assigned to modern features (such as restoration marks, etc.) because these features focus more on the current state of the object rather than the age judgment. In the weighted summation process, different weights are assigned to the features according to their importance, and finally a comprehensive recognition result that integrates historical and modern features is output.
[0071] Finally, after all levels of processing and fusion, the outputs from the first, second, and third recognition layers are fused through three layers of full connection to generate the final information recognition results, which include multiple dimensions such as the age, style, material, texture features, and restoration of the item. For example, when analyzing an ancient porcelain, the output recognition results include its age (such as "Ming Dynasty blue and white porcelain"), style (such as "blue and white decoration"), material (such as "porcelain clay"), texture (such as "cracks"), history (such as "natural wear"), and restoration (such as "modern restoration traces"). This information provides valuable data support for antique identification, style analysis, historical research, etc., and ultimately provides accurate and comprehensive recognition results through meticulous feature extraction and comprehensive analysis.
[0072] Furthermore, the system provided in the application embodiment also includes:
[0073] The information recognition result includes a first recognition result and a second recognition result. The first recognition result is obtained by reorganizing the ancient characteristics of the target antique, and the second recognition result is obtained by converting the first recognition result from ancient to modern times.
[0074] In an embodiment of the present application, the information recognition result includes a first recognition result and a second recognition result. The first recognition result is obtained based on the reorganization of the ancient characteristics of the target antique. Specifically, the target antique image is first preprocessed to enhance the image quality and extract historical features. Through technologies such as convolutional neural networks (CNN), features related to age and style in the image, such as texture, cracks, weathering marks, etc., are identified. These historical features help to infer the age and source of antique items. In order to accurately identify these features, the image is divided into blocks, and a suitable block division method is selected according to the regularity of the image texture. Each block is processed independently to improve the accuracy of feature extraction. Through these steps, the first recognition result is finally generated, which accurately reflects the historical characteristics in the antique image.
[0075] The second recognition result is obtained by converting the first recognition result from ancient to modern. At this stage, the historical features in the first recognition result are fused with modern restoration traces. The purpose of the ancient to modern conversion is to combine the historical features and modern restoration features in the antique image to generate a more comprehensive recognition result. Through cross-modal information fusion, the historical information in the image is combined with modern data sources such as text descriptions and expert labels to ensure that modern elements such as restoration traces can be recognized. Then, the historical and modern features are weighted fused, and different weights are assigned according to their importance in recognition. Through weighted summation, historical features usually have a higher weight, while modern restoration features have a relatively low weight, ensuring that the recognition result can reflect both the historical background of the object and its modern status. Finally, a second recognition result containing historical and modern features is output to provide users with a comprehensive analysis of the object, covering multiple dimensions such as age, style, and restoration status.
[0076] Furthermore, the system provided in the application embodiment also includes:
[0077] If the antique is damaged, the third recognition layer is triggered to perform self-attention restoration calculations on anchor point positioning and stripe patterns to determine fuzzy restoration information; based on the fuzzy restoration information and the first recognition result, the antique is three-dimensionally simulated to determine the simulated restoration of the antique and perform interface visualization.
[0078] In the embodiment of the present application, if the antique is damaged, the third recognition layer is triggered, and the damaged area is determined by anchor point positioning. First, edge detection is performed on the image, and the boundaries of the damaged area are automatically identified using algorithms such as Canny edge detection. By locating obvious feature points in the image, the starting and ending parts of the damaged area are found, and these points are called anchor points. Anchor points are obvious feature points in the damaged area, such as the starting point of a crack or a damaged edge. They serve as a reference for the subsequent repair process to ensure that the repaired part can be accurately docked with the surrounding area.
[0079] After the anchor point is located, the self-attention restoration calculation of the stripe pattern is performed. In this process, the texture of the damaged area of the image is restored, and the self-attention mechanism is used to calculate the similarity and relationship between the damaged area and the surrounding area. The self-attention mechanism restores the texture and details of the damaged part by focusing on the dependency between the damaged area and the intact area. Specifically, by analyzing the texture, shape and other features of the damaged area and the surrounding area in the image, the texture pattern that the missing part should have is inferred, and then the restoration is completed. This process pays special attention to the stripes and texture patterns of the damaged area to ensure that the restored image can be naturally connected to the original image.
[0080] After obtaining the fuzzy restoration information, a three-dimensional simulation is performed in combination with the first recognition result. The three-dimensional structure of the antique image is restored by mapping the repaired texture and shape data into a three-dimensional space using depth mapping technology. Specifically, the details in the fuzzy restoration information are combined with the historical features of the image, and a three-dimensional reconstruction algorithm is used to construct a virtual three-dimensional model of the antique item. This process accurately restores the three-dimensional form of the damaged area, especially cracks, defects, etc., by extracting information such as the shape, size, and surface details of the item, ensuring that the restored model can truly reflect the appearance of the antique.
[0081] Finally, the restored 3D model is presented to the user using interface visualization technology. The 3D restored model is rendered and displayed to the user through OpenGL technology, and the user can observe the restored antique by rotating, zooming, etc.
[0082] In the embodiments of the present application, in summary, the embodiments of the present application have at least the following technical effects:
[0083] This application introduces multiple self-attention dimensions, performs hierarchical structured deployment, constructs a self-attention identifier and deploys it in an embedded manner on an antique information platform; uploads antique scans according to the multi-threaded port of the antique information platform, pre-processes the antique scans and imports the self-attention identifier, performs self-attention meta-triggering and decision-making based on recognition needs, and determines information recognition results; wherein, the recognition steps based on the self-attention identifier include: by performing image recognition and block segmentation based on fuzzy textures, performing directional recognition under meta-concurrent self-attention for block images, aggregating and reorganizing the concurrent recognition results, and determining information recognition results, wherein meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations. The present invention solves the technical problems of low accuracy, inability to effectively extract global features, and inability to process complex damaged areas in the prior art in antique image recognition, and achieves the technical effect of improving antique image recognition accuracy and comprehensively extracting feature information by introducing multiple self-attention dimensions, performing hierarchical structured deployment, and concurrent self-attention mechanisms.
[0084] Embodiment 2 is based on the same inventive concept as the antique information recognition system that introduces the self-attention mechanism in the aforementioned embodiment. Figure 2 As shown, the embodiment of the present application provides an antique information recognition method that introduces a self-attention mechanism, and the method includes:
[0085] Introduce multiple self-attention dimensions, perform hierarchical structured deployment, build a self-attention recognizer and embed it in an antique information platform; upload antique scans according to the multi-threaded port of the antique information platform, pre-process the antique scans and import them into the self-attention recognizer, perform self-attention meta-triggering and decision-making based on recognition needs, and determine information recognition results; wherein, the recognition steps based on the self-attention recognizer include: by performing fuzzy texture-based image recognition and block segmentation, performing directional recognition under meta-concurrent self-attention on the block image, aggregating and reorganizing the concurrent recognition results, and determining the information recognition results, wherein the meta-concurrent self-attention at least includes block feature elements, inter-block relations, and cross-modal relations.
[0086] Furthermore, a multi-dimensional self-attention dimension is introduced, and the method further includes:
[0087] Introduce fuzzy texture attention, deploy the block mode, and set the first attention node; introduce feature attention and set the second attention node; introduce relational attention and set the third attention node; introduce cross-modal attention and set the fourth attention node; introduce global attention and set the fifth attention node.
[0088] Furthermore, hierarchical structured deployment is performed, and the method further includes:
[0089] According to the first attention node, a first recognition layer is constructed; the second attention node, the third attention node and the fourth attention node are deployed in parallel to construct a second recognition layer; according to the fifth attention node, a third recognition layer is constructed; the first recognition layer, the second recognition layer and the third recognition layer are fully connected in three layers to determine a self-attention recognizer.
[0090] Furthermore, the method further comprises:
[0091] The blocking mode is determined according to the texture features; if it is a regular texture, non-uniform geometric blocking is adopted, and if it is an irregular texture, uniform geometric blocking is adopted.
[0092] Furthermore, the method further comprises:
[0093] Through permission constraints, a multi-threaded port is deployed, wherein the multi-threaded port includes at least a user side and an expert side; the antique scan image is uploaded according to the multi-threaded port; wherein the uploading method includes antique scanning and uploading based on the built-in data acquisition card of the antique information platform, and image document retrieval and uploading.
[0094] Furthermore, the self-attention element triggering and decision making are performed based on the recognition needs, and the method further includes:
[0095] Recognition requirements are received, and the recognition requirements are converted into self-attention elements, wherein the self-attention element identifier has a focus direction; based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer.
[0096] Further, based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer, and the method further includes:
[0097] Trigger the first recognition layer, identify the fuzzy texture of the antique scan image, and determine the block image; import the block image into the second recognition layer, trigger the concurrent block image directional recognition based on the second attention node, the third attention node and the fourth attention node, and introduce lateral interaction compensation, and output the concurrent recognition result; import the concurrent recognition result into the third recognition layer, aggregate and reorganize it according to the ancient and modern characteristics, and determine the information recognition result.
[0098] Furthermore, the method further comprises:
[0099] The information recognition result includes a first recognition result and a second recognition result. The first recognition result is obtained by reorganizing the ancient characteristics of the target antique, and the second recognition result is obtained by converting the first recognition result from ancient to modern times.
[0100] Furthermore, the method further comprises:
[0101] If the antique is damaged, the third recognition layer is triggered to perform self-attention restoration calculations on anchor point positioning and stripe patterns to determine fuzzy restoration information; based on the fuzzy restoration information and the first recognition result, the antique is three-dimensionally simulated to determine the simulated restoration of the antique and perform interface visualization.
[0102] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. The processes depicted in the accompanying drawings do not necessarily require the specific order and continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0103] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0104] This specification and drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.
Claims
1. The antique information recognition system using the self-attention mechanism is characterized by: The system comprises: The recognizer deployment module is used to introduce multiple self-attention dimensions, perform hierarchical structured deployment, build a self-attention recognizer and embed it in the antique information platform; An information recognition module is used to upload the antique scanned image according to the multi-threaded port of the antique information platform, pre-process the antique scanned image and import it into the self-attention recognizer, trigger and make decisions based on the recognition requirements, and determine the information recognition result; Among them, the recognition steps based on the self-attention recognizer include: By performing fuzzy texture-based image recognition and segmentation, directional recognition under meta-concurrent self-attention is performed on the segmented images, the concurrent recognition results are aggregated and reorganized to determine the information recognition results, wherein the meta-concurrent self-attention at least includes segmentation feature elements, inter-block relations, and cross-modal relations.
2. The antique information recognition system introducing the self-attention mechanism as claimed in claim 1, characterized in that: Introducing multiple self-attention dimensions, including: Introduce fuzzy texture attention, deploy the block mode, and set the first attention node; Introduce feature attention and set the second attention node; Introduce relational attention and set the third attention node; Introduce cross-modal attention and set the fourth attention node; Introduce global attention and set the fifth attention node.
3. The antique information recognition system introducing the self-attention mechanism as claimed in claim 2 is characterized in that: Perform layered structured deployment, including: Constructing a first recognition layer according to the first attention node; The second attention node, the third attention node and the fourth attention node are deployed in parallel to construct a second recognition layer; Constructing a third recognition layer according to the fifth attention node; The first recognition layer, the second recognition layer and the third recognition layer are fully connected in three layers to determine a self-attention recognizer.
4. The antique information recognition system introducing the self-attention mechanism as claimed in claim 3 is characterized in that: Determine the blocking mode according to the texture features; Among them, if it is a regular texture, non-uniform geometric blocking is used, and if it is an irregular texture, uniform geometric blocking is used.
5. The antique information recognition system introducing the self-attention mechanism as claimed in claim 1, characterized in that: Deploy a multi-threaded port through authority constraints, wherein the multi-threaded port includes at least a user side and an expert side; Uploading the antique scan image according to the multi-threaded port; Among them, the uploading method includes scanning and uploading antiques based on the built-in data acquisition card of the antique information platform, and retrieving and uploading image documents.
6. The antique information recognition system introducing the self-attention mechanism as claimed in claim 3, characterized in that: Self-attention meta-triggering and decision-making guided by recognition needs, including: Receiving a recognition requirement, converting the recognition requirement into a self-attention element, wherein the self-attention element identifier has a focus direction; Based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer.
7. The antique information recognition system introducing the self-attention mechanism as claimed in claim 6, characterized in that: Based on the self-attention element, the self-attention identifier is layer-triggered and node-triggered within the layer, including: triggering the first recognition layer to recognize the fuzzy texture of the antique scan image and determine the segmented image; Importing the segmented image into the second recognition layer, triggering concurrent segmented image directional recognition based on the second attention node, the third attention node and the fourth attention node, introducing lateral interaction compensation, and outputting concurrent recognition results; The concurrent recognition results are introduced into the third recognition layer, aggregated and reorganized based on ancient and modern characteristics, and the information recognition results are determined.
8. The antique information recognition system introducing the self-attention mechanism as claimed in claim 7, characterized in that: The information recognition result includes a first recognition result and a second recognition result. The first recognition result is obtained by reorganizing the ancient characteristics of the target antique, and the second recognition result is obtained by converting the first recognition result from ancient to modern times.
9. The antique information recognition system introducing the self-attention mechanism as claimed in claim 8, characterized in that: The system further comprises: If the antique is damaged, the third recognition layer is triggered to perform self-attention restoration calculation for anchor point positioning and stripe pattern to determine fuzzy restoration information; According to the fuzzy restoration information and the first recognition result, the antique is simulated in three dimensions, the simulated restoration antique is determined and the interface is visualized.
10. The method for identifying antique information by introducing a self-attention mechanism is characterized in that: The method is performed by the antique information recognition system introducing the self-attention mechanism as described in any one of claims 1 to 9, comprising: Introducing multiple self-attention dimensions, performing hierarchical structured deployment, constructing a self-attention recognizer and embedding it in the antique information platform; According to the multi-threaded port of the antique information platform, the antique scanned image is uploaded, the antique scanned image is pre-processed and imported into the self-attention recognizer, and the self-attention element is triggered and decided based on the recognition demand to determine the information recognition result; Among them, the recognition steps based on the self-attention recognizer include: By performing fuzzy texture-based image recognition and segmentation, directional recognition under meta-concurrent self-attention is performed on the segmented images, the concurrent recognition results are aggregated and reorganized to determine the information recognition results, wherein the meta-concurrent self-attention at least includes segmentation feature elements, inter-block relations, and cross-modal relations.
Citation Information
Patent Citations
Bronze ware identification system
CN115393848A
Attention operation processing method and device
CN118585249A
Multi-modal information fusion site identification method and device
CN119025958A
Jade category identification method and device, electronic equipment and storage medium
CN119296098A
KR20240146429A
Cited By
Antique age determination method and system based on artificial intelligence and big data
CN120495797A
Antique age determination method and system based on artificial intelligence and big data
CN120495797B