A method and system for accurate identification of cultural features based on AIGC

By constructing a five-dimensional fusion path through AIGC and a joint spatial-temporal-semantic modeling mechanism, the problem of judging the polysemy of cultural graphics is solved, and the accurate attribution and interpretable recognition results of cultural graphics are achieved, which is applicable to the recognition of complex cultural graphics.

CN121303147BActive Publication Date: 2026-04-03HANGZHOU NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine the true cultural affiliation of cultural graphics that are ambiguous, regionally diverse, or have historical evolutionary characteristics, and lack interpretable reasoning paths.

Method used

By introducing AIGC and a joint spatial-temporal-semantic modeling mechanism, a five-dimensional fusion path of 'image features - semantic text - geographic tags - time tags - oral cultural corpus' is constructed to perform semantic interpretation and attribution reasoning of polysemous cultural graphics, including preprocessing, cultural image feature extraction, spatial consistency and temporal consistency matching, semantic similarity matching, and semantic path graph construction.

Benefits of technology

It improves the accuracy of cultural graphic attribution determination, enhances the ability to identify non-standard cultural symbols, provides interpretable recognition results, is suitable for complex graphic recognition needs, and expands the application scope of cultural graphic recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303147B_ABST
    Figure CN121303147B_ABST
Patent Text Reader

Abstract

This invention discloses a method for accurate cultural feature recognition based on AIGC, comprising: acquiring raw data; extracting cultural image features from the raw images using a pre-trained AIGC semantic generation model to generate an initial set of semantic candidate texts; performing spatial consistency matching and temporal consistency matching between the geographical information and spoken semantic text data in the raw data and the initial set of semantic candidate texts to obtain a matched first set of semantic candidate texts; performing semantic similarity matching between the spoken semantic text data in the raw data and the first set of semantic candidate texts to construct a final second set of semantic candidate texts; and constructing a corresponding semantic path graph based on the second set of semantic candidate texts and the corresponding raw data. This invention also provides a system for accurate cultural feature recognition based on AIGC. The method provided by this invention can effectively solve the problem of ambiguity judgment in cultural image recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a method and system for accurate identification of cultural features based on AIGC. Background Technology

[0002] With the advancement of digital protection of cultural heritage, the automatic identification and attribution of cultural images have become key technologies in digital museums and the cultural industry. Cultural image recognition aims to analyze cultural elements in images and determine their cultural system through computer vision and artificial intelligence technologies.

[0003] Patent document CN120671683A discloses a knowledge graph-based security control method. This method constructs a fire prevention database by crawling various types of data based on keywords using a web crawler. Natural language processing (NLP) is used to perform semantic analysis on the data in the fire prevention database. Word vector modeling and entity relation extraction are then used to process the semantically analyzed fire prevention database. An internal large-scale language model is constructed using the processed fire prevention database, and Leiden technology is employed to perform hierarchical clustering on this model. A multi-dimensional mapping model is established based on the spatial features of ancient buildings, cultural relics, and disaster-causing factors from the hierarchical clustering results. Based on this multi-dimensional mapping model, knowledge graph technology is used to construct a multimodal, dynamically related fire prevention knowledge graph. Based on user query keywords, a search mode is dynamically selected, and based on the search mode, security control results are generated according to the query keywords.

[0004] Patent document CN118626872A discloses a multimodal method and apparatus for text-image matching of cultural relics. This method determines the initial matching cultural relics based on the similarity between the semantic features of each candidate cultural relic and the semantic features of the target cultural relic, and achieves high-precision matching and association between text-image, image-image, and text-text by comparing the similarity between candidate sub-images and template sub-images. However, this type of method has significant shortcomings when dealing with cultural graphics that have polysemy, regional differences, or historical evolution characteristics. It is difficult to accurately determine their true cultural affiliation and lacks an interpretable reasoning path. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for accurate identification of cultural features based on AIGC, which can effectively solve the problem of ambiguity judgment in cultural graphic recognition.

[0006] To achieve the first objective of this invention, the following technical solution is provided: a method for accurate identification of cultural features based on AIGC, comprising the following steps:

[0007] Acquire raw data, including raw images, as well as corresponding geographic information and spoken semantic text data;

[0008] The original image is subjected to cultural image feature extraction by a pre-trained AIGC semantic generation model to generate a corresponding initial semantic candidate text set. Each semantic candidate text in the initial semantic candidate text set contains graphic meaning text, cultural attribution information and corresponding confidence score.

[0009] Using the geographic information and spoken semantic text data in the original data, spatial consistency matching and temporal consistency matching are performed with the initial set of semantic candidate texts to obtain the first set of semantic candidate texts after matching.

[0010] The oral semantic text data in the original data is used to perform semantic similarity matching with the first set of semantic candidate texts. Semantic candidate texts with similarity greater than the matching threshold are retained to construct the final set of second set of semantic candidate texts.

[0011] Based on the second set of semantic candidate texts and the corresponding original data, a corresponding semantic path graph is constructed.

[0012] This invention introduces AIGC and a space-time-semantic joint modeling mechanism, and constructs a five-dimensional fusion path of "image features - semantic text - geographic tags - time tags - oral cultural corpus" to achieve semantic interpretation and attribution reasoning of polysemous cultural graphics.

[0013] Specifically, the spoken semantic text data includes speech transcription data and / or historical annotation text data about cultural graphics in the original image.

[0014] Specifically, the original image needs to be preprocessed before being input into the AIGC semantic generation model. The preprocessing steps are as follows:

[0015] The texture features of the original image are enhanced using an image local contrast enhancement algorithm;

[0016] The preliminary outline of the target cultural graphic is extracted using a gradient-based edge detection algorithm;

[0017] The preliminary contour is closed by performing a region clustering algorithm to extract the complete graphic region;

[0018] The graphic region is standardized to a preset resolution and shape ratio to obtain a target image region containing the target cultural graphic.

[0019] Specifically, the AIGC semantic generation model is constructed by fusing a visual encoder and a language generator. The visual encoder is used to extract cultural image features from the input image, and the language generator is used to generate semantic candidate text data based on the cultural image features.

[0020] Specifically, the spatial consistency matching process is as follows:

[0021] The corresponding cultural distribution area is obtained based on the cultural affiliation information;

[0022] Calculate the geographical distance between the shooting location coordinates in the geographic tag data and the cultural distribution area;

[0023] Determine whether the geographical distance exceeds a preset spatial consistency threshold; if so, filter out the corresponding semantic candidate text data.

[0024] Specifically, the time consistency matching process is as follows:

[0025] Extract the time node information from each semantic candidate text data and spoken semantic text data;

[0026] The system calculates similarity based on time node information and determines whether the time similarity calculation result is lower than the preset time consistency threshold. If so, the corresponding semantic candidate text data is filtered out.

[0027] Specifically, the construction process of the semantic path graph is as follows:

[0028] Each semantic candidate text in the second set of semantic candidate texts is used as a graph node, the cultural image features corresponding to the original image extracted by the AIGC semantic generation model are used as the starting node, and the associated content between the cultural image features and the semantic candidate texts is used as the associated edge to construct a semantic path graph.

[0029] Specifically, when constructing a semantic path graph, if there are graph nodes that cannot establish effective semantic association edges through existing semantic candidate texts, the AIGC semantic generation model is called to generate breakpoint completion text for completion, using the target image region corresponding to the breakpoint node and the contextual semantic information in the constructed semantic path graph. The breakpoint completion text is then inserted into the corresponding breakpoint position in the semantic path graph.

[0030] To achieve the second objective of this invention, the following technical solution is provided: a precise cultural feature recognition system based on AIGC, comprising:

[0031] An image receiving module is used to receive raw images containing graphics of the target culture;

[0032] The data acquisition module is used to acquire geographic information and spoken semantic text data associated with the original image;

[0033] The time tag extraction module is used to extract time node information related to the target cultural graphic from spoken semantic text data;

[0034] The target image extraction module is used to extract the target image region of the target cultural graphic;

[0035] The candidate text generation module generates multiple semantic candidate texts based on the target image region of the input original image;

[0036] The first filtering module is used to perform spatial consistency and temporal consistency matching on multiple semantic candidate texts in order to construct a first semantic candidate text set.

[0037] The second filtering module is used to perform semantic matching on the first set of semantic candidate texts in order to construct the second set of semantic candidate texts.

[0038] The semantic path generation module generates a corresponding semantic path graph based on the second set of semantic candidate texts.

[0039] The results output module is used to visualize the semantic path graph.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] This invention improves the accuracy of cultural graphic attribution judgment by combining geographic-temporal constraints for semantic path filtering, solving the problems of poor accuracy and high ambiguity in semantic understanding and attribution judgment of traditional image recognition or text comparison methods. It supplements the gaps in mainstream corpora with oral semantics, enhancing the ability to recognize non-standard cultural symbols and effectively addressing the lack of systematic modeling of oral folk cultural materials and the failure of mainstream corpora to cover a large number of regional and marginal cultural semantics. It provides complete semantic paths and attribution links, avoiding "black box" results and making the recognition results interpretable. It can automatically construct historical evolution paths, suitable for cultural dissemination research scenarios, realizing semantic completion and evolutionary reasoning of cultural symbols. It has high scalability, applicable to complex graphic recognition needs such as marginal cultures, composite cultures, and polysemous cultures, expanding the application scope of cultural graphic recognition. Attached Figure Description

[0042] Figure 1 A block diagram illustrating a method for accurate identification of cultural features based on AIGC provided in this embodiment;

[0043] Figure 2 This embodiment provides a flowchart for extracting a target image region containing the target cultural graphic.

[0044] Figure 3 A flowchart for spatial consistency matching provided in this embodiment;

[0045] Figure 4 A flowchart for time consistency matching provided in this embodiment;

[0046] Figure 5This is a flowchart illustrating the construction process of the semantic path graph provided in this embodiment;

[0047] Figure 6 This is a flowchart of the breakpoint node completion process provided in this embodiment;

[0048] Figure 7 This is a block diagram of the cultural feature accurate recognition system provided in this embodiment. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0050] The same cultural image may have different meanings in different historical periods and regions. Relying solely on image features or semantic similarity matching is prone to misjudgment. In addition, existing technologies usually ignore structured and unstructured information such as geographic location data, historical period labels, and regional oral cultural semantics associated with the image, resulting in low accuracy and poor adaptability in attribution judgment.

[0051] To address the aforementioned issues, this embodiment introduces an AIGC and spatial-temporal-semantic joint modeling mechanism. By constructing a five-dimensional fusion path of "image features - semantic text - geographic tags - time tags - oral cultural corpus", it achieves semantic interpretation and attribution reasoning for polysemous cultural graphics.

[0052] like Figure 1 As shown, this embodiment provides a method for accurate identification of cultural features based on AIGC, including:

[0053] Acquire raw data, including raw images, as well as corresponding geographic information and spoken semantic text data;

[0054] The original image is subjected to cultural image feature extraction by a pre-trained AIGC semantic generation model to generate a corresponding initial semantic candidate text set. Each semantic candidate text in the initial semantic candidate text set contains graphic meaning text, cultural attribution information and corresponding confidence score.

[0055] Using the geographic information and spoken semantic text data in the original data, spatial consistency matching and temporal consistency matching are performed with the initial set of semantic candidate texts to obtain the first set of semantic candidate texts after matching.

[0056] The oral semantic text data in the original data is used to perform semantic similarity matching with the first set of semantic candidate texts. Semantic candidate texts with similarity greater than the matching threshold are retained to construct the final set of second set of semantic candidate texts.

[0057] Based on the second set of semantic candidate texts and the corresponding original data, a corresponding semantic path graph is constructed.

[0058] More specifically, in step S100, an original image containing the target cultural graphic is received.

[0059] The original image is obtained through image acquisition equipment or from a database. The original image can be a digital image in various formats, such as JPG, PNG, TIFF, etc., with a resolution of no less than 1024×768 pixels to ensure that the cultural graphic details in the image are clearly distinguishable. The original image may contain multiple cultural graphic elements, of which at least one is the target cultural graphic that needs to be identified.

[0060] In step S200, geographic tag data and spoken semantic text data associated with the original image are obtained. The geographic tag data includes the coordinate information of the shooting location, and the spoken semantic text data includes speech transcription data and / or historical annotation text data associated with the target cultural graphic.

[0061] Geographically labeled data refers to GPS coordinate data extracted from the EXIF ​​information of an image, or obtained through manual input of the shooting location information by the user. Geographically labeled data typically includes coordinate information in three dimensions: longitude, latitude, and altitude. Oral semantic text data is obtained in two ways: first, by converting the user's verbal description of the cultural graphic into text data using speech recognition technology; second, by extracting annotation text related to the cultural graphic from sources such as historical documents and museum explanatory texts. The length of oral semantic text data is typically between 50 and 500 words, containing descriptions of the graphic's appearance, historical background, and cultural significance.

[0062] In step S300, the first historical period information of the target cultural graphic association is extracted based on the oral semantic text data.

[0063] Extracting information from the first historical period involves using natural language processing technology to identify time-related keywords and phrases from spoken semantic text, such as dynasty names, reign titles, and specific years. Specifically, named entity recognition algorithms can be used to extract this time information and standardize it into a unified time representation format, such as "BC / BC + year" or "dynasty + period".

[0064] For example, extract "mid-Tang Dynasty" from "this is a pottery figurine from the mid-Tang Dynasty" and convert it into the standard time range of "713-755 AD".

[0065] In step S400, the target image region containing the target cultural graphic is extracted from the original image.

[0066] When generating semantic descriptions, AIGC models are highly sensitive to the semantic purity and boundary clarity of the input region. If the target region is not extracted and the entire image is directly input, it may contain background interference elements, such as text, watermarks and other images, as well as irrelevant visual information, such as the shooting angle and the photographer's shadow.

[0067] In step S500, the target image region is semantically generated using an AIGC-based semantic generation model to obtain multiple semantic candidate texts containing graphic meaning, cultural affiliation information, and corresponding confidence scores.

[0068] Specifically, the semantic generation model is a multimodal generation network that integrates a visual encoder and a language generator. The visual encoder is used to extract the semantic content representation vector of image sub-regions, and the language generator is used to generate semantic candidate text data based on the semantic content representation vector.

[0069] Specifically, the visual encoder uses a pre-trained convolutional neural network (such as ResNet-101 or VisionTransformer) as the backbone network to extract deep visual features of the target image region. The encoder contains multiple convolutional layers and attention mechanisms, which can capture local texture features and global structural features in the image. Its output is a high-dimensional feature vector with a dimension of 2048, representing the semantic content representation of the image.

[0070] The language generator is based on the Transformer decoder architecture, which contains 12 decoder layers, each with 8 attention heads and a hidden layer dimension of 768. It receives the feature vector output by the visual encoder as a conditional input, and fuses the visual features with the text features through a cross-attention mechanism. Then, it generates semantic text describing cultural graphics in an autoregressive manner. The generated text contains three main parts: a description of the meaning of the graphics (such as "double-eared pottery jar" and "flying apsara mural"), cultural attribution information (such as "Han Dynasty Central Plains culture" and "Buddhist art of Dunhuang Mogao Grottoes"), and a corresponding confidence score (a floating-point number between 0 and 1, representing the model's degree of confidence in the judgment).

[0071] In step S600, each semantic candidate text data is matched with the geographic tag data and the first historical period information for spatial and temporal consistency. Candidate text data that do not meet the preset spatial consistency threshold and temporal consistency threshold are filtered out to obtain the first semantic candidate set.

[0072] The fundamental purpose of spatial and temporal consistency matching is to improve the credibility of semantic candidate results and the accuracy of cultural attribution judgment. In cultural image recognition scenarios, a certain image may have multiple interpretations. For example, the same symbol may represent completely different cultural meanings in different regions or historical periods. If only image or semantic generation results are relied upon, ambiguity or incorrect attribution is very likely to occur. Therefore, by introducing the geographic label data and temporal information attached to the image, and performing dual spatial and temporal constraints on each semantic candidate text, it is possible to effectively eliminate those misidentified semantics that do not match the image background, and retain only the highly credible candidate results that match the actual geographical distribution and historical background. This provides more accurate basic semantic units for the subsequent construction of semantic path maps and avoids the generation of erroneous links in the cultural attribution reasoning process due to semantic deviation.

[0073] In step S700, the first semantic candidate set is matched with the spoken semantic text data for semantic similarity, and the data in the first semantic candidate set that does not meet the preset semantic similarity threshold is filtered out to obtain the second semantic candidate set.

[0074] Specifically, firstly, a pre-trained language model (such as BERT or RoBERTa) is used to encode the candidate text and the spoken semantic text into high-dimensional semantic vectors. Then, the cosine similarity between the two vectors is calculated to obtain a similarity score between 0 and 1. A preset semantic similarity threshold of 0.6 is set. If the calculated similarity is lower than this threshold, the candidate text is considered to be semantically mismatched with the spoken information and is removed from the first semantic candidate set. The remaining candidate texts constitute the second semantic candidate set.

[0075] In step S800, a semantic path graph representing the semantic link from image features to cultural affiliation is constructed based on the second semantic candidate set.

[0076] Constructing a semantic path graph representing the semantic link from image features to cultural attribution aims to organize scattered semantic candidate information into interpretable and inferable cultural attribution paths. In the second semantic candidate set, although relatively reliable semantic fragments have been obtained through multi-dimensional screening based on spatial, temporal, and semantic similarity, these fragments may still lack a clear logical order and hierarchical structure, making them difficult to use directly for judging cultural attribution. Therefore, by constructing a semantic path graph, semantic elements such as graphic meaning, regional association, and historical period are used as nodes, and their semantic relationships are used as edges to form a directed path structure from image content to cultural attribution. This not only clearly demonstrates the attribution reasoning logic of cultural symbols but also identifies and completes the breakpoints in the semantic chain, improving the interpretability and systematicity of the attribution results and providing a structured reasoning basis for the intelligent recognition of complex and polysemous cultural graphics.

[0077] In step S900, a structured recognition result containing graphic meaning, cultural affiliation, semantic path, and confidence score is generated based on the semantic path graph.

[0078] Specifically, the highest-confidence complete paths are extracted from the semantic path graph and organized into structured data in JSON format. The structured recognition result includes the following fields:

[0079] Graphic meaning: A description of the physical form and function of the graphic in the target culture, such as "blue and white porcelain bowl" or "stone Buddha statue";

[0080] Cultural attribution: The cultural system and historical background to which the target cultural graphic belongs, such as "Ming Dynasty Jingdezhen porcelain culture" or "Northern Wei Yungang Grottoes Buddhist art";

[0081] Semantic path: The reasoning path from graphic features to cultural affiliation, represented by a sequence of nodes and edges;

[0082] Confidence score: The degree of confidence the system has in the recognition result, with a value ranging from 0 to 1, and is calculated by weighted average of the confidence scores of each node in the path.

[0083] The final identification results can be directly used in applications such as digital cataloging of cultural heritage, generation of museum exhibit descriptions, and production of cultural and educational content.

[0084] like Figure 2 The flowchart shown is for extracting the target image region containing the target cultural graphic provided in this embodiment, including steps S410, S420, S430 and S440.

[0085] In step S410, the texture features of the original image are enhanced by an image local contrast enhancement algorithm.

[0086] Specifically, an adaptive histogram equalization method can be used to divide the image into multiple 8×8 pixel small regions, perform histogram equalization on each region separately, and then merge the processing results through bilinear interpolation to avoid artificial traces at the region boundaries. This step makes the texture details in the image more obvious, which is beneficial for subsequent edge detection.

[0087] In step S420, the preliminary outline of the target cultural graphic is extracted using a gradient-based edge detection algorithm.

[0088] Specifically, the process involves: first, calculating the horizontal and vertical gradients of the image, performing convolution operations using the Sobel operator, then calculating the gradient magnitude and direction, applying a non-maximum suppression algorithm to the gradient magnitude to retain local maximum gradient points, and forming a preliminary edge contour.

[0089] In this step, a multi-scale edge detection strategy can be adopted, that is, edge detection is performed under different Gaussian smoothing parameters, and then the results are fused, which can realize the processing of cultural graphics at different scales.

[0090] In step S430, a closure judgment is performed on the preliminary contour using a region clustering algorithm to extract the complete graphic region.

[0091] Specifically, the density-based clustering algorithm DBSCAN is used to cluster closely spaced edge points into one class, forming continuous contour segments. For regions with incomplete contours, the active contour model (Snake algorithm) is applied to complete the contours, and the contours gradually fit the graphic boundaries through the principle of minimizing energy.

[0092] In step S440, the graphic region is standardized to a preset resolution and shape ratio to obtain a target image region containing the target cultural graphic.

[0093] Specifically, the process involves: first, calculating the minimum bounding rectangle of the outline; then, cropping the region and adjusting it to a standard size using an affine transformation; maintaining the original aspect ratio of the graphic during the adjustment process; and adding black or white fill to the image edges if necessary. Finally, the image is standardized for brightness and contrast to meet the input requirements of the subsequent semantic generation model.

[0094] like Figure 3 The diagram shown is a flowchart of spatial consistency matching provided in this embodiment, including steps S610, S620 and S630.

[0095] In step S610, the corresponding cultural distribution area is obtained based on the cultural affiliation information.

[0096] In practical implementation, a cultural geographic distribution database can be established, containing the geographic distribution range of various cultural types. Each culture in the database stores its distribution area in the form of a geographic polygon, which is composed of a series of geographic coordinate points. For example, the distribution area of ​​"Han Dynasty Central Plains Culture" may be a polygonal area centered on Luoyang, covering parts of present-day Henan and Shaanxi provinces.

[0097] In step S620, the geographical distance between the shooting location coordinates and the cultural distribution area in the geographic tag data is calculated.

[0098] Specifically, the Haversine formula is used to calculate the spherical distance between two geographic coordinate points, in kilometers. For cultural distribution areas, the system calculates the shortest distance from the shooting location coordinates to the boundary of the distribution area. If the shooting location coordinates are located inside the distribution area, the distance is zero.

[0099] In step S630, it is determined whether the geographical distance exceeds a preset spatial consistency threshold. If so, the corresponding semantic candidate text data is filtered out.

[0100] The preset spatial consistency threshold varies depending on the cultural type. For example, for local cultures with strong regional characteristics, the threshold may be set at 50 kilometers, while for imperial cultures with a wide range of influence, the threshold may be set at 500 kilometers. If the calculated geographical distance exceeds the threshold of the corresponding cultural type, the semantic candidate text is considered to be spatially inconsistent with the shooting location and is removed from the candidate set.

[0101] like Figure 4 The diagram shown is a time consistency matching flowchart provided in this embodiment, including steps S640, S650 and S660.

[0102] In step S640, the second historical period information contained in each semantic candidate text data is extracted.

[0103] Specifically, the same named entity recognition algorithm used in step S300 to extract the first historical period information is employed to identify time-related expressions from the semantic candidate text. For example, "early Song Dynasty" is extracted from "this may be an early Song Dynasty celadon" as the second historical period information and standardized to the time range of "960-1050 AD".

[0104] In step S650, the time similarity between the information of the second historical period and the information of the first historical period is calculated.

[0105] In the specific calculation, the standardized time information is first converted into time intervals, and the degree of overlap between the two time intervals is calculated. The formula is: Time similarity = Overlapping time length / Union length of the two time intervals. For example, if the first historical period is "900-1000 AD" and the second historical period is "950-1050 AD", then the overlap is 50 years, the union length is 150 years, and the time similarity is 50 / 150 = 0.33.

[0106] In step S660, it is determined whether the time similarity calculation result is lower than the preset time consistency threshold. If so, the corresponding semantic candidate text data is filtered out.

[0107] The preset time consistency threshold is usually set to 0.2, which means that the two time intervals must have at least 20% overlap. If the calculated time similarity is lower than the threshold, the semantic candidate text is considered to be inconsistent with the spoken information in time and is removed from the candidate set.

[0108] After the dual screening of spatial and temporal consistency in steps S610 to S660 above, the remaining semantic candidate texts constitute the first semantic candidate set.

[0109] like Figure 5 The diagram shown is a flowchart of the semantic path graph construction provided in this embodiment, including steps S810 and S820.

[0110] In step S810, the graphic meaning, cultural affiliation information, and corresponding semantic relationship features are extracted based on the data fields of each semantic candidate text in the second semantic candidate set.

[0111] Specifically, natural language processing techniques are used to analyze each candidate text, identifying entities (such as object names, materials, and stylistic features) and relationships (such as "belongs to," "originates from," and "influences"). For example, from the sentence "This is a tri-colored pottery figurine made in Chang'an during the Tang Dynasty, reflecting the characteristics of cultural exchange along the Silk Road," the system extracts "tri-colored pottery figurine" (graphic meaning), "Chang'an during the Tang Dynasty" (cultural affiliation), and "reflects cultural exchange along the Silk Road" (semantic relationship features).

[0112] In step S820, based on preset semantic association rules, multiple semantic candidate texts with similar semantic relationship features or logical order are used as graph nodes to construct a graph structure of semantic path graph. The graph structure includes graph nodes and associated edges between nodes. Graph nodes represent intermediate semantic meanings from target image regions to cultural affiliation, and associated edges represent semantic transition relationships between nodes.

[0113] Specifically, a directed acyclic graph (DAG) structure is used to represent semantic paths. Starting from image feature nodes, the path passes through multiple intermediate semantic nodes and finally reaches the cultural attribution node. For example, a semantic path might be: "ceramic container" → "double-eared design" → "celadon glaze technique" → "Longquan kiln system of the Song Dynasty" → "Zhejiang culture of the Southern Song Dynasty". Each node is accompanied by a confidence score, and each edge is labeled with the semantic relationship type, such as "possesses features", "adopts techniques", "belongs to a school of thought", etc.

[0114] In the process of constructing the semantic path graph, although the semantic candidate texts have undergone multi-dimensional screening, due to the complexity and ambiguity of the cultural graphics themselves and the incompleteness of the data, there may still be some graph nodes that cannot establish an effective connection with the preceding and following semantics, which are called "breakpoints". These breakpoints will cause the semantic chain to be interrupted, making it impossible to smoothly infer the cultural affiliation of the image content, reducing the accuracy and interpretability of the recognition system. Therefore, identifying and judging these breakpoint nodes is a key step to realize subsequent semantic completion, restore the semantic path structure, and ensure the closed loop of the reasoning chain, thereby improving the system's ability to process complex cultural graphics and its recognition robustness.

[0115] like Figure 6 The diagram shown is a flowchart of the breakpoint completion process provided in this embodiment, including steps S830, S840 and S850.

[0116] In step S830, it is determined whether there are breakpoint nodes in the semantic path graph. Breakpoint nodes are graph nodes that cannot establish effective semantic association edges through existing semantic candidate texts.

[0117] In step S840, if a breakpoint node exists, the semantic generation model is invoked to generate breakpoint completion text based on the target image region corresponding to the breakpoint node and the contextual semantic information in the constructed semantic path graph.

[0118] Specifically, the system inputs the semantic information before and after the breakpoint node as prompts into the semantic generation model, guiding the model to generate semantic text that can connect the breakpoint. For example, if there is a breakpoint in the path from "bronze surface decoration" to "Western Zhou ritual culture", the system may generate "animal mask pattern with typical Western Zhou style" as an intermediate connecting node.

[0119] In step S850, the breakpoint completion text is inserted into the corresponding breakpoint position in the semantic path graph.

[0120] This embodiment also provides a precise cultural feature recognition system based on AIGC, used to execute the steps of the precise cultural feature recognition system based on AIGC provided in the above embodiment, such as... Figure 7 As shown, it includes:

[0121] The system comprises the following modules: an image receiving module for receiving the original image containing the target cultural graphic; a data acquisition module for acquiring geographic tag data and spoken semantic text data associated with the original image, wherein the geographic tag data includes the coordinates of the shooting location, and the spoken semantic text data includes speech-to-text data and / or historical annotation text data associated with the target cultural graphic; a time tag extraction module for extracting the first historical period information associated with the target cultural graphic based on the spoken semantic text data; a target image extraction module for extracting the target image region containing the target cultural graphic from the original image; and a candidate text generation module for generating semantics for the target image region using an AIGC-based semantic generation model to obtain multiple semantic candidate texts containing graphic meaning, cultural affiliation information, and corresponding confidence scores. The process involves several modules: a text selection module and a second selection module. The first selection module matches each semantic candidate text data with geographic tag data and first historical period information for spatial and temporal consistency, filtering out candidate text data that do not meet preset spatial and temporal consistency thresholds to obtain a first semantic candidate set. The second selection module matches the first semantic candidate set with spoken semantic text data for semantic similarity, filtering out data in the first semantic candidate set that do not meet preset semantic similarity thresholds to obtain a second semantic candidate set. The third selection module generates a semantic path graph based on the second semantic candidate set, representing the semantic link from image features to cultural affiliation. The fourth selection module outputs a result based on the semantic path graph, generating a structured recognition result that includes graphic meaning, cultural affiliation, semantic path, and confidence score.

[0122] More specifically, the target image extraction module extracts the target image region containing the target cultural graphic by: enhancing the texture features of the original image through a local image contrast enhancement algorithm; extracting the preliminary outline of the target cultural graphic through a gradient-based edge detection algorithm; performing a closure judgment on the preliminary outline through a region clustering algorithm to extract the complete graphic region; and standardizing the graphic region to a preset resolution and shape ratio to obtain the target image region containing the target cultural graphic.

[0123] The semantic generation model used in the candidate text generation module is a multimodal generation network that integrates a visual encoder and a language generator. The visual encoder is used to extract the image semantic content representation vector of the image sub-region, and the language generator is used to generate semantic candidate text data based on the image semantic content representation vector.

[0124] The first filtering module performs spatial consistency matching by: obtaining the corresponding cultural distribution area based on cultural affiliation information; calculating the geographical distance between the shooting location coordinates in the geographic tag data and the cultural distribution area; determining whether the geographical distance exceeds the preset spatial consistency threshold, and if so, filtering out the corresponding semantic candidate text data.

[0125] The first screening module performs time consistency matching by: extracting the second historical period information contained in each semantic candidate text data; calculating the time similarity between the second historical period information and the first historical period information; determining whether the time similarity calculation result is lower than the preset time consistency threshold, and if so, filtering out the corresponding semantic candidate text data.

[0126] The semantic path generation module constructs a semantic path graph by: extracting the graphic meaning, cultural affiliation information, and corresponding semantic relationship features from the data fields of each semantic candidate text in the second semantic candidate set; and constructing a graph structure of the semantic path graph by taking multiple semantic candidate texts with similar semantic relationship features or logical order as graph nodes based on preset semantic association rules. The graph structure includes graph nodes and the associated edges between nodes, where graph nodes represent the intermediate semantic meaning from the target image region to the cultural affiliation, and associated edges represent the semantic transition relationship between nodes.

[0127] The semantic path generation module also includes the following steps in constructing the semantic path graph: determining whether there are breakpoint nodes in the semantic path graph. Breakpoint nodes are graph nodes that cannot establish effective semantic association edges through existing semantic candidate texts. If breakpoint nodes exist, the semantic generation model is called to generate breakpoint completion text based on the target image region corresponding to the breakpoint node and the contextual semantic information in the constructed semantic path graph. The breakpoint completion text is then inserted into the corresponding breakpoint position in the semantic path graph.

[0128] Furthermore, the terms "upper," "lower," "inner," "outer," "front," and "rear" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0129] Of course, the above description is only a specific embodiment of the present invention and is not intended to limit the scope of the present invention. All equivalent changes or modifications made to the structure, features and principles described in the claims of the present invention should be included in the scope of the claims of the present invention.

[0130] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for accurate identification of cultural features based on AIGC, characterized in that, Includes the following steps: Acquire raw data, which includes raw images containing graphics of the target culture, as well as corresponding geographic information and oral semantic text data; The original image is subjected to cultural image feature extraction by a pre-trained AIGC semantic generation model to generate a corresponding initial semantic candidate text set. Each semantic candidate text in the initial semantic candidate text set contains graphic meaning text, cultural attribution information and corresponding confidence score. Using the geographic information and spoken semantic text data in the original data, spatial consistency matching and temporal consistency matching are performed with the initial set of semantic candidate texts to obtain the first set of semantic candidate texts after matching. The oral semantic text data in the original data is used to perform semantic similarity matching with the first set of semantic candidate texts. Semantic candidate texts with similarity greater than the matching threshold are retained to construct the final set of second set of semantic candidate texts. Based on the second set of semantic candidate texts and the corresponding original data, a corresponding semantic path graph is constructed. The construction process of the semantic path graph is as follows: Each semantic candidate text in the second set of semantic candidate texts is used as a graph node, the cultural image features corresponding to the original image extracted by the AIGC semantic generation model are used as the starting node, and the associated content between the cultural image features and the semantic candidate texts is used as the associated edge to construct a semantic path graph.

2. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, The spoken semantic text data includes speech transcription data and / or historical annotation text data about cultural graphics in the original image.

3. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, The original image needs to be preprocessed before being input into the AIGC semantic generation model. The preprocessing steps are as follows: The texture features of the original image are enhanced using an image local contrast enhancement algorithm; The preliminary outline of the target cultural graphic is extracted using a gradient-based edge detection algorithm; The preliminary contour is closed by performing a region clustering algorithm to extract the complete graphic region; The graphic region is standardized to a preset resolution and shape ratio to obtain a target image region containing the target cultural graphic.

4. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, The AIGC semantic generation model is constructed by fusing a visual encoder and a language generator. The visual encoder is used to extract cultural image features from the input image, and the language generator is used to generate semantic candidate text data based on the cultural image features.

5. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, The process of spatial consistency matching is as follows: The corresponding cultural distribution area is obtained based on the cultural affiliation information; Calculate the geographical distance between the shooting location coordinates in the geographic tag data and the cultural distribution area; Determine whether the geographical distance exceeds a preset spatial consistency threshold; if so, filter out the corresponding semantic candidate text data.

6. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, The time consistency matching process is as follows: Extract the time node information from each semantic candidate text data and spoken semantic text data; The system calculates similarity based on time node information and determines whether the time similarity calculation result is lower than the preset time consistency threshold. If so, the corresponding semantic candidate text data is filtered out.

7. The method for accurate identification of cultural features based on AIGC according to claim 1, characterized in that, When constructing a semantic path graph, if there are graph nodes that cannot establish effective semantic association edges through existing semantic candidate texts, the AIGC semantic generation model is called to generate breakpoint completion text for completion, using the target image region corresponding to the breakpoint node and the contextual semantic information in the constructed semantic path graph. The breakpoint completion text is then inserted into the corresponding breakpoint position in the semantic path graph.

8. A system for accurate identification of cultural features based on AIGC, characterized in that, The steps for performing the AIGC-based accurate cultural feature identification method as described in any one of claims 1 to 7 include: An image receiving module is used to receive raw images containing graphics of the target culture; The data acquisition module is used to acquire geographic information and spoken semantic text data associated with the original image; The time tag extraction module is used to extract time node information related to the target cultural graphic from spoken semantic text data; The target image extraction module is used to extract the target image region of the target cultural graphic; The candidate text generation module generates multiple semantic candidate texts based on the target image region of the input original image; The first filtering module is used to perform spatial consistency and temporal consistency matching on multiple semantic candidate texts in order to construct a first semantic candidate text set. The second filtering module is used to perform semantic matching on the first set of semantic candidate texts in order to construct the second set of semantic candidate texts. The semantic path generation module generates a corresponding semantic path graph based on the second set of semantic candidate texts. The results output module is used to visualize the semantic path graph.

Citation Information

Patent Citations

  • Safety control method based on knowledge graph

    CN120671683A

  • Cultural relic text image matching method and device based on multiple modes

    CN118626872A

  • File travel content analysis method and system based on semantic recognition

    CN120235732A