Method and system for extracting and visualizing hierarchical semantics of traditional chinese painting images based on large models
By employing multi-level structural features and semantic extraction methods, combined with visual language models and large language models, this study addresses the issues of limited semantic information extraction dimensions and historical context understanding in traditional Chinese painting images. It provides a platform for in-depth mining and exploration of multi-level semantics, thereby enhancing the analytical capabilities of traditional Chinese painting images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies for semantic understanding of traditional Chinese paintings have a relatively singular dimension in extracting semantic information, ignoring the co-occurrence relationship between facial expressions, body postures, and background objects, lacking a unified semantic feature modeling mechanism, making it difficult to deeply understand the historical context, and lacking an interactive analysis platform.
A large model-based approach is adopted, which extracts facial features, pose features and co-occurrence features from Chinese painting images through multi-level structural feature extraction and semantic extraction. Multi-level semantic extraction is performed by combining visual language model and large language model, and a three-layer feature clustering view, semantic association view and temporal evolution analysis view are constructed for visualization.
It enhances the understanding of the historical context and artistic intent of traditional Chinese painting images, provides a platform for in-depth mining and exploration of multi-level semantics, solves the problems of single-dimensional semantic information extraction and lack of unified modeling, and realizes in-depth analysis of traditional Chinese painting images.
Smart Images

Figure CN121366431B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data processing technology, and in particular to a method and system for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model. Background Technology
[0002] Traditional Chinese painting (Guohua) refers to paintings created using traditional Chinese tools and materials such as brushes, ink, and Xuan paper. It emphasizes artistic conception, spirit, and vividness, encompassing various subjects including landscapes, flowers and birds, and figures. Guohua is an important form of Chinese cultural inheritance and is widely used in art exhibitions, education, and cultural digitization. In Guohua works, figures, as a core visual element, are often used to express the artist's intentions, depict historical scenes, and convey social concepts; they often carry key semantic and contextual information within the entire painting. Therefore, accurately identifying and analyzing the figure elements and the multi-layered semantics contained within the paintings is of significant value in the process of intelligent understanding and deep digital modeling of Guohua.
[0003] To understand artistic images, current technologies primarily employ traditional computer vision methods, such as image classification, style recognition, and object detection. These methods largely rely on low-level visual features like color, texture, and contours, or use convolutional neural networks to extract deep representations for style analysis or image generation tasks. In analyzing human elements, some methods incorporate emotion recognition and pose estimation techniques, such as using OpenPose to extract keypoints to determine a person's action state, or identifying facial expressions and emotional signals through local features. These methods have achieved preliminary perception of the static structure of people in images.
[0004] Furthermore, the rise of Vision-Language Models (VLMs) in recent years has provided new methods for image-text information fusion. For example, models like CLIP, through large-scale image-text pairing training, enable images and natural language to align in a unified semantic space. These models are increasingly being applied to the classification and retrieval of museum collections, improving the semantic structuring capabilities and accessibility of artworks to some extent.
[0005] Meanwhile, in systems focused on interpreting image content, the DARK system provides researchers with tools to identify common compositional patterns in ancient images by detecting recurring motifs and symbolic structures in images; the InTaVia system constructs a transnational cultural knowledge graph, assisting experts in establishing semantic connections between historical figures, events, and visual art, supporting the reconstruction of cultural narratives across works; the Virtual Rosetta project attempts to visually cluster historical images, enhancing the immersive display effect with virtual reality; and the Arnold system focuses on the display and management of multi-dimensional image sets, helping users organize and classify cultural image resources. Although the above systems provide some analytical support in the temporal, spatial, and stylistic dimensions of artistic images, most of their functions still rely on external metadata of images (such as creation year, location of the work, artist information, etc.), lacking the ability to systematically model the semantic structure within images (such as facial expressions, action states, and the relationship between figures and backgrounds). Furthermore, these systems generally lack mechanisms for the structured extraction and semantic reconstruction of key visual elements (especially figures) in images, making it difficult to meet users' needs for in-depth exploration in areas such as cultural understanding, style analysis, and narrative reasoning.
[0006] In summary, the following technical challenges remain to be addressed in the semantic understanding and visualization analysis of traditional Chinese paintings: First, the extraction of semantic information is often limited to a single dimension, neglecting multi-layered structural features such as facial expressions, body postures, and co-occurrence relationships with background objects. Second, there is a lack of a unified semantic feature modeling mechanism, resulting in fragmented and unsystematic semantic expression that lacks hierarchy and completeness. Third, the semantics inherent in traditional Chinese paintings are highly dependent on specific historical and cultural backgrounds, and traditional methods struggle to achieve effective semantic construction through superficial visual features, leading to an understanding of traditional Chinese paintings detached from their historical context. Fourth, there is a lack of interactive analysis platforms and mechanisms to provide users with in-depth exploration paths for structural and semantic features, limiting users' understanding and application of potential semantic information in images. Summary of the Invention
[0007] The technical problem to be solved by this invention is to provide a method and system for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model, so as to solve the problems of the existing technology having a relatively single dimension of semantic information extraction, understanding of traditional Chinese painting detached from historical context, and lack of interactive analysis methods.
[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0009] This invention provides a method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model, comprising the following steps: extracting multi-level structural features from the traditional Chinese painting image, extracting facial features, posture features, and co-occurrence features of figures in the image; extracting multi-level semantics from the image based on a visual language model and a large language model to obtain the theme category and feature category of the image; and constructing and displaying a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view based on the results of multi-level structural feature extraction and multi-level semantic extraction.
[0010] Another aspect of the present invention provides a hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model, comprising: a multi-level structural feature extraction device for extracting multi-level structural features from traditional Chinese painting images, extracting facial features, posture features, and co-occurrence features of figures in the images; a multi-level semantic extraction device for extracting multi-level semantics from traditional Chinese painting images based on a visual language model and a large language model, obtaining the subject category and feature category of the images; and a multi-level semantic visualization device for constructing and displaying a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view based on the results of the multi-level structural feature extraction and the multi-level semantic extraction.
[0011] The beneficial technical effects of this invention are as follows: By extracting facial features, posture features, and co-occurrence features of figures in traditional Chinese paintings from three levels—micro, meso, and macro—a three-layer structural feature of the painting is obtained. Then, based on the extracted three-layer structural feature, a visual language model and a large language model are introduced to perform multi-level semantic extraction, resulting in the semantic features of the painting. This effectively solves the problem of the relatively single dimension of semantic information extraction in existing technologies and overcomes the technical limitations of traditional methods that rely on shallow visual features and are difficult to model deep semantics, thereby improving the ability to understand the historical context and artistic intent contained in traditional Chinese paintings. In addition, the extracted structural features and semantic features of the painting are visualized through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view, providing users with an interactive analysis platform and approach to assist users in conducting in-depth mining and exploration of multi-level semantics in traditional Chinese paintings. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating a method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model, according to one embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of the multi-level structural feature extraction process in one embodiment of the present invention;
[0014] Figure 3 A schematic diagram of a 68-key-point annotation system for faces;
[0015] Figure 4 This is a schematic diagram of the multi-level semantic extraction process in one embodiment of the present invention;
[0016] Figure 5 This is a schematic diagram illustrating the effect of the time evolution analysis view of the present invention in aggregation mode;
[0017] Figure 6 This is a schematic diagram illustrating the effect of the time evolution analysis view of the present invention in detail mode;
[0018] Figure 7 This is a schematic diagram of the structure of a hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model, according to one embodiment of the present invention. Detailed Implementation
[0019] To enable those skilled in the art to more clearly understand the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0020] This invention provides a method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model. For example... Figure 1 As shown, in one embodiment of the present invention, the method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model includes steps S10 to S30:
[0021] S10: Perform multi-level structural feature extraction on the traditional Chinese painting image, extracting facial features, posture features, and co-occurrence features of the figures in the painting image.
[0022] Specifically, step S10 extracts facial features, posture features, and co-occurrence features of figures in the traditional Chinese painting image from three levels: micro, meso, and macro, respectively, to obtain a three-layer structural feature of the traditional Chinese painting image: at the micro level, facial features are based on the facial key point detection model to identify the geometric features of the facial region, and then key parameters related to the emotions and facial expressions of the figure, such as the tilt angle of the eyebrows, the opening and closing of the eyes, and the shape of the mouth, are calculated and extracted; at the meso level, posture features are obtained by obtaining the coordinates of key points of the human skeletal structure through the human key point estimation model, and then the length and direction of the bone connection are calculated to depict the body movements and behavioral features of the figure; at the macro level, object detection model is used to detect objects in the traditional Chinese painting image, and then object entity pairs are constructed. Finally, the object entity names and the normalized relative distance between object entities are used as the co-occurrence features at the macro level.
[0023] like Figure 2As shown, the multi-level structural feature extraction in step S10 is divided into three stages: facial feature extraction, pose feature extraction, and co-occurrence feature extraction.
[0024] 1.1 Facial Feature Extraction
[0025] The facial feature extraction process at the microscopic level mainly consists of facial landmark detection and facial feature calculation:
[0026] 1.1.1 First, 42 key points on the face are selected as target key points from the 68-key-point annotation system for the human face. For example... Figure 3 As shown, the 68-key-point annotation system for the face accurately describes the geometric structure of the face by annotating 68 key points in specific facial regions (such as eyes, eyebrows, nose, mouth, and contours). The 42 key points include key points 18-27 for describing eyebrows, key points 37-48 for describing eyes, and key points 49-68 for describing the mouth. This embodiment of the invention selects 42 key points from the 68-key-point annotation system as target key points, which can reduce the computational load in subsequent facial key point detection and improve detection efficiency. Of course, in other embodiments of the invention, the 68 key points of the face can also be directly used as target key points for subsequent facial key point detection.
[0027] 1.1.2 Then, a facial keypoint estimation model is used to detect facial keypoints in the traditional Chinese painting image to obtain the facial keypoints of the figures in the painting image. Specifically, using the target keypoints as the detection target, a facial keypoint estimation model is used to detect facial keypoints in the traditional Chinese painting image to obtain the facial keypoints of the figures in the painting image. Facial keypoint estimation models include Openpose, FAN, Sapiens-1b-pose-133-keypoints, etc. In this embodiment of the invention, Sapiens-1b-pose-133-keypoints is used for facial keypoint detection.
[0028] 1.1.3 Finally, the coordinates of the acquired facial key points are used to calculate facial features, resulting in the final facial features. Specifically, after acquiring the facial key points, the coordinates of the key points are used to calculate the tilt of the eyebrows, eyes, and mouth. The tilt (angular features) of the eyebrows, eyes, and mouth are uniformly represented as follows: ,in, The orientation angles between selected facial key points are calculated using the coordinates between them. The degree of expression, such as the opening and closing of the eyes and mouth, is calculated using the distances between these key points. The degree of expression is expressed as... ,in, These represent the size of the eyes when they open and close, and the size of the mouth when they open and close. The value is taken as the vertical distance between selected facial key points; the final facial feature is represented as: .
[0029] 1.2 Pose Feature Extraction
[0030] The meso-level pose feature extraction process mainly consists of human keypoint detection and pose feature calculation:
[0031] 1.2.1 First, a human keypoint estimation model is used to detect human keypoints in the traditional Chinese painting image, obtaining the skeletal keypoints of the figures in the image. Specifically, 17 human keypoints are used as detection targets. The human keypoint estimation model is used to detect human keypoints in the traditional Chinese painting image, obtaining the skeletal keypoints of the figures in the image. Each skeletal keypoint is represented as follows: ,in , and These represent the horizontal and vertical coordinates, respectively. Human keypoint detection refers to detecting important joints (such as the head, shoulders, elbows, and knees) in an image and inferring the person's posture or movement state based on these points. Human keypoint estimation models include OpenPose, PoseNet, AlphaPose, and Sapiens-1b-pose-133-keypoints. In this embodiment of the invention, Sapiens-1b-pose-133-keypoints is used for human keypoint detection.
[0032] 1.2.2 Then, the 17 human body key points (skeleton key points) were used to calculate the pose features and obtain the skeleton length and skeleton orientation.
[0033] Specifically, bone length is determined by connecting bone keypoints. and The Euclidean distance between them is used for calculation, and the formula is:
[0034] ;
[0035] Where E represents the predefined connection relationship of skeletal key points.
[0036] Bone orientation is calculated by determining the orientation angle between the connected bone keypoint pairs and the horizontal axis, using the following formula:
[0037] ;
[0038] The final pose feature is represented as:
[0039] ;
[0040] 1.3 Co-occurrence Feature Extraction
[0041] 1.3.1 First, an object detection model is used to perform entity detection on the traditional Chinese painting image to obtain the objects in the image. Specifically, object detection is a computer vision task used to locate and classify different objects in an image. Object detection models include YOLO, DINO-X, Grounding DINO, etc. In this embodiment of the invention, DINO-X is used to perform entity detection on the traditional Chinese painting image.
[0042] 1.3.2 Next, the acquired objects are filtered to remove non-holistic objects and retain holistic objects. Specifically, for the acquired objects, individual parts such as hands and eyes are removed, and only complete holistic entities such as human bodies, animals, plants, and buildings are retained to ensure the effectiveness of the extracted object information and to ensure that key elements in the painting can be effectively extracted.
[0043] 1.3.3 Then, the filtered object entities are paired into object entity pairs, and frequent but insufficient informational object entity pairs are filtered out. Specifically, after pairing the filtered object entities into object entity pairs, the term frequency-inverse document frequency method is used to filter out frequent but insufficient informational combinations (object entity pairs), retaining only the top 70% of object entity pairs.
[0044] 1.3.4 Finally, Euclidean distance is used to represent the spatial relationship between pairs of object entities, and the final co-occurrence feature is represented as follows:
[0045] ;
[0046] in, and The names of the objects in a traditional Chinese painting. for and The normalized relative distance between them This is the set of object entities obtained through entity detection.
[0047] S20: Based on visual language models and large language models, multi-level semantic extraction is performed on traditional Chinese painting images to obtain the subject category and feature category of the images.
[0048] Specifically, step S20, based on the structural features extracted in step S10, introduces a visual language model and a large language model for semantic modeling, and extracts semantic features of the traditional Chinese painting image according to the theme feature classification system obtained from the modeling. Semantic modeling includes two steps: semantic generation based on the visual language model and semantic clustering modeling based on the large language model. In the semantic generation step based on the visual language model, the original traditional Chinese painting image, the structural features extracted in step S10, and externally acquired metadata of the traditional Chinese painting image are input into the visual language model to automatically generate natural language descriptions of the face, body, and spatial layout of the figures. In the semantic clustering modeling step based on the large language model, the semantic themes of the traditional Chinese painting image are clustered using the large language model to obtain a theme feature classification system. After constructing the theme feature classification system, the semantic features of the traditional Chinese painting image are extracted according to the modeled theme feature classification system, describing the semantics of the traditional Chinese painting image from multiple levels: micro, meso, macro, and theme.
[0049] like Figure 4 As shown, step S20, namely, performing multi-level semantic extraction on the Chinese painting image based on the visual language model and the large language model to obtain the theme category and feature category of the Chinese painting image, includes the following steps S21 to S23:
[0050] S21: Based on the multi-level structural feature extraction results, a natural language description of the traditional Chinese painting image is generated using a visual language model. Specifically, the traditional Chinese painting image, metadata (including painting name, author, creation year, category label, etc.), and the structural features extracted in step S10 are used as input to the visual language model. Then, the visual language model, acting as the artist, provides a natural language description of facial, body, and spatial features, and generates a summary of the overall visual content of the traditional Chinese painting. A visual language model is a pre-trained multimodal model capable of simultaneously understanding image and text information. This type of model is trained using large-scale image-text alignment data and has the ability to establish a correspondence between image content and natural language. Its main functions include image content recognition, text generation, image-text matching, and cross-modal reasoning. Visual language models include CLIP, DALL·E series, BLIP, Flamingo, LLaVA, DeepSeek-VL, GPT series, etc. In this embodiment, the visual language model used is GPT-4o-mini.
[0051] S22: Based on the natural language description of traditional Chinese painting images, a large language model is used to cluster and model the semantic topics in the images to obtain a topic feature classification system. Specifically, the images and their natural language descriptions are used as input to the large language model. The model extracts the semantic topics from the image data and performs semantic topic clustering to generate corresponding topic labels. Finally, the topic feature classification system is constructed by establishing a mapping relationship between semantic topics and feature categories.
[0052] Large language models are large-scale natural language processing models trained on massive text corpora, possessing powerful language understanding and generation capabilities. They can integrate contextual and inference information, transforming structured input into semantic descriptions. Large language models include GPT-4o-mini, Claude 2 / 3, PaLM 2, Llama 2 / 3, Kosmos-2, etc. In this embodiment, GPT-4o-mini is used.
[0053] In this embodiment, to address the context length limitation issue, step S22 employs a bottom-up batch processing strategy, first performing local clustering and then merging and integrating the results to ultimately construct a clear and reusable topic feature classification system. Specifically, this includes the following steps:
[0054] 2.2.1 Random batch sampling and candidate theme extraction are performed on the traditional Chinese painting image data. Specifically, n paintings are randomly selected, and the natural language descriptions generated in step S21 are used as input to the large language model. The large language model then extracts the themes of the nth paintings. Candidate themes and feature categories in batches of traditional Chinese painting image data.
[0055] 2.2.2 The LLM (Local Language Model) is invoked to perform semantic topic clustering on each batch of traditional Chinese painting images, generating cluster labels and classification rationale. Specifically, the text description of each batch of traditional Chinese painting images is input into the large language model, which then infers and analyzes the commonalities and differences among similar images, generating several candidate topic labels (i.e., cluster labels), and providing the clustering logic and rationale for each topic category.
[0056] 2.2.3 Summarize and re-cluster all candidate topics extracted from all batches. Specifically, the candidate topics extracted from all batches of traditional Chinese painting image data are summarized and input into a large language model for secondary clustering to generate corresponding topic labels and classification reasons. Merging and analyzing the topic labels generated by the secondary clustering with the topic labels generated by the semantic topic clustering of each batch of traditional Chinese painting images can improve the consistency, accuracy, and coverage of the clustering results.
[0057] 2.2.4 Merge similar topic tags and remove ambiguous topic tags to construct an initial topic feature classification system. Specifically, the topic tags generated by secondary clustering are integrated with the topic tags generated by semantic topic clustering of each batch of Chinese painting images. Topic tags with similar semantics or overlapping expressions are merged, and topic tags with ambiguous expressions or no practical meaning are removed, so that the subsequent topic feature classification system is clearer and more explicit. Finally, a mapping relationship is established between the topic tags obtained after merging and removal and the corresponding feature categories to construct the initial topic feature classification system.
[0058] 2.2.5 The initial topic feature classification system was optimized and updated through expert manual verification and model-assisted optimization to obtain the final topic feature classification system. Through expert manual verification and model-assisted optimization, a topic feature classification system with broad coverage and clear semantic boundaries was constructed as a standardized image label reference.
[0059] Step S22 constructs a clear and reusable topic feature classification system, effectively solving the problem of the lack of a unified semantic feature modeling mechanism in existing technologies.
[0060] S23: Based on the aforementioned subject feature classification system, the traditional Chinese painting images are automatically classified and interpreted to obtain their subject category and feature category. Specifically, the subject feature classification system, the structural features of each traditional Chinese painting image, and the text description are input into a large language model. The large language model automatically infers and assigns the most matching subject category and feature category to each traditional Chinese painting image, while generating a clear classification explanation. Finally, the subject category and feature category of the traditional Chinese painting image are output for subsequent analysis and display. The structural features of the traditional Chinese painting images are extracted in step S10, including facial features, pose features, and co-occurrence features; the text description of the traditional Chinese painting images is generated by the visual language model in step S21.
[0061] S30: Based on the multi-level structural feature extraction results and multi-level semantic extraction results, construct and display a three-level feature clustering view, semantic association view, time evolution analysis view and detail presentation view.
[0062] Specifically, step S30 visualizes the structural features extracted in step S10 and the semantic features of the traditional Chinese painting image extracted in step S20. It displays the structural and semantic features of the painting image through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view, assisting users in deeply exploring and mining the multi-layered semantics of the painting image. Specifically, the three-layer feature clustering view visualizes and clusters the three types of structural features extracted—facial features, pose features, and co-occurrence features—and provides natural language search functionality; the semantic association view reduces the dimensionality of the relationships between multi-layered semantic features; the temporal evolution analysis view constructs a connection between the painting's semantics and its historical background, helping users understand the evolutionary trends of structural features and semantic themes over time; and the detail presentation view allows for in-depth viewing of the visual and semantic details of a single painting.
[0063] 3.1 Three-layer feature clustering view:
[0064] The three-layer feature clustering view is used to present the distribution of the three-layer structural features (facial features, pose features, and co-occurrence features) of the Chinese painting image extracted in step S10 in the embedding space, and supports natural language search and interactive exploration operations.
[0065] The three-layer feature clustering view includes three sub-views: facial feature clustering, pose feature clustering, and co-occurrence feature clustering. All three sub-views are clustering scatter plots. Specifically, the t-SNE dimensionality reduction method is used to project high-dimensional features (facial or pose features) onto a two-dimensional plane, generating either the facial feature clustering sub-view or the pose feature clustering sub-view. In the facial feature clustering and pose feature clustering sub-views, spatial proximity is used to represent the structural feature similarity between images. In the co-occurrence feature clustering sub-view, the similarity between image pairs is first calculated by combining Jaccard similarity and normalized Euclidean distance, and then the t-SNE dimensionality reduction method is used to reduce the dimensionality of the co-occurrence features to a two-dimensional plane to draw the clustering scatter plot.
[0066] In the three-layer feature clustering view, users can hover or zoom to view representative works of a specific cluster, and it supports linked highlighting operation. That is, if you circle the scatter points in any of the three sub-views, the corresponding scatter points in the selected scatter point will also be highlighted in the remaining two sub-views.
[0067] The three-layer feature clustering view supports natural language search. Specifically, the three-layer feature clustering view is equipped with a natural language search box. Users can enter natural language text containing semantic intent in the natural language search box. In response to the user's input operation in the natural language search box, the query intent is parsed by the large language model and automatically mapped to the three-layer structural feature dimensions obtained by multi-level structural feature extraction for feature matching. The most relevant set of Chinese painting images is returned according to semantic similarity.
[0068] 3.2 Semantic Relationship View:
[0069] The structure and functions of semantically related views are described below:
[0070] 3.2.1 Chord Diagram: Represents the proportion and correlation of micro, meso, and macro features within a cluster. Specifically, the chord diagram includes an outer ring and internal scattered points. Each arc segment of the ring represents a specific feature category within that structural level (e.g., "solemn" expression in the facial feature level, or "kneeling" posture in the posture level). The longer the arc, the greater the proportion of that feature at the current level, meaning more works possess this feature. Each scattered point in the diagram represents a traditional Chinese painting, distributed in two-dimensional space, reflecting the structural and semantic dimensions of the artwork. The color of the scattered points represents the thematic category of the artwork, such as "Lady Painting," "Taoist and Buddhist Figure Painting," and "Portrait Painting."
[0071] 3.2.2 Topic Statistics Panel: A large language model is used to embed the semantic labels and classification reasons for each image, calculate their semantic similarity at the topic level, and map it to a two-dimensional plane using spatial distance. The color of the bars represents the topic category, and the length of the bars indicates the number of Chinese paintings included in the current topic category.
[0072] 3.2.3 Connections between Chords: For clusters of structural features that are similar, connections with adjustable thresholds visualize their commonality. To avoid visual clutter, only the top 25% of highly similar connections are displayed by default. Specifically, a connection extending from the outer ring of a chord to another chord represents the association between the current feature and two other structural levels. For example, a specific pose feature (such as "looking back") may frequently co-occur with a specific facial expression (such as "smiling") or a co-occurring object (such as "plum blossom"). The thickness (or number) of the connections indicates the strength of the association, helping users identify structural commonalities or differences across levels.
[0073] The semantic association view is used to identify the commonalities and differences in the structural patterns of characters under different themes, making it easier for users to conduct association analysis between themes and structural features.
[0074] 3.3 Time Evolution Analysis View:
[0075] The time evolution analysis view combines historical timeline information to show the temporal evolution trends of structural features and semantic themes, including the following two modes:
[0076] Aggregation Mode: Aggregates and displays paintings with time intervals less than a preset threshold. The area and saturation of the dots represent the number of aggregated paintings. This mode is suitable for macro-level overviews, such as... Figure 5 As shown.
[0077] Detailed Mode: Displays each painting by its corresponding year in a vertical line format. Hovering over the painting allows you to view its title and the original image, suitable for detailed analysis and zooming in on specific areas. For example... Figure 6 As shown, each vertical line in the image represents a painting. The timeline feature column on the left lists the top 10 objects in the selected traditional Chinese painting, such as ladies, courtyards, garden rocks, maids, and trees. A dot on a vertical line indicates that the object to the left of that dot appears in the painting. For example, if there are four dots on the third vertical line, it means that the painting represented by that line contains four objects: ladies, courtyards, maids, and trees.
[0078] 3.4 Detailed View:
[0079] The detailed view displays the original Chinese painting image on the left and a natural language description of the image on the right. Clicking on the image on the left will take you to the metadata display page, which shows the metadata of the Chinese painting, including the year of creation, location of the work, and artist information.
[0080] For steps S10-S30, after receiving the traditional Chinese painting image, the facial features, posture features, and co-occurrence features of the figures in the painting image are first extracted from three levels: micro, meso, and macro, respectively, to obtain the three-layer structural features of the painting image. Then, based on the extracted three-layer structural features, a visual language model and a large language model are introduced to perform multi-level semantic extraction to obtain the semantic features of the painting image. Finally, the extracted structural features and semantic features of the painting image are visualized. The structural and semantic features of the painting image are displayed through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view, assisting users in conducting in-depth multi-level semantic mining and exploration of the painting image.
[0081] The hierarchical semantic extraction and visualization method for traditional Chinese painting images based on a large model in this invention extracts facial features, posture features, and co-occurrence features of figures in traditional Chinese painting images from three levels: micro, meso, and macro, respectively, to obtain the three-layer structural features of the images. Then, based on the extracted three-layer structural features, a visual language model and a large language model are introduced to perform multi-layer semantic extraction, obtaining the semantic features of the images. This effectively solves the problem of the relatively single dimension of semantic information extraction in existing technologies and overcomes the technical limitations of traditional methods that rely on shallow visual features and are difficult to model deep semantics. This enhances the understanding of the historical context and artistic intent contained in traditional Chinese paintings. Furthermore, it constructs a clear and reusable thematic feature classification system, effectively addressing the lack of a unified semantic feature modeling mechanism in existing technologies. Finally, it visualizes the extracted structural features and semantic features of the extracted traditional Chinese painting images, displaying the structural and semantic features of the images through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view. This provides users with an interactive analysis platform and approach, assisting them in conducting in-depth multi-level semantic mining and exploration of traditional Chinese painting images.
[0082] This invention also provides a system for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model. For example... Figure 7 As shown, in one embodiment of the present invention, the hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model includes a multi-level structural feature extraction device 10, a multi-level semantic extraction device 20, and a multi-level semantic visualization device 30. Detailed descriptions of each device are as follows:
[0083] The multi-level structural feature extraction device 10 is used to extract multi-level structural features from Chinese painting images, extracting facial features, posture features, and co-occurrence features of figures in the Chinese painting images.
[0084] Specifically, the multi-level structural feature extraction device 10 extracts facial features, posture features, and co-occurrence features of figures in traditional Chinese paintings from three levels: microscopic, mesoscopic, and macroscopic, respectively, to obtain three-level structural features of the paintings: at the microscopic level, facial features are identified based on a facial keypoint detection model to identify the geometric features of the facial region, and then key parameters related to the emotions and facial expressions of the figures, such as the tilt angle of the eyebrows, the opening and closing of the eyes, and the shape of the mouth, are calculated and extracted; at the mesoscopic level, posture features are obtained by obtaining the coordinates of key points of the human skeletal structure through a human keypoint estimation model, and then calculating the length and direction of the bone connections to depict the body movements and behavioral characteristics of the figures; at the macroscopic level, object detection models are used to detect objects in the paintings, and then object pairs are constructed. Finally, the names of the objects and the normalized relative distances between the objects are used as the co-occurrence features at the macroscopic level.
[0085] The multi-level semantic extraction device 20 is used to perform multi-level semantic extraction on Chinese painting images based on visual language models and large language models to obtain the subject category and feature category of the Chinese painting images.
[0086] Specifically, the multi-level semantic extraction device 20, based on the structural features extracted by the multi-level structural feature extraction device 10, introduces a visual language model and a large language model for semantic modeling, and extracts the semantic features of the traditional Chinese painting image according to the theme feature classification system obtained from the modeling. Semantic modeling includes two steps: semantic generation based on the visual language model and semantic clustering modeling based on the large language model. In the semantic generation step based on the visual language model, the original traditional Chinese painting image, the structural features extracted by the multi-level structural feature extraction device 10, and externally acquired metadata of the traditional Chinese painting image are input into the visual language model to automatically generate natural language descriptions of the face, body, and spatial layout of the figures. In the semantic clustering modeling step based on the large language model, the large language model is used to cluster the semantic themes and structural features of the traditional Chinese painting image to obtain a theme feature classification system. After constructing the theme feature classification system, the semantic features of the traditional Chinese painting image are extracted according to the modeled theme feature classification system, describing the semantics of the traditional Chinese painting image from multiple levels: micro, meso, macro, and theme.
[0087] The multi-level semantic visualization device 30 is used to construct and display a three-level feature clustering view, a semantic association view, a time evolution analysis view, and a detail presentation view based on the multi-level structural feature extraction results and the multi-level semantic extraction results.
[0088] Specifically, the multi-level semantic visualization device 30 works in conjunction with the multi-level structural feature extraction device 10 and the multi-level semantic extraction device 20 to visualize the structural features extracted by the multi-level structural feature extraction device 10 and the semantic features of the traditional Chinese painting image extracted by the multi-level semantic extraction device 20. This visualization is presented through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view, assisting users in deeply exploring and mining the multi-level semantics of the traditional Chinese painting image. Specifically, the three-layer feature clustering view visualizes and clusters the traditional Chinese painting image based on three types of structural features extracted: facial features, pose features, and co-occurrence features, and provides a natural language search function; the semantic association view reduces the dimensionality of the relationships between multi-level semantic features; the temporal evolution analysis view constructs a connection between the semantics of the painting and its historical background, helping users understand the evolutionary trends of structural features and semantic themes over time; and the detail presentation view allows for in-depth viewing of the visual and semantic details of a single painting.
[0089] The hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model provided in this invention, after receiving a traditional Chinese painting image, extracts facial features, posture features, and co-occurrence features of figures in the image from three levels: micro, meso, and macro, respectively, to obtain the three-layer structural features of the image. Then, based on the extracted three-layer structural features, a visual language model and a large language model are introduced to perform multi-layer semantic extraction, obtaining the semantic features of the image. This effectively solves the problem of the relatively single dimension of semantic information extraction in existing technologies and overcomes the technical limitations of traditional methods that rely on shallow visual features and are difficult to model deep semantics, thereby improving the ability to understand the historical context and artistic intentions contained in traditional Chinese paintings. In addition, the extracted structural features and semantic features of the image are visualized through a three-layer feature clustering view, a semantic association view, a temporal evolution analysis view, and a detail presentation view, providing users with an interactive analysis platform and approach to assist users in conducting in-depth mining and exploration of multi-layer semantics in traditional Chinese paintings.
[0090] Refer again Figure 7 In one embodiment, the multi-level structural feature extraction device 10 includes:
[0091] The micro-feature extraction module 11 is used to extract facial features at the micro level of the Chinese painting image;
[0092] The mesoscopic feature extraction module 12 is used to extract the pose features at the mesoscopic level of the Chinese painting image;
[0093] The macro-feature extraction module 13 is used to extract co-occurrence features at the macro level of traditional Chinese painting images.
[0094] Among them, the micro-feature extraction module 11 extracts facial features from the Chinese painting image through facial key point detection and facial feature calculation; the meso-feature extraction module 12 extracts posture features from the Chinese painting image through human body key point detection and posture feature calculation; and the macro-feature extraction module 13 extracts co-occurrence features from the Chinese painting image through object entity detection and object entity spatial distance calculation. The specific limitations of each functional module within the multi-level feature extraction device 10 can be found in the limitations of step S10 above, and will not be repeated here.
[0095] Refer again Figure 7 In one embodiment, the multi-level semantic extraction device 20 includes:
[0096] Natural language description generation module 21 is used to generate natural language descriptions of Chinese painting images based on the results of multi-level structure feature extraction using a visual language model.
[0097] The topic feature classification system construction module 22 is used for natural language description based on Chinese painting images. It uses a large language model to cluster and model the semantic topics in Chinese painting images to obtain a topic feature classification system.
[0098] The multi-level semantic extraction module 23 is used to automatically classify and interpret the Chinese painting image based on the theme feature classification system, and obtain the theme category and feature category of the Chinese painting image.
[0099] The specific limitations of each functional module within the multi-level semantic extraction device 20 can be found in the limitations of step S20 above, and will not be repeated here. Specifically, the topic feature classification system construction module 22 constructs a clear and reusable topic feature classification system, effectively solving the problem of the lack of a unified semantic feature modeling mechanism in the prior art.
[0100] Refer again Figure 7 In one embodiment, the multi-level semantic visualization device 30 includes:
[0101] The feature distribution view module 31 is used to generate facial feature clustering subviews, posture feature clustering subviews, and co-occurrence feature clustering subviews based on the facial features, posture features, and co-occurrence features of the figures in the Chinese painting image extracted by the multi-level structure feature extraction device.
[0102] The semantic association view generation module 32 is used to generate the semantic association view based on the multi-level structural feature extraction results and the multi-level semantic extraction results.
[0103] The time evolution analysis view generation module 33 is used to generate the time evolution analysis view based on the multi-level structural feature extraction results, the multi-level semantic extraction results, and the creation year of the Chinese painting.
[0104] The detail presentation view generation module 34 is used to generate the detail presentation view based on the multi-level semantic extraction results and the metadata of the Chinese painting.
[0105] In one embodiment, the multi-level semantic visualization device 30 further includes:
[0106] The natural language search module generates a natural language search box and displays it in the three-layer feature clustering view. It then performs feature matching based on the natural language text entered in the search box and returns the most relevant set of traditional Chinese paintings based on the matching results. Specifically, users can enter natural language text containing semantic intent into the search box. Responding to this input, the module parses the query intent using a large language model and automatically maps it to a three-layer feature dimension obtained through multi-level feature extraction for feature matching. Finally, it returns the most relevant set of traditional Chinese paintings based on semantic similarity.
[0107] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0108] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model, characterized in that, Includes the following steps: Multi-level structural feature extraction is performed on traditional Chinese painting images, extracting facial features, posture features, and co-occurrence features of figures in the images. Based on visual language models and large language models, multi-level semantic extraction is performed on traditional Chinese painting images to obtain the subject category and feature category of the images. Based on the multi-level structural feature extraction results and multi-level semantic extraction results, a three-level feature clustering view, semantic association view, temporal evolution analysis view and detail presentation view are constructed and displayed. Extracting facial features from figures in traditional Chinese paintings involves the following steps: Forty-two key points on the face are selected as target key points from the 68-key point annotation system. The 42 key points include key points 18-27, 37-48 and 49-68 in the 68-key point annotation system. Using the target key points as the detection target, the facial key point estimation model is used to detect facial key points on the Chinese painting image to obtain the facial key points of the figures in the Chinese painting image. The facial features are calculated using the coordinates of the acquired facial key points to obtain the final facial features. Extracting co-occurrence features of figures in traditional Chinese paintings involves the following steps: Entity detection is performed on traditional Chinese painting images using an object detection model to obtain the objects in the images. The acquired objects are filtered to remove non-holistic objects and retain holistic objects. The filtered object entities are paired into object entity pairs, and frequent but insufficient object entity pairs are filtered out. Using Euclidean distance as a representation of the spatial relationship between pairs of objects, the final co-occurrence feature is represented as follows: , in, and The names of the objects in a traditional Chinese painting. for and The normalized relative distance between them It is a set of object entities obtained through entity detection; The method of performing multi-level semantic extraction on traditional Chinese painting images based on visual language models and large language models to obtain the subject category and feature category of the traditional Chinese painting images includes the following steps: Based on the multi-level structural feature extraction results, a natural language description of a traditional Chinese painting image is generated using a visual language model. Based on the natural language description of traditional Chinese painting images, a large language model is used to cluster and model the semantic themes in the images to obtain a theme feature classification system. Based on the aforementioned subject feature classification system, the traditional Chinese painting images are automatically classified and interpreted to obtain their subject categories and feature categories.
2. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, The step of using the coordinates of the acquired facial key points to calculate facial features and obtain the final facial features includes the following steps: Use formula Calculate the tilt of the eyebrows, eyes, and mouth, among which... To select the orientation angle between facial key points, ; Use formula Calculate the degree of facial expression, among which, These represent the size of the eyes when they open and close, and the size of the mouth when they open and close, respectively. The final facial features are represented as follows: .
3. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, Extracting the posture features of figures in traditional Chinese paintings involves the following steps: Using 17 human keypoints as detection targets, a human keypoint estimation model was used to detect human keypoints in a traditional Chinese painting image, obtaining the skeletal keypoints of the figures in the image. Each skeletal keypoint is represented as follows: ,in , and These are the horizontal and vertical coordinates, respectively. By connecting key points of the skeleton and The Euclidean distance between them is used to calculate bone length, using the following formula: ; Where E represents the predefined connection relationship of skeletal key points; Bone orientation is calculated by determining the orientation angle between the connected bone keypoint pairs and the horizontal axis, using the following formula: ; final The pose features are represented as follows: 。 4. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, The natural language description based on traditional Chinese painting images uses a large language model to cluster and model the semantic themes in the images to obtain a theme feature classification system, including the following steps: Random batch sampling of traditional Chinese painting image data and extraction of candidate themes; The LLM algorithm is used to perform semantic topic clustering on each batch of traditional Chinese painting images, generating cluster labels and classification reasons. Summarize all candidate topics extracted from all batches and re-cluster them for optimization; Merge similar topic tags and remove ambiguous topic tags, and construct an initial topic feature classification system; The initial topic feature classification system is optimized and updated through expert manual verification and model-assisted optimization to obtain the final topic feature classification system.
5. A hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model, characterized in that, include: A multi-level structural feature extraction device is used to extract multi-level structural features from Chinese painting images, specifically extracting facial features, posture features, and co-occurrence features of figures in the Chinese painting images. A multi-level semantic extraction device is used to perform multi-level semantic extraction on Chinese painting images based on visual language models and large language models to obtain the subject category and feature category of the Chinese painting images. A multi-level semantic visualization device is used to construct and display a three-level feature clustering view, a semantic association view, a time evolution analysis view, and a detail presentation view based on the multi-level structural feature extraction results and the multi-level semantic extraction results. The multi-level structural feature extraction device is specifically used for: Forty-two key points on the face are selected as target key points from the 68-key-point annotation system. These 42 key points include key points 18-27, 37-48, and 49-68 in the 68-key-point annotation system. Using these target key points as detection targets, a facial key point estimation model is used to detect facial key points in the traditional Chinese painting image, thereby obtaining the facial key points of the figures in the painting image. The coordinates of the obtained facial key points are then used to calculate facial features, resulting in the final facial features. Entity detection is performed on traditional Chinese painting images using an object detection model to obtain object entities. These entities are then filtered, removing non-holistic entities and retaining holistic ones. The filtered entities are paired into entity pairs, and frequent but insufficiently informative pairings are removed. Euclidean distance is used to represent the spatial relationships between entity pairs, and the final co-occurrence feature is represented as follows: ,in, and The names of the objects in a traditional Chinese painting. for and The normalized relative distance between them It is a set of object entities obtained through entity detection; The multi-level semantic extraction device is specifically used for: Based on the multi-level structural feature extraction results, a natural language description of a traditional Chinese painting image is generated using a visual language model. Based on the natural language description of traditional Chinese painting images, a large language model is used to cluster and model the semantic themes in the images to obtain a theme feature classification system. Based on the aforementioned subject feature classification system, the traditional Chinese painting images are automatically classified and interpreted to obtain their subject categories and feature categories.
6. The hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model as described in claim 5, characterized in that, The multi-level semantic visualization device includes: The feature distribution view module is used to generate facial feature clustering subviews, posture feature clustering subviews, and co-occurrence feature clustering subviews based on the facial features, posture features, and co-occurrence features of figures in the Chinese painting image extracted by the multi-level structure feature extraction device. The natural language search module is used to generate a natural language search box and display the natural language search box in the three-layer feature clustering view, and to perform feature matching based on the natural language text entered in the natural language search box, and return the most relevant set of Chinese painting images according to the feature matching results; The semantic association view generation module is used to generate the semantic association view based on the multi-level structural feature extraction results and the multi-level semantic extraction results. The time evolution analysis view generation module is used to generate the time evolution analysis view based on the multi-level structural feature extraction results, multi-level semantic extraction results, and the creation year of the Chinese painting. The detail presentation view generation module is used to generate the detail presentation view based on the multi-level semantic extraction results and the metadata of the Chinese painting.