Traditional Chinese painting image hierarchical semantic extraction and visualization method and system based on large model

By combining multi-level structural feature extraction with visual language models and large language models, the problem of single-dimensional semantic information extraction from traditional Chinese painting images is solved. This enables multi-level semantic understanding and visualization of traditional Chinese painting images, and enhances the ability to understand historical context and artistic intent.

CN121366431AActive Publication Date: 2026-01-20ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511952562.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-01-20
Estimated Expiration
2045-12-23

AI Technical Summary

Technical Problem

Existing technologies for semantic understanding of traditional Chinese paintings have a relatively singular dimension in extracting semantic information, ignoring the co-occurrence relationship between facial expressions, body postures, and background objects, lacking a unified semantic feature modeling mechanism, making it difficult to deeply understand the historical context, and lacking an interactive analysis platform.

Method used

A multi-level structural feature extraction method is adopted to extract facial features, posture features and co-occurrence features of traditional Chinese painting images. Combined with visual language model and large language model, multi-level semantic extraction is performed, and a three-level feature clustering view, semantic association view and temporal evolution analysis view are constructed for visualization.

Benefits of technology

It enhances the understanding of the historical context and artistic intent of traditional Chinese painting images, provides a platform for in-depth mining and exploration of multi-level semantics, overcomes the limitations of semantic modeling in traditional methods, and realizes the display of multi-level semantic features of traditional Chinese painting images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366431A_ABST
    Figure CN121366431A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese painting image hierarchical semantic extraction and visualization method and system based on a large model. According to the method, facial features, posture features and co-occurrence features of figures in a traditional Chinese painting image are extracted from the microscopic level, the mesoscopic level and the macroscopic level respectively, and then a visual language model and a large language model are introduced based on the extracted three-layer structure features to carry out multi-level semantic extraction; according to the method, the problem of relatively single semantic information extraction dimension in the prior art is effectively solved, the technical limitation that deep semantics is difficult to model due to the fact that a traditional method depends on shallow visual features is overcome, and the ability of understanding historical contexts and artistic intentions contained in traditional paintings is improved; besides, the structural features and the semantic features of the traditional Chinese painting image are displayed through the three-layer feature clustering view, the semantic association view, the time evolution analysis view and the detail presentation view, an interactive analysis platform and approach are provided for the user, and the user is assisted in performing multi-level semantic deep mining and exploration on the traditional Chinese painting image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image data processing, and particularly relates to a Chinese painting image hierarchical semantic extraction and visualization method and system based on a large model. BACKGROUND

[0002] Chinese painting refers to traditional Chinese painting, which is usually drawn using tools and materials such as a brush pen, ink, and rice paper. Chinese painting emphasizes mood, rhythm, and vitality, and covers a wide range of subjects such as landscapes, flowers and birds, and figures. Chinese painting is an important form of Chinese cultural heritage and is widely used in art exhibitions, education and communication, and cultural digitization. In Chinese painting, figures are often used to express the author's intentions, depict historical scenes, and convey social concepts, and they often carry key semantic and situational information in the entire painting. Therefore, in the process of intelligent understanding and deep digital modeling of Chinese painting, accurately identifying and analyzing the figure elements in the painting and the multi-level semantics contained in the painting have important value.

[0003] To achieve understanding of artistic images, current technologies mainly use traditional computer vision methods such as image classification, style recognition, and object detection. These methods mostly rely on low-level visual features such as color, texture, and contour, or use convolutional neural networks to extract deep representations for style analysis or image generation tasks. In the analysis of figure elements, some methods introduce emotion recognition and pose estimation techniques, such as using OpenPose to extract key points to determine the action state of the figure, or recognizing facial expressions and emotional signals through local feature recognition. These methods have initially realized the static structural perception of figures in the picture.

[0004] In addition, in recent years, the rise of vision-language models (VLM) has provided a new method for the fusion of image and text information. For example, models such as CLIP are trained through large-scale image-text pairing, allowing images and natural language to be aligned in a unified semantic space. This type of model has been gradually applied to the classification and retrieval of museum collections, to some extent, improving the semantic structuring ability and accessibility of art.

[0005] Meanwhile, in the system oriented to image content interpretation, the DARK system provides tool support for researchers to identify common composition patterns in ancient images by detecting repeated motifs and symbolic structures in images; the InTaVia system constructs a cross-cultural knowledge graph, assists experts in establishing semantic connections between historical figures, events, and visual arts, and supports the reconstruction of cultural narratives across works; the Virtual Rosetta project attempts to visually cluster historical images and enhance immersive display effects with virtual reality; and the Arnold system is oriented to the display and management of multi-dimensional image sets, helping users organize and classify cultural image resources. Although the above systems provide certain analysis support in the time, space, and style dimensions of artistic images, most of their functions still rely on external metadata of images (such as creation year, work location, artist information, etc.), and lack systematic modeling capabilities for the internal semantic structure of images (such as character expression, action state, and relationship between characters and background, etc.). At the same time, these systems generally lack mechanisms for structured extraction and semantic reconstruction of key visual elements in images, especially character elements, making it difficult to meet the needs of users in cultural understanding, style analysis, and narrative reasoning.

[0006] In summary, in the field of Chinese painting image semantic understanding and visual analysis, the following technical problems need to be solved: first, the semantic information extraction dimension is relatively single, often ignoring the multi-level structural characteristics of characters in facial expressions, body postures, and co-occurrence relationships with background objects; second, there is a lack of unified semantic feature modeling mechanism, resulting in scattered and unsystematic semantic expression, which leads to a lack of hierarchy and completeness in semantic expression; third, the semantics contained in Chinese painting works are highly dependent on specific historical and cultural backgrounds, and traditional methods are difficult to effectively construct semantics through shallow visual features, resulting in a lack of understanding of Chinese painting in historical context; fourth, there is a lack of interactive analysis platform and mechanism to provide users with deep mining and exploration paths for structural features and semantic features, limiting users' in-depth understanding and application of potential semantic information in images. SUMMARY

[0007] The technical problem to be solved by the present application is to provide a Chinese painting image hierarchical semantic extraction and visualization method and system based on a large model, to solve the problems of single semantic information extraction dimension, lack of understanding of Chinese painting in historical context, and lack of interactive analysis approach in the prior art.

[0008] To solve the above technical problems, the technical solution adopted by the present application is as follows: In one aspect of the present application, a method for extracting and visualizing hierarchical semantics of a Chinese painting image based on a large model is provided, comprising the following steps: performing multi-level structure feature extraction on the Chinese painting image, and extracting facial features, posture features, and co-occurrence features of characters in the Chinese painting image respectively; performing multi-level semantic extraction on the Chinese painting image based on a visual language model and a large language model, and obtaining theme categories and feature categories of the Chinese painting image; and constructing and displaying three-layer feature clustering views, semantic association views, time evolution analysis views, and detail presentation views based on the multi-level structure feature extraction results and the multi-level semantic extraction results.

[0009] In another aspect of the present application, a system for extracting and visualizing hierarchical semantics of a Chinese painting image based on a large model is provided, comprising: a multi-level structure feature extraction device configured to perform multi-level structure feature extraction on the Chinese painting image, and extract facial features, posture features, and co-occurrence features of characters in the Chinese painting image respectively; a multi-level semantic extraction device configured to perform multi-level semantic extraction on the Chinese painting image based on a visual language model and a large language model, and obtain theme categories and feature categories of the Chinese painting image; and a multi-level semantic visualization device configured to construct and display three-layer feature clustering views, semantic association views, time evolution analysis views, and detail presentation views based on the multi-level structure feature extraction results and the multi-level semantic extraction results.

[0010] The present application has the beneficial technical effects that: by extracting facial features, posture features, and co-occurrence features of characters in the Chinese painting image from micro, meso, and macro levels respectively, three-layer structure features of the Chinese painting image are obtained, then, based on the extracted three-layer structure features of the Chinese painting image, a visual language model and a large language model are introduced to perform multi-level semantic extraction, and semantic features of the Chinese painting image are obtained, effectively solving the problem of single dimension of semantic information extraction in the prior art, overcoming the technical limitation that traditional methods rely on shallow visual features and are difficult to model deep semantics, thereby improving the understanding ability of historical context and artistic intention contained in the Chinese painting; in addition, the extracted structure features and the extracted semantic features of the Chinese painting image are visualized, and the structure features and the semantic features of the Chinese painting image are displayed through three-layer feature clustering views, semantic association views, time evolution analysis views, and detail presentation views, providing an interactive analysis platform and approach for users, assisting users in performing multi-level semantic deep mining and exploration on the Chinese painting image. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 FIG. 1 is a flowchart of a method for extracting and visualizing hierarchical semantics of a Chinese painting image based on a large model according to an embodiment of the present application; Figure 2 FIG. 2 is a flowchart of multi-level structure feature extraction according to an embodiment of the present application; Figure 3A schematic diagram of a face 68 key point labeling system; Figure 4 A flowchart of multi-level semantic extraction in an embodiment of the present application; Figure 5 An effect schematic diagram of the time evolution analysis view in the aggregation mode of the present application; Figure 6 An effect schematic diagram of the time evolution analysis view in the detail mode of the present application; Figure 7 An effect schematic diagram of the detail presentation view of the present application; Figure 8 A structural schematic diagram of a Chinese painting image hierarchical semantic extraction and visualization system based on a large model in an embodiment of the present application. DETAILED DESCRIPTION

[0012] In order for those skilled in the art to more clearly understand the purpose, technical solution and advantages of the present application, the present application will be further described below in conjunction with the drawings and embodiments.

[0013] The present application provides a Chinese painting image hierarchical semantic extraction and visualization method based on a large model. As shown in the figure, Figure 1 In an embodiment of the present application, the Chinese painting image hierarchical semantic extraction and visualization method based on a large model includes steps S10 to S30: S10: Multi-level structure feature extraction is performed on the Chinese painting image, and facial features, posture features, co-occurrence features of the characters in the Chinese painting image are extracted respectively.

[0014] Specifically, step S10 extracts facial features, posture features, and co-occurrence features of the characters in the Chinese painting image from three levels of micro level, meso level, and macro level, respectively, to obtain three-level structure features of the Chinese painting image: the facial features of the micro level are based on a face key point detection model to identify the geometric features of the face region, and then calculate and extract key parameters related to character emotions and facial expressions such as eyebrow tilt angle, eye opening degree, and mouth shape; the posture features of the meso level are obtained by a human body key point estimation model to obtain human body skeleton structure key point coordinates, and then calculate the length and direction of the skeleton connection to depict the body movements and behavior characteristics of the characters; the macro level detects the object entities in the Chinese painting image through a target detection model, and then constructs object entity pairs, and finally takes the object entity names and the normalized relative distances between the object entities as the co-occurrence features of the macro level.

[0015] As shown in the figure, Figure 2 The multi-level structure feature extraction in step S10 includes three links of facial feature extraction, posture feature extraction, and co-occurrence feature extraction.

[0016] 1.1 Facial feature extraction The facial feature extraction process at the microscopic level mainly consists of facial landmark detection and facial feature calculation: 1.1.1 First, 42 key points on the face are selected as target key points from the 68-key-point annotation system for the human face. For example... Figure 3 As shown, the 68-key-point annotation system for the face accurately describes the geometric structure of the face by annotating 68 key points in specific facial regions (such as eyes, eyebrows, nose, mouth, and contours). The 42 key points include key points 18-27 for describing eyebrows, key points 37-48 for describing eyes, and key points 49-68 for describing the mouth. This embodiment of the invention selects 42 key points from the 68-key-point annotation system as target key points, which can reduce the computational load in subsequent facial key point detection and improve detection efficiency. Of course, in other embodiments of the invention, the 68 key points of the face can also be directly used as target key points for subsequent facial key point detection.

[0017] 1.1.2 Then, a facial keypoint estimation model is used to detect facial keypoints in the traditional Chinese painting image to obtain the facial keypoints of the figures in the painting image. Specifically, using the target keypoints as the detection target, a facial keypoint estimation model is used to detect facial keypoints in the traditional Chinese painting image to obtain the facial keypoints of the figures in the painting image. Facial keypoint estimation models include Openpose, FAN, Sapiens-1b-pose-133-keypoints, etc. In this embodiment of the invention, Sapiens-1b-pose-133-keypoints is used for facial keypoint detection.

[0018] 1.1.3 Finally, the coordinates of the acquired facial key points are used to calculate facial features, resulting in the final facial features. Specifically, after acquiring the facial key points, the coordinates of the key points are used to calculate the tilt of the eyebrows, eyes, and mouth. The tilt (angular features) of the eyebrows, eyes, and mouth are uniformly represented as follows: ,in, The orientation angles between selected facial key points are calculated using the coordinates between them. The degree of expression, such as the opening and closing of the eyes and mouth, is calculated using the distances between these key points. The degree of expression is expressed as... ,in, These represent the size of the eyes when they open and close, and the size of the mouth when they open and close. The value is taken as the vertical distance between selected facial key points; the final facial feature is represented as: .

[0019] 1.2 Pose Feature Extraction The pose feature extraction link at the mesoscopic level mainly includes human key point detection and pose feature calculation: 1.2.1 First, the human key point estimation model is used for human key point detection of the Chinese painting image to obtain the skeletal key points of the figure in the Chinese painting image. Specifically, 17 human key points are taken as the detection target, the human key point estimation model is used for human key point detection of the Chinese painting image to obtain the skeletal key points of the figure in the Chinese painting image, and each skeletal key point is represented as , wherein , and are horizontal and vertical coordinates respectively. Human key point detection refers to detecting important joint points (such as head, shoulder, elbow, knee, etc.) of the human body in the image, and inferring the posture or action state of the figure according to the joint points. The human key point estimation model includes OpenPose, PoseNet, AlphaPose, Sapiens-1b-pose-133-keypoints, etc. In the embodiment of the present application, Sapiens-1b-pose-133-keypoints is used for human key point detection.

[0020] 1.2.2 Then, the 17 human key points (skeletal key points) obtained are used for pose feature calculation to obtain the skeletal length and skeletal direction.

[0021] Specifically, the skeletal length is calculated by the Euclidean distance between the skeletal key points and , and the formula is: ; wherein E represents a predefined skeletal key point connection relationship.

[0022] The skeletal direction is calculated by calculating the direction angle between the connected skeletal key point pair and the horizontal axis, and the formula is: ; The final pose feature is represented as: .

[0023] 1.3 Co-occurrence feature extraction 1.3.1 First, the target detection model is used for entity detection of the Chinese painting image to obtain the object entity in the Chinese painting image. Specifically, target detection is a computer vision task for locating and classifying different objects in an image. The target detection model includes YOLO, DINO-X, Grounding DINO, etc. In the embodiment of the present application, DINO-X is used for entity detection of the Chinese painting image.

[0024] 1.3.2 Then, the obtained object entities are screened to remove non-integral object entities and retain integral object entities. Specifically, for the obtained object entities, the individual parts of human organs such as hands and eyes are removed, and only integral entities such as complete human bodies, animals, plants, buildings, etc. are retained to ensure the effectiveness of the extracted object information and ensure that the key elements in the painting can be effectively extracted.

[0025] 1.3.3 Then, the screened object entities are combined into object entity pairs two by two, and object entity pairs with insufficient information quantity are filtered out. Specifically, after the screened object entities are combined into object entity pairs two by two, the term frequency-inverse document frequency method is used to filter out combinations (object entity pairs) with insufficient information quantity, and only the top 70% of object entity pairs are retained.

[0026] 1.3.4 Finally, the Euclidean distance is used as the spatial relationship between the object entity pairs, and the final co-occurrence feature is represented as: ; wherein, and are the object entity names in the Chinese painting image, are the normalized relative distances between and .

[0027] S20: Multi-level semantic extraction of the Chinese painting image based on a visual language model and a large language model to obtain the theme category and feature category of the Chinese painting image.

[0028] Specifically, the step S20 introduces a visual language model and a large language model for semantic modeling based on the structural features extracted in the step S10, and extracts the semantic features of the Chinese painting image according to the modeling of the theme feature classification system. Semantic modeling includes two steps of semantic generation based on a visual language model and semantic clustering modeling based on a large language model: in the semantic generation step based on the visual language model, the original Chinese painting image, the structural features extracted in the step S10 and the externally obtained Chinese painting image metadata are input into the visual language model to automatically generate natural language descriptions of human faces, bodies and spatial layouts; in the semantic clustering modeling step based on the large language model, the semantic theme of the Chinese painting image is modeled by using the large language model to obtain the theme feature classification system. After the theme feature classification system is constructed, the semantic features of the Chinese painting image are extracted according to the modeling of the theme feature classification system to describe the semantic of the Chinese painting image from micro, meso, macro and theme levels.

[0029] As Figure 4 ​As shown, the step S20, i.e., the multi-level semantic extraction of the Chinese painting image based on the visual language model and the large language model, to obtain the theme category and feature category of the Chinese painting image, includes the following steps S21 to S23: S21: Based on the multi-level structure feature extraction result, a visual language model is used to generate a natural language description of the Chinese painting image. Specifically, the Chinese painting image, metadata (including painting name, author, creation year, category label, etc.), and the structure features extracted in step S10 are used as the input of the visual language model, and then the visual language model is used to generate a natural language description of the face, body, and spatial features as an artist, and generate a summary of the overall visual content of the Chinese painting. The visual language model refers to a pre-trained multi-modal model that can understand both image and text information. This type of model is trained on large-scale image-text aligned data and has the ability to establish a correspondence between image content and natural language. Its main functions include image content recognition, text generation, image-text matching, and cross-modal reasoning. The visual language model includes CLIP, DALL·E series, BLIP, Flamingo, LLaVA, DeepSeek-VL, GPT series, etc. In this embodiment, the visual language model uses GPT-4o-mini.

[0030] S22: Based on the natural language description of the Chinese painting image, a large language model is used to cluster and model the semantic themes in the Chinese painting image to obtain a theme feature classification system. Specifically, the Chinese painting image and its natural language description are used as the input of the large language model, and the semantic themes in the Chinese painting image data are extracted by the large language model, and the semantic theme clustering is performed to generate corresponding theme labels. Finally, a semantic theme-feature category mapping relationship is established to construct the theme feature classification system.

[0031] The large language model is a large-scale natural language processing model trained on a large amount of text corpus, with strong language understanding and generation capabilities. The large language model can integrate context and reasoning information, and convert structured input into semantic description. The large language model includes GPT-4o-mini, Claude 2 / 3, PaLM 2, Llama 2 / 3, Kosmos-2, etc. In this embodiment, the large language model uses GPT-4o-mini.

[0032] In this embodiment, to solve the context length limitation problem, step S22 uses a bottom-up batch processing strategy, which first performs local clustering and then integrates and integrates to finally construct a clear and reusable theme feature classification system, which includes the following steps: 2.2.1 Random batch sampling of Chinese painting image data and extracting candidate themes. Specifically, n paintings are randomly selected, and the natural language description generated in step S21 is combined as input to the large language model. The large language model extracts candidate themes and feature categories from the batch of Chinese painting image data.

[0033] 2.2.2 Invoking LLM for semantic theme clustering of each batch of Chinese painting images, generating clustering labels and classification reasons. Specifically, the text description of each batch of Chinese painting images is input into the large language model. The large language model analyzes the similarities and differences of similar Chinese painting images and generates several candidate theme labels (i.e., clustering labels), and provides the clustering logic and reason explanation of each theme.

[0034] 2.2.3 Summarizing all batch-extracted candidate themes and re-clustering optimization. Specifically, the candidate themes extracted from all batch Chinese painting image data are summarized and input into the large language model for secondary clustering, generating corresponding theme labels and classification reasons. Merging and analyzing the theme labels generated by secondary clustering and the theme labels generated by semantic theme clustering of each batch of Chinese painting images can improve the consistency, accuracy, and coverage of the clustering results.

[0035] 2.2.4 Merging similar theme labels and removing ambiguous theme labels, and constructing an initial theme feature classification system. Specifically, the theme labels generated by secondary clustering and the theme labels generated by semantic theme clustering of each batch of Chinese painting images are integrated together. Similar or overlapping theme labels are merged, and ambiguous or meaningless theme labels are removed to make the subsequent theme feature classification system more clear and explicit. Finally, the theme labels obtained after merging and removing are mapped to the corresponding feature categories to construct an initial theme feature classification system.

[0036] 2.2.5 Optimizing and updating the initial theme feature classification system through expert manual verification and model-assisted optimization to obtain the final theme feature classification system. Through expert manual verification and model-assisted optimization, a theme feature classification system with wide coverage and clear semantic boundaries is constructed, serving as a standardized image label reference.

[0037] Through step S22, a clear and reusable theme feature classification system is constructed, effectively solving the problem of lack of semantic feature unified modeling mechanism in the prior art.

[0038] ​S23: Automatically classify and interpret the Chinese painting image based on the theme-feature classification system, to obtain the theme category and feature category of the Chinese painting image. Specifically, the theme-feature classification system, the structural features of each Chinese painting image, and the text description are input into a large language model, the large language model automatically reasons and assigns the most matching theme category and feature category to each Chinese painting image, and generates a clear classification explanation, finally outputs the theme category and feature category of the Chinese painting image for subsequent analysis and display. Wherein the structural features of the Chinese painting image are extracted in step S10, including facial features, posture features and co-occurrence features; the text description of the Chinese painting image is generated by the visual language model in step S21.

[0039] S30: Based on the multi-level structure feature extraction result and the multi-level semantic extraction result, construct and display three-layer feature clustering view, semantic association view, time evolution analysis view and detail presentation view.

[0040] Specifically, step S30 visualizes the structural features extracted in step S10 and the semantic features of the Chinese painting image extracted in step S20, and displays the structural features and semantic features of the Chinese painting image through three-layer feature clustering view, semantic association view, time evolution analysis view and detail presentation view, to assist users in deep mining and exploring the multi-level semantics of the Chinese painting image. Wherein, the three-layer feature clustering view is based on the three types of structural features of facial features, posture features and co-occurrence features, and visualizes and clusters the three types of structural features of the Chinese painting image, and provides natural language search function; the semantic association view reduces the dimensionality of the association relationship between the multi-level semantic features; the time evolution analysis view constructs the association between the painting semantics and the historical background, to assist users in understanding the evolution trend of the structural features and semantic themes in the time dimension; the detail presentation view supports in-depth viewing of the visual and semantic details of a single painting.

[0041] 3.1 Three-layer feature clustering view: The three-layer feature clustering view is used to present the distribution of the three-layer structural features (facial features, posture features, and co-occurrence features) of the Chinese painting image extracted in step S10 in the embedding space, and supports natural language search and interactive exploration operations.

[0042] The three-layer feature clustering view includes three sub-views of facial feature clustering, posture feature clustering, and co-occurring feature clustering, and the three sub-views are all clustering scatter plots. Among them, the t-SNE dimension reduction method is used to project high-dimensional features (facial features or posture features) to a two-dimensional plane to generate a facial feature clustering sub-view or a posture feature clustering sub-view. In the facial feature clustering and posture feature clustering sub-views, the spatial proximity represents the structural feature similarity between images; in the co-occurring feature clustering sub-view, the image pair similarity is first calculated by combining the Jaccard similarity and the normalized Euclidean distance, and then the t-SNE dimension reduction method is used to reduce the co-occurring features to a two-dimensional plane to draw a clustering scatter plot.

[0043] In the three-layer feature clustering view, the user can hover or zoom to view the representative works of a specific cluster, and support the linkage highlighting operation, that is, when a scatter point is selected in any one of the three sub-views, the corresponding painting of the selected scatter point will also be highlighted in the corresponding scatter points in the remaining two sub-views.

[0044] The three-layer feature clustering view supports natural language search. Specifically, a natural language search box is provided in the three-layer feature clustering view, and the user can input natural language text containing semantic intent in the natural language search box. In response to the user's input operation in the natural language search box, the query intent is analyzed by a large language model, and is automatically mapped to the three-layer structure feature dimensions obtained through multi-level structure feature extraction for feature matching, and the most relevant set of Chinese painting image works is returned according to the semantic similarity.

[0045] 3.2 Semantic association view: The composition and function of the semantic association view are as follows: 3.2.1 Chord diagram: represents the proportion and association of micro, meso and macro features in the cluster. Specifically, the chord diagram includes an outer ring and an inner scatter point, wherein each arc segment of the ring represents a specific feature category in the structure level (for example, in the facial feature level, it may be a "dignified" expression, and in the posture level, it may be a "kneeling posture", etc.), and the longer the arc length, the greater the proportion of the feature in the current level, that is, more works have this feature; each scatter point in the diagram represents a Chinese painting, and the scatter points are scattered in a two-dimensional space, reflecting the distribution position of the work in the structure and semantic dimensions; the color of the scatter point represents the theme classification of the work, such as "woman painting", "Daoist and Buddhist figure painting", "portrait painting", etc.

[0046] 3.2.2 Theme statistics panel: embeds the semantic label and classification reason of each image using a large language model, calculates the semantic similarity in the theme level, and maps it to a two-dimensional plane in a spatial distance manner. The color of the bar represents the theme category, and the length of the bar represents the number of Chinese paintings included in the current theme category.

[0047] 3.2.3 Link between chord diagrams: For pairs of clusters with similar structural characteristics, links with adjustable thresholds visualize the degree of their commonality. To avoid visual clutter, only the top 25% of the most similar links are displayed by default. Specifically, a link extending from the outer ring of one chord diagram to another represents the association between the current feature and another two structural levels, e.g., a specific posture feature (e.g., "turning to look back") might frequently co-occur with a particular facial expression (e.g., "smiling") or a co-occurring object (e.g., "plum blossoms"). The thickness (or number) of the link represents the strength of the association, helping users identify structural commonalities or differences across levels.

[0048] The semantic association view is used to identify the commonalities and differences in the structural patterns of characters under different themes, facilitating user analysis of the association between themes and structural features. For example: it is found that the feature of a solemn and dignified face occupies the largest proportion in Taoist figure paintings and historical figure paintings. To explore the reason, in Taoist figure paintings, the solemn expression symbolizes the spiritual state of the characters, such as their cultivation, tranquility, and transcendence; while in historical figure paintings, it is used to show the loyalty and righteousness of heroic characters or their sense of responsibility. Although the themes are different, both of them use solemn expressions to convey authority, spirituality, and moral power, thus having similarities in visual expression. For another example, when the user hovers the mouse over an arc segment of the chord diagram, it is found that the depiction of dance scenes in pleasure-seeking paintings in different periods of time has significant differences in character posture, expression, and surrounding objects (such as musical instruments, gardens, attendants, etc.): Tang Dynasty works mostly represent bold and open movements and a joyful atmosphere, while Song Dynasty works tend to be subtle and delicate, emphasizing rhythm and elegance, reflecting the aesthetic orientation and social atmosphere of different eras.

[0049] 3.3 Time evolution analysis view: The time evolution analysis view combines historical timeline information to show the time evolution trends of structural features and semantic themes, including the following two modes: Aggregated mode: Paintings with a time interval less than a pre-set threshold are displayed in an aggregated manner, with the area and saturation of the circle representing the number of aggregations, suitable for macro overview, as shown in Figure 5 .

[0050] Detailed mode: Each painting is displayed in a vertical line manner at the corresponding year, and the work name and original image can be viewed by hovering, suitable for detailed analysis and local enlargement. As shown in Figure 6 , each vertical line in the figure represents a painting, and the time line feature column on the left side of the figure lists the top 10 objects in the selected Chinese painting image, such as ladies, gardens, garden stones, maids, and trees. The circle on the vertical line indicates that the object on the left side of the circle appears in the painting, for example, there are four circles on the third vertical line, indicating that the painting represented by the vertical line simultaneously contains four objects: ladies, gardens, maids, and trees.

[0051] 3.4 Detail presentation view: As shown in Figure 7 The left side of the detail presentation view shows the original image of the Chinese painting, and the right side shows the natural language description of the Chinese painting image. At the same time, by clicking the image on the left side, the user can jump to the metadata display page to show the metadata of the Chinese painting, including the creation year of the work, the location of the work, the artist information, etc.

[0052] For steps S10-S30, after receiving the Chinese painting image, first, the facial features, posture features, and co-occurrence features of the characters in the Chinese painting image are extracted from the micro level, the meso level, and the macro level respectively to obtain the three-layer structural features of the Chinese painting image. Then, based on the extracted three-layer structural features of the Chinese painting image, a visual language model and a large language model are introduced for multi-level semantic extraction to obtain the semantic features of the Chinese painting image. Finally, the extracted structural features and the extracted semantic features of the Chinese painting image are visualized, and the structural features and the semantic features of the Chinese painting image are displayed through the three-layer feature clustering view, the semantic association view, the time evolution analysis view, and the detail presentation view to assist the user in deep mining and exploration of the multi-level semantics of the Chinese painting image.

[0053] The method for extracting and visualizing hierarchical semantics of Chinese painting images based on a large model in the embodiment of the application extracts the facial features, posture features, and co-occurrence features of the characters in the Chinese painting image from the micro level, the meso level, and the macro level respectively to obtain the three-layer structural features of the Chinese painting image, and then extracts the semantic features of the Chinese painting image based on the extracted three-layer structural features of the Chinese painting image by introducing a visual language model and a large language model for multi-level semantic extraction. This effectively solves the problem of single dimension of semantic information extraction in the prior art, overcomes the technical limitation of traditional methods that rely on shallow visual features and are difficult to model deep semantics, and thus improves the understanding ability of the historical context and artistic intention contained in the Chinese painting. In addition, a clear and reusable theme feature classification system is constructed, effectively solving the problem of lack of unified modeling mechanism for semantic features in the prior art. Finally, the extracted structural features and the extracted semantic features of the Chinese painting image are visualized, and the structural features and the semantic features of the Chinese painting image are displayed through the three-layer feature clustering view, the semantic association view, the time evolution analysis view, and the detail presentation view to provide an interactive analysis platform and approach for the user, assisting the user in deep mining and exploration of the multi-level semantics of the Chinese painting image.

[0054] The application also provides a system for extracting and visualizing hierarchical semantics of Chinese painting images based on a large model. As shown in Figure 8As shown, in one embodiment of the present application, the large model-based Chinese painting image hierarchical semantic extraction and visualization system includes a multi-level structure feature extraction device 10, a multi-level semantic extraction device 20, and a multi-level semantic visualization device 30. Each device is described in detail as follows: The multi-level structure feature extraction device 10 is used for multi-level structure feature extraction of Chinese painting images, and respectively extracts facial features, posture features, and co-occurrence features of characters in the Chinese painting images.

[0055] Specifically, the multi-level structure feature extraction device 10 extracts facial features, posture features, and co-occurrence features of characters in the Chinese painting images from three levels of microscopic, mesoscopic, and macroscopic levels, respectively, to obtain three-level structure features of the Chinese painting images: the facial features at the microscopic level are based on a face key point detection model to identify the geometric features of the face region, and then calculate and extract key parameters related to character emotions and facial expressions such as eyebrow tilt angle, eye opening degree, and mouth shape; the posture features at the mesoscopic level obtain human skeleton structure key point coordinates through a human key point estimation model, and then calculate the length and direction of the skeleton connection to depict the body movements and behavior features of the characters; the macroscopic level detects object entities in the Chinese painting image through a target detection model, and then constructs object entity pairs, and finally takes the object entity names and the normalized relative distances between the object entities as the co-occurrence features at the macroscopic level.

[0056] The multi-level semantic extraction device 20 is used for multi-level semantic extraction of Chinese painting images based on a visual language model and a large language model, to obtain the theme category and feature category of the Chinese painting image.

[0057] Specifically, the multi-level semantic extraction device 20 introduces a visual language model and a large language model for semantic modeling based on the structure features extracted by the multi-level structure feature extraction device 10, and extracts semantic features of the Chinese painting image according to the theme feature classification system obtained by modeling. Semantic modeling includes two steps of semantic generation based on a visual language model and semantic clustering modeling based on a large language model: in the semantic generation based on a visual language model step, the original Chinese painting image, the structure features extracted by the multi-level structure feature extraction device 10, and the externally obtained Chinese painting image metadata are input into the visual language model together to automatically generate natural language descriptions of the characters' faces, bodies, and spatial layout; in the semantic clustering modeling based on a large language model step, a large language model is used to cluster model the semantic theme and structure features of the Chinese painting image to obtain a theme feature classification system. After the theme feature classification system is constructed, the semantic features of the Chinese painting image are extracted according to the theme feature classification system obtained by modeling, and the Chinese painting image semantics are described from the microscopic, mesoscopic, macroscopic, and theme levels.

[0058] The multi-level semantic visualization device 30 is used to construct and display a three-level feature clustering view, a semantic association view, a time evolution analysis view and a detail presentation view based on the multi-level structure feature extraction result and the multi-level semantic extraction result.

[0059] Specifically, the multi-level semantic visualization device 30 is linked with the multi-level structure feature extraction device 10 and the multi-level semantic extraction device 20, and performs visualization processing on the structure features extracted by the multi-level structure feature extraction device 10 and the semantic features of the Chinese painting images extracted by the multi-level semantic extraction device 20, and displays the structure features and the semantic features of the Chinese painting images through the three-level feature clustering view, the semantic association view, the time evolution analysis view and the detail presentation view, thereby assisting users in deep mining and exploring the multi-level semantics of the Chinese painting images. The three-level feature clustering view is used to visually cluster and display the Chinese painting images based on the three types of structure features, i.e., the face features, the posture features and the co-occurrence features, and provides a natural language search function; the semantic association view is used to dimensionally display the association relationship between the multi-level semantic features; the time evolution analysis view is used to assist users in understanding the evolution trend of the structure features and the semantic themes in the time dimension by constructing the association between the painting semantics and the historical background; and the detail presentation view is used to support in-depth viewing of the visual and semantic details of a single painting.

[0060] The multi-level semantic extraction and visualization system based on a large model provided in the embodiment of the present application receives a Chinese painting image, extracts face features, posture features and co-occurrence features of a person in the Chinese painting image from a microscopic level, a mesoscopic level and a macroscopic level, obtains three-level structure features of the Chinese painting image, introduces a visual language model and a large language model based on the three-level structure features of the Chinese painting image to perform multi-level semantic extraction, and obtains semantic features of the Chinese painting image, thereby effectively solving the problem of single dimension of semantic information extraction in the prior art, overcoming the technical limitation that a traditional method is difficult to model deep semantics due to the dependence on shallow visual features, and thereby improving the understanding ability of historical context and artistic intention contained in the Chinese painting. In addition, the structure features and the semantic features of the Chinese painting image are subjected to visualization processing, and the structure features and the semantic features of the Chinese painting image are displayed through the three-level feature clustering view, the semantic association view, the time evolution analysis view and the detail presentation view, thereby providing an interactive analysis platform and approach for users, and assisting users in deep mining and exploring the multi-level semantics of the Chinese painting image.

[0061] Referring again to Figure 8 In an embodiment, the multi-level structure feature extraction device 10 comprises: The microscopic feature extraction module 11 is used to extract face features of a microscopic level of the Chinese painting image. The mesoscopic feature extraction module 12 is used to extract posture features of a mesoscopic level of the Chinese painting image. The macro feature extraction module 13 is configured to extract the co-occurrence features of the Chinese painting image at a macro level.

[0062] The micro feature extraction module 11 is configured to extract the facial features of the Chinese painting image by facial key point detection and facial feature calculation; the meso feature extraction module 12 is configured to extract the posture features of the Chinese painting image by human key point detection and posture feature calculation; and the macro feature extraction module 13 is configured to extract the co-occurrence features of the Chinese painting image by object entity detection and object entity-to-space distance calculation. The specific definitions of the functional modules inside the multi-level structure feature extraction device 10 can be found in the above definitions of step S10, and will not be repeated here.

[0063] Referring again to Figure 8 In an embodiment, the multi-level semantic extraction device 20 comprises: The natural language description generation module 21 is configured to generate a natural language description of the Chinese painting image using a visual language model based on the multi-level structure feature extraction result. The theme feature classification system construction module 22 is configured to use a large language model to cluster and model the semantic themes in the Chinese painting image based on the natural language description of the Chinese painting image, to obtain a theme feature classification system. The multi-level semantic extraction module 23 is configured to automatically classify and interpret the Chinese painting image based on the theme feature classification system, to obtain the theme category and feature category of the Chinese painting image.

[0064] The specific definitions of the functional modules inside the multi-level semantic extraction device 20 can be found in the above definitions of step S20, and will not be repeated here. Through the theme feature classification system construction module 22, a clear and reusable theme feature classification system is constructed, effectively solving the problem of lack of semantic feature unified modeling mechanism in the prior art.

[0065] Referring again to Figure 8 In an embodiment, the multi-level semantic visualization device 30 comprises: The feature distribution view module 31 is configured to generate a facial feature clustering sub-view, a posture feature clustering sub-view, and a co-occurrence feature clustering sub-view based on the facial features, posture features, and co-occurrence features of the characters in the Chinese painting image extracted by the multi-level structure feature extraction device. The semantic association view generation module 32 is configured to generate the semantic association view based on the multi-level structure feature extraction result and the multi-level semantic extraction result. The time evolution analysis view generation module 33 is configured to generate the time evolution analysis view based on the multi-level structure feature extraction result, the multi-level semantic extraction result, and the creation year of the Chinese painting work. The detail presentation view generation module 34 is configured to generate the detail presentation view based on the multi-level semantic extraction result and the metadata of the Chinese painting work.

[0066] In an embodiment, the multi-level semantic visualization device 30 further comprises: The natural language search module is configured to generate a natural language search box and display the natural language search box in the three-level feature clustering view, perform feature matching based on the natural language text input in the natural language search box, and return a most relevant Chinese painting image set according to the feature matching result. Specifically, the user can input natural language text containing semantic intent in the natural language search box, and the natural language search module responds to the input operation of the user in the natural language search box, analyzes the query intent through a large language model, automatically maps to the three-level structure feature dimensions obtained through multi-level structure feature extraction to perform feature matching, and returns a most relevant Chinese painting image work set according to the semantic similarity.

[0067] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is exemplified, and in actual application, the above functions can be completed by different functional units or modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above described functions.

[0068] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model, characterized in that, Includes the following steps: Multi-level structural feature extraction is performed on traditional Chinese painting images, extracting facial features, posture features, and co-occurrence features of figures in the images. Based on visual language models and large language models, multi-level semantic extraction is performed on traditional Chinese painting images to obtain the subject category and feature category of the images. Based on the results of multi-level structural feature extraction and multi-level semantic extraction, a three-level feature clustering view, semantic association view, temporal evolution analysis view, and detail presentation view are constructed and displayed.

2. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, Extracting facial features from figures in traditional Chinese paintings involves the following steps: The facial key point estimation model is used to detect facial key points in Chinese painting images to obtain the facial key points of the figures in the Chinese painting images. The facial features are calculated using the coordinates of the obtained facial key points to obtain the final facial features.

3. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 2, characterized in that, The method of using a facial key point estimation model to detect facial key points in a traditional Chinese painting image, and obtaining the facial key points of the figures in the painting image, includes the following steps: Forty-two key points of the face were selected as target key points from the 68 key point annotation system for the face. The 42 key points include key points 18-27, key points 37-48 and key points 49-68 in the 68 key point annotation system for the face. Using the target key points as the detection target, a facial key point estimation model is used to detect facial key points in the traditional Chinese painting image, thereby obtaining the facial key points of the figures in the traditional Chinese painting image.

4. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 2, characterized in that, The step of using the coordinates of the acquired facial key points to calculate facial features and obtain the final facial features includes the following steps: Use formula Calculate the tilt of the eyebrows, eyes, and mouth, among which... To select the directional angle between facial key points; Use formula Calculate the degree of facial expression, among which, These represent the size of the eyes when they open and close, and the size of the mouth when they open and close, respectively. The final facial features are represented as follows: .

5. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, Extracting the posture features of figures in traditional Chinese paintings involves the following steps: Using 17 human keypoints as detection targets, a human keypoint estimation model was used to detect human keypoints in a traditional Chinese painting image, obtaining the skeletal keypoints of the figures in the image. Each skeletal keypoint is represented as follows: ,in , and These are the horizontal and vertical coordinates, respectively. By connecting key points of the skeleton and The Euclidean distance between them is used to calculate bone length, using the following formula: , Where E represents the predefined connection relationship of skeletal key points; Bone orientation is calculated by determining the orientation angle between the connected bone keypoint pairs and the horizontal axis, using the following formula: ; final The pose features are represented as follows: 。 6. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, Extracting co-occurrence features of figures in traditional Chinese paintings involves the following steps: Entity detection is performed on traditional Chinese painting images using an object detection model to obtain the objects in the images. The acquired objects are filtered to remove non-holistic objects and retain holistic objects. The filtered object entities are paired into object entity pairs, and frequent but insufficient object entity pairs are filtered out. Using Euclidean distance as a representation of the spatial relationship between pairs of objects, the final co-occurrence feature is represented as follows: , in, and The names of the objects in a traditional Chinese painting. for and The normalized relative distance between them This is the set of object entities obtained through entity detection.

7. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 1, characterized in that, The method of performing multi-level semantic extraction on traditional Chinese painting images based on visual language models and large language models to obtain the subject category and feature category of the traditional Chinese painting images includes the following steps: Based on the multi-level structural feature extraction results, a natural language description of a traditional Chinese painting image is generated using a visual language model. Based on the natural language description of traditional Chinese painting images, a large language model is used to cluster and model the semantic themes in the images to obtain a theme feature classification system. Based on the aforementioned subject feature classification system, the traditional Chinese painting images are automatically classified and interpreted to obtain their subject categories and feature categories.

8. The method for hierarchical semantic extraction and visualization of traditional Chinese painting images based on a large model as described in claim 7, characterized in that, The natural language description based on traditional Chinese painting images uses a large language model to cluster and model the semantic themes in the images to obtain a theme feature classification system, including the following steps: Random batch sampling of traditional Chinese painting image data and extraction of candidate themes; The LLM algorithm is used to perform semantic topic clustering on each batch of traditional Chinese painting images, generating cluster labels and classification reasons. Summarize all candidate topics extracted from all batches and re-cluster them for optimization; Merge similar topic tags and remove ambiguous topic tags, and construct an initial topic feature classification system; The initial topic feature classification system is optimized and updated through expert manual verification and model-assisted optimization to obtain the final topic feature classification system.

9. A hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model, characterized in that, include: A multi-level structural feature extraction device is used to extract multi-level structural features from Chinese painting images, specifically extracting facial features, posture features, and co-occurrence features of figures in the Chinese painting images. A multi-level semantic extraction device is used to perform multi-level semantic extraction on Chinese painting images based on visual language models and large language models to obtain the subject category and feature category of the Chinese painting images. A multi-level semantic visualization device is used to construct and display a three-level feature clustering view, a semantic association view, a time evolution analysis view, and a detail presentation view based on the multi-level structural feature extraction results and the multi-level semantic extraction results.

10. The hierarchical semantic extraction and visualization system for traditional Chinese painting images based on a large model as described in claim 9, characterized in that, The multi-level semantic visualization device includes: The feature distribution view module is used to generate facial feature clustering subviews, posture feature clustering subviews, and co-occurrence feature clustering subviews based on the facial features, posture features, and co-occurrence features of figures in the Chinese painting image extracted by the multi-level structure feature extraction device. The natural language search module is used to generate a natural language search box and display the natural language search box in the three-layer feature clustering view, and perform feature matching based on the natural language text entered in the natural language search box, and return the most relevant set of Chinese painting images according to the feature matching results; The semantic association view generation module is used to generate the semantic association view based on the multi-level structural feature extraction results and the multi-level semantic extraction results. The time evolution analysis view generation module is used to generate the time evolution analysis view based on the multi-level structural feature extraction results, multi-level semantic extraction results, and the creation year of the Chinese painting. The detail presentation view generation module is used to generate the detail presentation view based on the multi-level semantic extraction results and the metadata of the Chinese painting.

Citation Information

Patent Citations

  • Face sketch generation method based on drawing stroke guidance

    CN112633288A

  • Personnel relation extraction and inference analysis method and system based on view knowledge graph

    CN118607626A

  • Image display system, image display device and program

    JP2013235596A