Image document body language multi-dimensional analysis method and system
By repairing incomplete images using a mirror-symmetric completion algorithm and a spatiotemporal weighted model, and combining dynamic sequence and propagation network analysis, the robustness and spatiotemporal analysis problems of the relationship between human body postures, language and symbols in early civilization image documents were solved, achieving high-precision reconstruction of cultural logic.
Patent Information
- Application Number
- CN202510556271.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing technologies lack robustness in analyzing the relationship between human postures, language, and symbols in early civilization image documents. They cannot repair incomplete images, ignore spatiotemporal dynamic factors, lack cross-modal verification, and have poor adaptability.
The missing key points are repaired using a mirror-symmetric completion algorithm. A dynamic sequence model is constructed by combining spatiotemporal weighted K-means clustering and Markov chain model. Association analysis is performed using chi-square test and PageRank algorithm to generate body language-symbol semantic mapping table and propagation heat map, and an interactive 3D visualization resource library is developed.
It achieves high-precision restoration of incomplete images, quantifies the probability of body movement transfer, constructs a symbolic propagation network, provides a quantifiable path for cultural logic reconstruction, and enhances the spatiotemporal context awareness of the analysis.
Smart Images

Figure CN120656174B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition and analysis technology, and more specifically, to a method and system for multi-dimensional analysis of body language in image documents. Background Technology
[0002] Currently, in the field of archaeology, the analysis of the relationship between human postures and symbols in early civilization image documents (such as bronze decorations and pictorial bricks) mainly relies on archaeologists' experience and manual statistical methods. Existing techniques typically use basic image processing tools (such as edge detection and color segmentation) to extract image features and combine them with simple frequency statistics (such as chi-square tests) to analyze the co-occurrence relationship between postures and symbols.
[0003] However, such methods have significant drawbacks: First, they lack robustness to incomplete or low-contrast archaeological images, failing to effectively repair missing key points (such as the missing arm of the Sanxingdui bronze figure); second, they ignore spatiotemporal dynamic factors (such as the time decay of cultural transmission and geographical barriers), causing symbolic association analysis to be detached from historical context; third, semantic understanding relies on subjective interpretation and lacks cross-modal verification between documentary text and image data. Although some studies have attempted to introduce traditional machine learning algorithms (such as SVM classification), they are poorly adapted to the specific characteristics of archaeological scenarios (small sample size, multimodal heterogeneous data) and struggle to quantify the temporal role of symbols in ritual action chains. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for multi-dimensional analysis of body language in image documents, in order to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows: Firstly, this application provides a method for multi-dimensional analysis of body language in image documents, including: The coordinates of key human joints in high-resolution image documents unearthed in the early Bashu region were detected. Missing joints were repaired using a mirror symmetry completion algorithm. The hand gesture opening and closing degree, torso tilt angle and gaze direction vector were calculated. The archaeological stratigraphic age label and site geographic coordinates were integrated to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 hand joints and 4 torso joints. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the axis of symmetry of the spine. Spatiotemporally weighted K-means clustering analysis was performed on the standardized body language spatiotemporal feature dataset to classify body feature categories, establish a morphological classification library, and use Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model; The morphological classification database and the acquired archaeological symbol data were correlated and analyzed. Based on the state transition probability constraint of the dynamic sequence model, the co-occurrence relationship was constrained. The chi-square test was used to screen significant correlation pairs. The semantic similarity was calculated by combining the semantic embedding model of literature, and a body language-symbol semantic mapping table was compiled. The spatiotemporal stability of symbols in the body language-symbolic semantic mapping table is evaluated by the entropy method. A propagation network model including cultural resistance coefficient is established, and the PageRank algorithm is applied to identify core propagation nodes and draw a symbol propagation heat map. By integrating a morphological classification database, a body language-symbol semantic mapping table, and a propagation heatmap, principal component features are extracted using singular value decomposition to develop an interactive three-dimensional visualization digital resource database, thus completing the spatiotemporal evolution analysis of body language in literature.
[0005] Preferably, the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period are detected, missing joints are repaired using a mirror symmetry completion algorithm, and gesture opening and closing degree, torso tilt angle, and gaze direction vector are calculated. This is then integrated with archaeological stratigraphic date tags and site geographic coordinates to form a standardized spatiotemporal feature dataset of body language, including: The coordinates of 25 key points were detected using an improved OpenPose network. Based on the set of joint coordinates, the missing joints on one side are completed along the axis of symmetry of the spine to obtain a complete set of joint coordinates. Based on the complete set of joint coordinates, the ratio of the Euclidean distance to the maximum physiological distance of the hand joints is calculated to obtain a normalized opening and closing value, which is denoted as the first posture parameter value. The tilt angle of the torso relative to the vertical direction is calculated using the vector angle between the lines connecting the shoulder and hip joints, and is denoted as the second posture parameter value. Based on the coordinate difference between the head and eye key points, a normalized direction vector is generated, which is denoted as the third posture parameter value. The first, second, and third posture parameter values are then combined into a posture parameter set. The obtained archaeological stratigraphic chronological labels and site geographic coordinates are fused to output spatiotemporal labels; the body posture parameter set and spatiotemporal labels are concatenated to obtain the original feature vector, and the original feature vector is standardized and optimized to obtain a standardized body posture spatiotemporal feature dataset.
[0006] Preferably, the step of performing spatiotemporally weighted K-means clustering analysis on the standardized body language spatiotemporal feature dataset to classify body language feature categories, establish a morphological classification library, and derive action transition probabilities using a Markov chain model to construct a dynamic sequence model includes: Based on a standardized spatiotemporal feature dataset, samples are weighted using a spatiotemporal weighting function. The time weight is calculated using an exponential decay method, while the spatial weight is calculated inversely proportional to the distance between sites. Using the weighted feature data as the target, clustering is performed by minimizing the weighted distance to generate a set of body feature categories. Each category is labeled with the range of typical morphological features and confidence samples. Extract the central feature vectors of each category and the top 5% of samples with the highest confidence in the body feature category set to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of trunk tilt angle, and the principal component of the gaze direction vector; Multiple consecutive images are extracted from high-resolution image literature, and after 3D scanning, a virtual continuous action sequence is generated, which is then arranged into a scene image sequence. Based on a structured classification library, the transition frequency between adjacent body posture categories is statistically analyzed, and the transition probability is calculated using Laplacian smoothing. Finally, a dynamic sequence model M=(P, S) is constructed, where... Let S be the state transition probability matrix, and S = {S1, ..., S10} be the set of body posture categories.
[0007] Preferably, the step of performing association analysis between the morphological classification library and the acquired archaeological symbol data, constraining co-occurrence relationships based on the state transition probability of a dynamic sequence model, using a chi-square test to screen significant association pairs, calculating semantic similarity using a document semantic embedding model, and compiling a body language-symbol semantic mapping table, including: Based on the morphological classification library and archaeological symbol dataset, we performed co-occurrence frequency statistics under adjacent state transition constraints using a dynamic sequence model. We then combined the state transition probabilities to weight the co-occurrence frequencies, resulting in a dynamically constrained co-occurrence frequency matrix. Using the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as inputs, an improved chi-square test method is employed. By introducing transition probabilities, the expected frequencies are dynamically weighted and adjusted to screen those meeting the significance threshold. >3.841 and transition probability constraint The set of association pairs with a value greater than 0.2 is then used to output a list of association pairs. Based on the list of associated pairs, a pre-trained language model is used to vectorize and embed body descriptions and symbolic semantics in historical documents. The basic semantic relevance is calculated by cosine similarity, and the state transition probability is introduced to weight and enhance the temporal contribution of symbols in the action chain, generating a fused body language-symbol semantic mapping table. The body language-symbol semantic mapping table includes body category, associated symbol, chi-square value, and enhanced similarity.
[0008] Preferably, the process of evaluating the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establishing a propagation network model including cultural resistance coefficients, and applying the PageRank algorithm to identify core propagation nodes and draw a symbol propagation heatmap includes: Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol in different sites and ages, output the spatiotemporal entropy table of symbols, and quantify the spatiotemporal stability of the propagation of each symbol. By combining the symbol spatiotemporal entropy value table and the collected site geographical environment data, the cultural resistance coefficient is calculated, the edge weights of the propagation network are constructed, and a symbol propagation network with cultural resistance is obtained. The edge weight matrix reflects the ease or difficulty of symbol propagation across sites. Using the edge weight matrix of the propagation network, the propagation influence of each site node is calculated through an improved PageRank algorithm; Spatial interpolation was performed using Gaussian kernel density estimation based on dissemination influence, PageRank value list, and site UTM coordinates. Output a symbol propagation heatmap and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color levels.
[0009] Secondly, this application also provides a multi-dimensional analysis system for body language in image documents, including: Integration Module: This module is used to detect the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period. It repairs missing joints using a mirror symmetry completion algorithm and calculates the hand gesture opening and closing degree, torso tilt angle, and gaze direction vector. It integrates archaeological stratigraphic age tags and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 hand joints and 4 torso joints. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the spinal axis of symmetry. Analysis and construction module: used to perform spatiotemporal weighted K-means clustering analysis on standardized body language spatiotemporal feature dataset, classify body feature categories, establish a morphological classification library, and use Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model; The association compilation module is used to perform association analysis between the morphological classification library and the acquired archaeological symbol data. It uses the chi-square test to screen significant association pairs, combines the semantic embedding model of literature to calculate semantic similarity, and compiles a body language-symbol semantic mapping table. Evaluation and mapping module: Used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establish a propagation network model including cultural resistance coefficients, and apply the PageRank algorithm to identify core propagation nodes and draw symbol propagation heatmaps; Extraction and Analysis Module: This module integrates a morphological classification library, a body language-symbol semantic mapping table, and a propagation heatmap. It extracts principal component features using singular value decomposition, develops an interactive three-dimensional visualization digital resource library, and completes the spatiotemporal evolution analysis of body language in literature.
[0010] Thirdly, this application also provides a device for multi-dimensional analysis of body language in image documents, including: Memory, used to store computer programs; A processor is used to implement the steps of the image document body language multidimensional analysis method when executing the computer program.
[0011] Fourthly, this application also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for multi-dimensional analysis of body language in image documents.
[0012] The beneficial effects of this invention are as follows: This invention designs a mirror-symmetric completion algorithm and a spatiotemporally weighted OpenPose model to solve the problem of high-precision restoration of missing key points in archaeological images. It combines UTM projection and stratigraphic data to construct standardized spatiotemporal features. Secondly, it develops a dynamic sequence Markov chain model to quantify the probability of body movement transitions and integrates cultural resistance coefficients to construct a symbolic propagation network, achieving spatiotemporal context-aware correlation analysis. Finally, it integrates BERT semantic embedding and kernel density estimation techniques to generate a body language-symbolic semantic mapping table and a propagation heatmap, forming a full-chain analysis framework of "image interpretation - dynamic modeling - semantic verification." This method is the first to systematically introduce deep learning, complex networks, and digital humanities methods into the interpretation of archaeological images, providing a quantifiable and verifiable technical path for reconstructing the cultural logic of pre-literate civilizations.
[0013] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1This is a schematic diagram of the multi-dimensional analysis method for image document body language described in this embodiment of the invention; Figure 2 This is a schematic diagram of the structure of the image document body language multi-dimensional analysis system described in this embodiment of the invention; Figure 3 This is a schematic diagram of the structure of the image document body language multi-dimensional analysis device described in this embodiment of the invention.
[0016] In the diagram: 701, Integration Module; 702, Analysis and Construction Module; 703, Association and Compilation Module; 704, Evaluation and Drawing Module; 705, Extraction and Analysis Module; 800, Multi-dimensional Analysis Device for Image Document Body Language; 801, Processor; 802, Memory; 803, Multimedia Component; 804, I / O Interface; 805, Communication Component. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0018] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0019] Example 1:
[0020] This embodiment provides a method for multi-dimensional analysis of body language in image documents.
[0021] See Figure 1 The figure shows that the method includes steps S100, S200, S300, S400 and S500.
[0022] S100. The coordinates of key human joints in high-resolution image documents unearthed in the early Bashu region are detected. Missing joints are repaired by a mirror symmetry completion algorithm. The hand gesture opening and closing degree, torso tilt angle and gaze direction vector are calculated. The archaeological stratigraphic age label and site geographical coordinates are integrated to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 joints of the hand and 4 joints of the torso. The mirror symmetry completion algorithm includes flipping the coordinates of the joints on the healthy side along the axis of symmetry of the spine.
[0023] It is understood that this step includes S101, S102, S103, and S104, wherein: S101. The coordinates of 25 key points are detected using an improved OpenPose network, and the calculation formula is as follows:
[0024] In the formula, J is the set of joint coordinates. The coordinates of the joint point image. To detect confidence levels, W is the width of the input image and H is the height of the input image; S102. Based on the set of joint coordinates, complete the missing joints on one side along the axis of symmetry of the spine to obtain a complete set of joint coordinates. The calculation formula for completing the missing joints on one side is as follows:
[0025] In the formula, The midpoint of the spine, The coordinates of the missing joints to be completed. These are the coordinates of the corresponding joints on the mirror side. For the left shoulder, For the right shoulder, Left hip, Right hip; S103. Based on the complete set of joint coordinates, calculate the ratio of the Euclidean distance to the maximum physiological distance of the hand joints to obtain a normalized opening and closing value, denoted as the first posture parameter value; calculate the tilt angle of the torso relative to the vertical direction using the vector angle between the lines connecting the shoulder and hip joints, denoted as the second posture parameter value; generate a normalized direction vector based on the coordinate difference between the head and eye key points, denoted as the third posture parameter value; summarize the first, second, and third posture parameter values into a posture parameter set. It should be noted that the process for calculating posture parameters is as follows: The Euclidean distance between the hand joints (e.g., the fingertip and palm) is calculated and divided by the maximum physiological distance (normalized to 200 pixels based on the size of the bronze figure) to obtain a normalized hand gesture opening value; the torso's tilt angle relative to the vertical direction is calculated using the vector angle between the lines connecting the shoulder and hip joints; a unitized direction vector is generated based on the coordinate difference between the head and key eye points to represent the direction of the gaze; and then, a set of posture parameters including hand gesture opening, torso tilt angle, and gaze direction vector is output. The formula for calculating hand gesture opening is as follows:
[0026] In the formula, The coordinates are the tips of the index finger. The coordinates are the palm center, and S represents the degree of hand opening / closing. Its torso tilt angle is: And take the average of the left and right sides, its The coordinates of the hip joint are given; the vector of its gaze direction is given. Therefore, considering the incomplete nature of archaeological images, geometric constraints based on the spine's axis of symmetry are used to complete the images and solve the problem of conventional pose estimation models failing under missing data.
[0027] S104. The obtained archaeological stratigraphic chronology labels and site geographic coordinates are fused to output spatiotemporal labels; the body posture parameter set and spatiotemporal labels are concatenated to obtain the original feature vector, and the original feature vector is standardized and optimized to obtain a standardized body posture spatiotemporal feature dataset.
[0028] Archaeological stratigraphic dating can be obtained in several ways, including extracting symbolic information from inscriptions, decorations, and patterns on artifacts unearthed from the same batch in the Sichuan Basin (such as Sanxingdui bronzes, Jinsha gold artifacts, and Han Dynasty pictorial bricks). This includes, but is not limited to: engraved symbols on the surface of artifacts (such as "feathered human figures" and "divine tree patterns"); graphic symbols in inscriptions; and decorative patterns on pictorial bricks. This includes: extracting symbolic features from archaeological excavation reports and artifact catalogues; using image recognition technology to extract symbolic features from artifact photographs; or combining professional annotations and classifications by archaeologists. During processing, the extracted symbols can be standardized and coded to establish a symbol classification system, and the excavation location, dating, and other contextual information of each symbol can be recorded. This symbolic data consists of archaeological materials unearthed at the same time and from the same source as the body language image documents, possessing a clear spatiotemporal correlation, rather than being pre-set theoretical symbols. The site's geographical coordinates are obtained by converting the site's GPS coordinates to Universal Transverse Mercator (UTM) projection coordinates, where UTM projection ensures spatial consistency between the geographical coordinates and the body language features.
[0029] The standardization and optimization of the original feature vectors, resulting in the standardized body language spatiotemporal feature dataset, is calculated column-by-column based on the feature values of all samples in the training set.
[0030] In the formula, N is the total number of samples, and i is the feature dimension.
[0031] In summary, the standardized body language spatiotemporal feature dataset generated through this process can simultaneously reflect the morphological features of human postures and the archaeological spatiotemporal context, laying a data foundation for subsequent multidimensional analysis.
[0032] S200. Perform spatiotemporal weighted K-means clustering analysis on the standardized body language spatiotemporal feature dataset to classify body feature categories, establish a morphological classification library, and use the Markov chain model to derive the action transition probability, thereby constructing a dynamic sequence model.
[0033] It is understood that this step includes S201, S202, and S203, wherein: S201. Based on the standardized spatiotemporal feature dataset, the samples are weighted by a spatiotemporal weight function. The time weight is calculated in an exponential decay form, while the spatial weight is calculated inversely proportional to the distance between sites. Taking the weighted feature data as the target, clustering is performed by minimizing the weighted distance to generate a set of body feature categories. Each category is labeled with the range of typical morphological features and confidence samples. It should be noted that the time weight is calculated using an exponential decay formula: wt = e, where λ = 0.01 is the time decay coefficient. This is the latest date reference value.
[0034] S202. Extract the central feature vector of each category in the body posture feature category set and the top 5% of samples with the highest confidence to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of trunk tilt angle and the principal component of the gaze direction vector. It should be noted that continuous motion images are extracted from a series of cultural relics in a certain scene (such as Han Dynasty portrait bricks and bronze decorations). For example, a Han Dynasty portrait brick depicts continuous dance movements. After 3D scanning of incomplete cultural relics (such as Sanxingdui bronzes), a virtual continuous motion sequence is generated. Then, according to the narrative order of the surface decorations of the artifacts or the scene reconstruction by archaeologists, the image sequence is arranged, and the long sequence is divided into short motion units (3-5 frames per unit) according to the posture changes, and then the scene image sequence is output.
[0035] S203. Extract multiple consecutive images from high-resolution image literature, perform 3D scanning, generate a virtual continuous action sequence, and then arrange it into a scene image sequence; based on a structured classification library, count the transition frequency between adjacent body posture categories, and calculate the transition probability through Laplace smoothing, finally constructing a dynamic sequence model M=(P, S), where Let S be the state transition probability matrix, and S = {S1, ..., S10} be the set of body posture categories.
[0036] Understandably, in this step, after 3D scanning of incomplete artifacts (such as bronze figures with only one-sided movements), tools like Blender / Maya are used to mirror and complete the missing parts, generating a complete 3D model. Alternatively, keyframe interpolation techniques can be used to generate virtual action sequences (such as a smooth transition from "standing" to "kneeling"). Then, the image sequence is arranged according to archaeological narrative logic (such as the reading order of artifact decorations and the chronological relationship of stratigraphy), and divided into units according to action changes. For example, the "holding an object high" action of the Sanxingdui K2②:35 bronze standing figure is decomposed into three sub-actions: "raising hand → extending arm → fixing posture." Each posture category in the morphology classification library (such as C1: "both arms raised high") directly corresponds to a state S1 of the Markov chain. The feature range provided by the classification library (such as S∈[0.7,0.9]) is used for state determination thresholds. In this step, the constructed Markov chain model not only reflects the actual observed action transitions but also restores the underlying logic of cultural transmission through spatiotemporal weights and smoothing strategies.
[0037] S300. The morphological classification library and the acquired archaeological symbol data are correlated and analyzed. Based on the state transition probability constraint of the dynamic sequence model, the co-occurrence relationship is constrained. The chi-square test is used to screen significant correlation pairs. The semantic similarity is calculated by combining the semantic embedding model of literature, and a body language-symbol semantic mapping table is compiled.
[0038] It is understood that step S300 includes S301, S302, and S303, wherein: S301. Based on the morphological classification library and archaeological symbol dataset, the co-occurrence frequency is statistically processed under adjacent state transition constraints through a dynamic sequence model. The co-occurrence frequency is weighted by combining the state transition probability to obtain the co-occurrence frequency matrix under dynamic constraints. S302. Using the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, an improved chi-square test method is employed. By introducing transition probabilities, the expected frequencies are dynamically weighted and adjusted to screen those meeting the significance threshold. >3.841 and transition probability constraint The set of association pairs with a value greater than 0.2 is then used to output a list of association pairs. It should be noted that only the posture-symbol co-occurrence of adjacent states within the same action chain is counted. If the symbol If it appears in state Sk, then only if When associated with a predecessor state Si or successor state Sl where the transition probability pi→k>0.1, record the body type Ci / Cl and The co-occurrence of these elements is calculated using the following formula with weighted frequency:
[0039] In the formula, Symbols In state The value is 1 if it appears in the condition, and 0 otherwise. For the transition from state Si to The transition probabilities are calculated, and the output is a dynamically constrained co-occurrence frequency matrix.
[0040] Understandably, traditional chi-square tests only calculate expected values based on frequency, but in archaeological scenarios, the association between symbols and body postures must conform to the temporal logic of actions. For example, if the symbol It only appears in low-probability transition paths (such as...) <0.1), their co-occurrence may be due to accidental contamination (such as the inclusion of fragments), and should be weighted down. In the above steps, the observed values Oij in the traditional chi-square formula need to be weighted by transition probability. If wij→0, even if Oij is high, the χ² value will be suppressed. Among them, the screening conditions are: Significance threshold: χ² > 3.841 (α = 0.05, degrees of freedom = 1); lower bound of transition probability: pi→j > 0.2, excluding noise associations in low-probability paths. Therefore, state transition probability is introduced as a frequency weight to ensure that the statistical test results are consistent with the logic of cultural behavior. For example, the "divine tree pattern" is strengthened in early high-probability transition paths, while the same symbol is weakened in later low-probability paths. That is, the Bashu model... Compared with the Shang-Zhou model, if a certain symbol (such as "dragon pattern") is prominent only in the high-probability path of the Shang-Zhou period, it may reflect the functional evolution of its cultural transmission.
[0041] S303. Based on the list of associated pairs, a pre-trained language model is used to vectorize and embed the body descriptions and symbolic semantics in historical documents. The basic semantic association degree is calculated by cosine similarity, and the state transition probability is introduced to weight and enhance the temporal contribution of the symbol in the action chain, generating a fused body language-symbol semantic mapping table, which includes body category, associated symbol, chi-square value and enhanced similarity.
[0042] It should be noted that semantic embedding and basic similarity calculation are performed by extracting content containing body shapes from the literature. and symbols The descriptive paragraphs (such as "raising hands to honor the sacred tree") are used to generate paragraph-level semantic vectors using the BERT model. The dimension is 768. If a symbol appears frequently in subsequent action chains, its semantic association needs to be enhanced. A temporal contribution factor is defined, and the similarity is enhanced accordingly. Then, a selection threshold is applied to retain... The association pairs are arranged in descending order of chi-square value. Therefore, combining textual semantics (literature description) with action sequence (transition probability) overcomes the limitations of single-modality analysis. For example, the symbolic meaning not explicitly stated in the literature (such as "feathered human pattern") can be revealed through its frequent occurrence in the action chain. Gao was inferred to be "the guide of the gods".
[0043] In summary, by deeply coupling dynamic sequence models with multimodal data analysis, we can eliminate accidental associations that do not conform to ritual logic, integrate textual descriptions and action sequences, improve the credibility of symbolic meaning analysis, and provide quantitative evidence for archaeological controversies such as symbolic function differentiation and ritual stage division.
[0044] S400. The spatiotemporal stability of symbols in the body language-symbol semantic mapping table is evaluated by the entropy method. A propagation network model including cultural resistance coefficient is established, and the PageRank algorithm is applied to identify core propagation nodes and draw a symbol propagation heat map.
[0045] It is understood that step S400 includes S401, S402, S403, S404, and S405, wherein: The PageRank algorithm is used to identify core propagation nodes, and a symbol propagation heatmap is drawn, including: S401. Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol in different sites and ages, output the symbol spatiotemporal entropy table, and quantify the spatiotemporal stability of each symbol's propagation. S402. Combining the symbol spatiotemporal entropy value table and the collected site geographical environment data, calculate the cultural resistance coefficient, construct the edge weights of the propagation network, and obtain the symbol propagation network with cultural resistance, where the edge weight matrix reflects the ease or difficulty of symbol propagation across sites. S403. Using the edge weight matrix of the propagation network, the propagation influence of each site node is calculated using the improved PageRank algorithm. S404. Based on the propagation influence, PageRank value list, and site UTM coordinates, spatial interpolation is performed using Gaussian kernel density estimation, and the calculation formula is as follows:
[0046] In the formula, Here, represents the estimated propagation intensity at location (x, y), N is the total number of archaeological nodes, and h is the bandwidth parameter. Let K represent the propagation influence of the archaeological site nodes, and K be the Gaussian kernel function. Let be the Euclidean distance between the target point (x,y) and the site vi; S405 outputs a symbol propagation heatmap and visualizes the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color levels.
[0047] It should be noted that the Silverman criterion is used for adaptive calculation to select the bandwidth. For example, if the standard deviation of the site distribution is 50km, then h is approximately 45km. The study area is divided into a 1000×1000 grid, and f(x,y) is calculated for each grid point (x,y). In this embodiment, the grid is locally refined in high-density areas such as the Chengdu Plain. Then, the intensity is normalized, and a color gradation is designed, such as a red-yellow-blue gradient: red (RGB 255,0,0). >0.8 (core propagation area); Yellow (RGB 255,255,0): 0.3< ≤0.8; Blue (RGB 0,0,255): ≤0.3 (edge area), and overlay the site location (black dot), mountains (dark gray fill), and rivers (blue line) on the map.
[0048] Using S500, a comprehensive morphological classification library, a body language-symbol semantic mapping table, and a propagation heatmap, principal component features are extracted through singular value decomposition to develop an interactive three-dimensional visualization digital resource library, thus completing the spatiotemporal evolution analysis of body language in literature.
[0049] Understandably, in this step, based on the morphological classification library, body language-symbol semantic mapping table, and propagation heatmap data, the principal component features of the joint feature matrix are extracted using Singular Value Decomposition (SVD). A three-dimensional principal component space is constructed (X / Y / Z axes correspond to the three principal components respectively, and the time dimension is mapped to chronological changes through color gradients). An interactive visualization engine is developed based on WebGL and Three.js to achieve virtual reconstruction of action chains (such as the dynamic sequence of "both arms raised → kneeling → turning around") and spatiotemporal overlay of symbol propagation heatmaps (Gaussian kernel density estimation generates red-...). The system utilizes a blue gradient color scheme and multi-dimensional data linkage (clicking on a heatmap node highlights related literature paragraphs); further, it employs B-spline curve fitting to model cultural evolution trajectories, and detects key abrupt change periods (such as the decline of the Sanxingdui culture at BC1000±50) through curvature extreme values. Ultimately, it forms a digital resource library that supports VR immersive observation, cross-period comparison (such as the difference in action chains between the Shang and Zhou dynasties and the Han dynasty), and sensitivity simulation (modifying geographical resistance parameters to predict propagation paths). This provides an intelligent platform for the study of civilizations without written language, combining quantitative analysis and scene reconstruction capabilities, breaking through the limitations of traditional archaeology that relies on static charts and empirical speculation.
[0050] Example 2:
[0051] like Figure 2 As shown, this embodiment provides a multi-dimensional body language analysis system for image documents. See [link / reference]. Figure 2 The system includes: Integration Module 701: This module is used to detect the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period. It repairs missing joints using a mirror symmetry completion algorithm and calculates the hand gesture opening and closing degree, torso tilt angle, and gaze direction vector. It integrates archaeological stratigraphic age tags and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 hand joints and 4 torso joints. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the symmetry axis of the spine. Analysis module 702: is used to perform spatiotemporal weighted K-means clustering analysis on the standardized body language spatiotemporal feature dataset, classify body feature categories, establish a morphological classification library, and use the Markov chain model to derive the action transition probability, thereby constructing a dynamic sequence model; The association compilation module 703 is used to perform association analysis between the morphological classification library and the acquired archaeological symbol data, use the chi-square test to screen significant association pairs, combine the semantic embedding model of literature to calculate semantic similarity, and compile a body language-symbol semantic mapping table. Evaluation and drawing module 704: Used to evaluate the spatiotemporal stability of symbols in the body language-symbolic semantic mapping table using the entropy method, establish a propagation network model including cultural resistance coefficients, and apply the PageRank algorithm to identify core propagation nodes and draw a symbol propagation heatmap; Extraction and Analysis Module 705: Used to integrate the morphological classification library, body language-symbol semantic mapping table and propagation heat map, extract principal component features through singular value decomposition, develop an interactive three-dimensional visualization digital resource library, and complete the spatiotemporal evolution analysis of body language in documents.
[0052] Specifically, the integration module 701 includes: Detection Unit: Used to detect the coordinates of 25 key points using an improved OpenPose network. Its calculation formula is as follows:
[0053] In the formula, J is the set of joint coordinates. The coordinates of the joint point image. To detect confidence levels, W is the width of the input image and H is the height of the input image; The completion unit is used to complete the set of joint coordinates along the axis of symmetry of the spine for unilateral missing joints, resulting in a complete set of joint coordinates. The calculation formula for completing unilateral missing joints is as follows:
[0054] In the formula, The midpoint of the spine, The coordinates of the missing joints to be completed. These are the coordinates of the corresponding joints on the mirror side. For the left shoulder, For the right shoulder, Left hip, Right hip; The first calculation unit is used to calculate the ratio of the Euclidean distance to the maximum physiological distance of the hand joints based on the complete set of joint coordinates, obtaining a normalized opening and closing value, denoted as the first posture parameter value; it calculates the tilt angle of the torso relative to the vertical direction through the vector angle between the lines connecting the shoulder and hip joints, denoted as the second posture parameter value; it generates a normalized direction vector based on the coordinate difference between the head and eye key points, denoted as the third posture parameter value; and it summarizes the first, second, and third posture parameter values into a posture parameter set. The optimization unit is used to fuse the obtained archaeological stratigraphic chronological labels and site geographic coordinates to output spatiotemporal labels; it also concatenates the body posture parameter set and spatiotemporal labels to obtain the original feature vector, and then standardizes and optimizes the original feature vector to obtain a standardized body posture spatiotemporal feature dataset.
[0055] Specifically, the analysis construction module 702 includes: The second computing unit is used to weight samples based on a standardized spatiotemporal feature dataset using a spatiotemporal weighting function. The time weight is calculated using an exponential decay method, while the spatial weight is calculated inversely proportional to the distance between sites. Using the weighted feature data as the target, clustering is performed by minimizing the weighted distance to generate a set of body feature categories. Each category is labeled with the range of typical morphological features and confidence samples. Establishment Unit: Used to extract the central feature vector of each category in the body posture feature category set and the top 5% of samples with the highest confidence to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of trunk tilt angle and the principal component of the gaze direction vector; Extraction Unit: This unit extracts multiple consecutive images from high-resolution image documents, performs 3D scanning, generates a virtual continuous action sequence, and then arranges them into a scene image sequence. Based on a structured classification library, it statistically analyzes the transition frequency between adjacent body posture categories and calculates the transition probability using Laplacian smoothing. Finally, it constructs a dynamic sequence model M=(P, S), where... Let S be the state transition probability matrix, and S = {S1, ..., S10} be the set of body posture categories.
[0056] Specifically, the association compilation module 703 includes: Processing unit: Based on the morphological classification library and archaeological symbol dataset, it performs co-occurrence frequency statistics under adjacent state transition constraints using a dynamic sequence model, and calculates the co-occurrence frequency by weighting the co-occurrence frequency with the state transition probability to obtain the co-occurrence frequency matrix under dynamic constraints. The correction unit uses the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as inputs. It employs an improved chi-square test method, introducing transition probabilities to dynamically weight the desired frequencies, and then filters those meeting the significance threshold. >3.841 and transition probability constraint The set of association pairs with a value greater than 0.2 is then used to output a list of association pairs. The generation unit is used to vectorize and embed body descriptions and symbolic semantics in historical documents based on a list of associated pairs using a pre-trained language model. It calculates the basic semantic association degree through cosine similarity and introduces state transition probability to weight and enhance the temporal contribution of symbols in the action chain, generating a fused body language-symbol semantic mapping table. The body language-symbol semantic mapping table includes body category, associated symbol, chi-square value, and enhanced similarity.
[0057] Specifically, the evaluation drawing module 704 includes: Quantization Unit: Based on the symbol distribution data in the body language-symbol semantic mapping table, it uses the entropy method to calculate the frequency distribution dispersion of each symbol in different sites and ages, outputs a symbol spatiotemporal entropy table, and quantifies the spatiotemporal stability of each symbol's propagation. Construction Unit: Used to combine the symbol spatiotemporal entropy value table and the collected site geographical environment data to calculate the cultural resistance coefficient, construct the edge weights of the propagation network, and obtain the symbol propagation network with cultural resistance. The edge weight matrix reflects the ease or difficulty of symbol propagation across sites. The third calculation unit is used to calculate the propagation influence of each site node using the edge weight matrix of the propagation network and the improved PageRank algorithm. The fourth calculation unit is used to perform spatial interpolation based on the propagation influence, the PageRank value list, and the site's UTM coordinates, using Gaussian kernel density estimation. Its calculation formula is as follows:
[0058] In the formula, Here, represents the estimated propagation intensity at location (x, y), N is the total number of archaeological nodes, and h is the bandwidth parameter. Let K represent the propagation influence of the archaeological site nodes, and K be the Gaussian kernel function. Let be the Euclidean distance between the target point (x,y) and the site vi; Diffusion unit: Used to output a symbol propagation heatmap and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color levels.
[0059] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0060] Example 3:
[0061] Corresponding to the above method embodiments, this embodiment also provides an image document body language multi-dimensional analysis device. The image document body language multi-dimensional analysis device described below and the image document body language multi-dimensional analysis method described above can be referred to in correspondence.
[0062] Figure 3 This is a block diagram illustrating a multi-dimensional analysis device 800 for image document body language according to an exemplary embodiment. For example... Figure 3 As shown, the image document body language multi-dimensional analysis device 800 includes a processor 801 and a memory 802. The image document body language multi-dimensional analysis device 800 also includes one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0063] The processor 801 controls the overall operation of the image document body language multi-dimensional analysis device 800 to complete all or part of the steps in the aforementioned image document body language multi-dimensional analysis method. The memory 802 stores various types of data to support the operation of the image document body language multi-dimensional analysis device 800. This data may include, for example, instructions for any application or method operating on the image document body language multi-dimensional analysis device 800, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, or buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the image document body language multidimensional analysis device 800 and other devices. Wireless communication includes, for example, Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or one or more combinations thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0064] In an exemplary embodiment, the image document body language multi-dimensional analysis device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described image document body language multi-dimensional analysis method.
[0065] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the above-described image document body language multi-dimensional analysis method. For example, the computer-readable storage medium may be the memory 802 including the program instructions, which may be executed by the processor 801 of the image document body language multi-dimensional analysis device 800 to complete the above-described image document body language multi-dimensional analysis method.
[0066] Example 4:
[0067] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the image document body language multi-dimensional analysis method described above.
[0068] A computer program is stored on a readable storage medium, and when the computer program is executed by a processor, it implements the steps of the image document body language multi-dimensional analysis method of the above method embodiments.
[0069] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0071] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for multi-dimensional analysis of body language in image documents, characterized in that, include: The coordinates of key human joints in high-resolution image documents unearthed in the early Bashu region were detected. Missing joints were repaired using a mirror symmetry completion algorithm. The hand gesture opening and closing degree, torso tilt angle and gaze direction vector were calculated. The archaeological stratigraphic age label and site geographic coordinates were integrated to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 hand joints and 4 torso joints. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the axis of symmetry of the spine. Spatiotemporally weighted K-means clustering analysis was performed on the standardized body language spatiotemporal feature dataset to classify body feature categories, establish a morphological classification library, and use Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model; The morphological classification database and the acquired archaeological symbol data were correlated and analyzed. Based on the state transition probability constraint of the dynamic sequence model, the co-occurrence relationship was constrained. The chi-square test was used to screen significant correlation pairs. The semantic similarity was calculated by combining the semantic embedding model of literature, and a body language-symbol semantic mapping table was compiled. The spatiotemporal stability of symbols in the body language-symbolic semantic mapping table is evaluated by the entropy method. A propagation network model including cultural resistance coefficient is established, and the PageRank algorithm is applied to identify core propagation nodes and draw a symbol propagation heat map. By integrating a morphological classification database, a body language-symbol semantic mapping table, and a propagation heatmap, principal component features are extracted using singular value decomposition to develop an interactive three-dimensional visualization digital resource database, thus completing the spatiotemporal evolution analysis of body language in literature.
2. The method for multi-dimensional analysis of body language in image documents according to claim 1, characterized in that, The method involves detecting the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period, repairing missing joints using a mirror-symmetric completion algorithm, and calculating gesture opening and closing, torso tilt angle, and gaze direction vector. This data is then integrated with archaeological stratigraphic date tags and site geographic coordinates to form a standardized spatiotemporal feature dataset of body language, including: The coordinates of 25 key points are detected using an improved OpenPose network, and the calculation formula is as follows: In the formula, J is the set of joint coordinates, (x i y i ) represents the coordinates of the joint image, c i To detect confidence, W is the width of the input image and H is the height of the input image; Based on the set of joint coordinates, the missing joints on one side are completed along the axis of symmetry of the spine to obtain a complete set of joint coordinates. The calculation formula for completing the missing joints on one side is as follows: In the formula, p s p is the midpoint of the spine. missing p represents the coordinates of the missing joints to be filled. mirror p represents the coordinates of the corresponding joint on the mirror side. ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh Right hip; Based on the complete set of joint coordinates, the ratio of the Euclidean distance to the maximum physiological distance of the hand joints is calculated to obtain a normalized opening and closing value, which is denoted as the first posture parameter value. The tilt angle of the torso relative to the vertical direction is calculated using the vector angle between the lines connecting the shoulder and hip joints, and is denoted as the second posture parameter value. Based on the coordinate difference between the head and eye key points, a normalized direction vector is generated, which is denoted as the third posture parameter value. The first, second, and third posture parameter values are then combined into a posture parameter set. The obtained archaeological stratigraphic chronological labels and site geographic coordinates are fused to output spatiotemporal labels; the body posture parameter set and spatiotemporal labels are concatenated to obtain the original feature vector, and the original feature vector is standardized and optimized to obtain a standardized body posture spatiotemporal feature dataset.
3. The method for multi-dimensional analysis of body language in image documents according to claim 1, characterized in that, The standardized body language spatiotemporal feature dataset is subjected to spatiotemporal weighted K-means clustering analysis to classify body feature categories, establish a morphological classification library, and use a Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model, including: Based on a standardized spatiotemporal feature dataset, samples are weighted using a spatiotemporal weighting function. The time weight is calculated using an exponential decay method, while the spatial weight is calculated inversely proportional to the distance between sites. Using the weighted feature data as the target, clustering is performed by minimizing the weighted distance to generate a set of body feature categories. Each category is labeled with the range of typical morphological features and confidence samples. Extract the central feature vectors of each category and the top 5% of samples with the highest confidence in the body feature category set to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of trunk tilt angle, and the principal components of the gaze direction vector; Multiple consecutive images are extracted from high-resolution image literature, and after 3D scanning, a virtual continuous action sequence is generated, which is then arranged into a scene image sequence. Based on a structured classification library, the transition frequency between adjacent body posture categories is statistically analyzed, and the transition probability is calculated using Laplacian smoothing. Finally, a dynamic sequence model M = (P, S) is constructed, where... Let S be the state transition probability matrix, and let S = {S1, ..., S10} be the set of body posture categories.
4. The method for multi-dimensional analysis of body language in image documents according to claim 1, characterized in that, The process involves correlation analysis between the morphological classification library and the acquired archaeological symbol data. Based on the state transition probability constraint of a dynamic sequence model, co-occurrence relationships are constrained. A chi-square test is used to screen significant association pairs. Semantic similarity is calculated using a document semantic embedding model, and a body language-symbol semantic mapping table is compiled, including: Based on the morphological classification library and archaeological symbol dataset, we performed co-occurrence frequency statistics under adjacent state transition constraints using a dynamic sequence model. We then combined the state transition probabilities to weight the co-occurrence frequencies, resulting in a dynamically constrained co-occurrence frequency matrix. Using the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as inputs, an improved chi-square test method is employed. By introducing transition probabilities, the expected frequencies are dynamically weighted and adjusted to screen those meeting the significance threshold χ². 2 >3.841 and the transition probability is constrained to p ij The set of association pairs with a value greater than 0.2 is then used to output a list of association pairs. Based on the list of associated pairs, a pre-trained language model is used to vectorize and embed body descriptions and symbolic semantics in historical documents. The basic semantic relevance is calculated by cosine similarity, and the state transition probability is introduced to weight and enhance the temporal contribution of symbols in the action chain, generating a fused body language-symbol semantic mapping table. The body language-symbol semantic mapping table includes body category, associated symbol, chi-square value, and enhanced similarity.
5. The method for multi-dimensional analysis of body language in image documents according to claim 1, characterized in that, The method involves evaluating the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establishing a propagation network model that includes cultural resistance coefficients, and applying the PageRank algorithm to identify core propagation nodes and draw a symbol propagation heatmap, including: Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol in different sites and ages, output the spatiotemporal entropy table of symbols, and quantify the spatiotemporal stability of the propagation of each symbol. By combining the symbol spatiotemporal entropy value table and the collected site geographical environment data, the cultural resistance coefficient is calculated, the edge weights of the propagation network are constructed, and a symbol propagation network with cultural resistance is obtained. The edge weight matrix reflects the ease or difficulty of symbol propagation across sites. Using the edge weight matrix of the propagation network, the propagation influence of each site node is calculated through an improved PageRank algorithm; Based on the dissemination influence, PageRank value list, and site UTM coordinates, spatial interpolation is performed using Gaussian kernel density estimation, and the calculation formula is as follows: In the formula, f(x, y) is the estimated propagation intensity at location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) represents the propagation influence of the site nodes, and K is the Gaussian kernel function. Let be the Euclidean distance between the target point (x,y) and the site vi; Output a symbol propagation heatmap and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color levels.
6. A multi-dimensional body language analysis system for image documents, based on the multi-dimensional body language analysis method for image documents as described in claim 1, characterized in that, include: Integration Module: This module is used to detect the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period. It repairs missing joints using a mirror symmetry completion algorithm and calculates the hand gesture opening and closing degree, torso tilt angle, and gaze direction vector. It integrates archaeological stratigraphic age tags and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 hand joints and 4 torso joints. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the spinal axis of symmetry. Analysis and construction module: used to perform spatiotemporal weighted K-means clustering analysis on standardized body language spatiotemporal feature dataset, classify body feature categories, establish a morphological classification library, and use Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model; The association compilation module is used to perform association analysis between the morphological classification library and the acquired archaeological symbol data. It uses the chi-square test to screen significant association pairs, combines the semantic embedding model of literature to calculate semantic similarity, and compiles a body language-symbol semantic mapping table. Evaluation and mapping module: Used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establish a propagation network model including cultural resistance coefficients, and apply the PageRank algorithm to identify core propagation nodes and draw symbol propagation heatmaps; Extraction and Analysis Module: This module integrates a morphological classification library, a body language-symbol semantic mapping table, and a propagation heatmap. It extracts principal component features using singular value decomposition, develops an interactive three-dimensional visualization digital resource library, and completes the spatiotemporal evolution analysis of body language in literature.
7. The image document body language multi-dimensional analysis system according to claim 6, characterized in that, The integration module includes: Detection Unit: Used to detect the coordinates of 25 key points using an improved OpenPose network. Its calculation formula is as follows: In the formula, J is the set of joint coordinates, (x i y i ) represents the coordinates of the joint image, c i To detect confidence, W is the width of the input image and H is the height of the input image; The completion unit is used to complete the set of joint coordinates along the axis of symmetry of the spine for unilateral missing joints, resulting in a complete set of joint coordinates. The calculation formula for completing unilateral missing joints is as follows: In the formula, p s p is the midpoint of the spine. missing p represents the coordinates of the missing joints to be filled. mirror p represents the coordinates of the corresponding joint on the mirror side. ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh Right hip; The first calculation unit is used to calculate the ratio of the Euclidean distance to the maximum physiological distance of the hand joints based on the complete set of joint coordinates, obtaining a normalized opening and closing value, denoted as the first posture parameter value; it calculates the tilt angle of the torso relative to the vertical direction through the vector angle between the lines connecting the shoulder and hip joints, denoted as the second posture parameter value; it generates a normalized direction vector based on the coordinate difference between the head and eye key points, denoted as the third posture parameter value; and it summarizes the first, second, and third posture parameter values into a posture parameter set. The optimization unit is used to fuse the obtained archaeological stratigraphic chronological labels and site geographic coordinates to output spatiotemporal labels; it also concatenates the body posture parameter set and spatiotemporal labels to obtain the original feature vector, and then standardizes and optimizes the original feature vector to obtain a standardized body posture spatiotemporal feature dataset.
8. The image document body language multi-dimensional analysis system according to claim 6, characterized in that, The analysis building module includes: The second computing unit is used to weight samples based on a standardized spatiotemporal feature dataset using a spatiotemporal weighting function. The time weight is calculated using an exponential decay method, while the spatial weight is calculated inversely proportional to the distance between sites. Using the weighted feature data as the target, clustering is performed by minimizing the weighted distance to generate a set of body feature categories. Each category is labeled with the range of typical morphological features and confidence samples. Establishment Unit: Used to extract the central feature vector of each category in the body posture feature category set and the top 5% of samples with the highest confidence to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of trunk tilt angle and the principal component of the gaze direction vector; Extraction Unit: This unit extracts multiple consecutive images from high-resolution image documents, performs 3D scanning, generates a virtual continuous action sequence, and then arranges them into a scene image sequence. Based on a structured classification library, it statistically analyzes the transition frequency between adjacent body posture categories and calculates the transition probability using Laplacian smoothing. Finally, it constructs a dynamic sequence model M = (P, S), where... Let S be the state transition probability matrix, and let S = {S1, ..., S10} be the set of body posture categories.
9. The image document body language multi-dimensional analysis system according to claim 6, characterized in that, The association compilation module includes: Processing unit: Based on the morphological classification library and archaeological symbol dataset, it performs co-occurrence frequency statistics under adjacent state transition constraints using a dynamic sequence model, and calculates the co-occurrence frequency by weighting the co-occurrence frequency with the state transition probability to obtain the co-occurrence frequency matrix under dynamic constraints. The correction unit uses the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as inputs. It employs an improved chi-square test method, dynamically weighting the desired frequencies by introducing transition probabilities, and then filters those that meet the significance threshold χ². 2 >3.841 and the transition probability is constrained to p ij The set of association pairs with a value greater than 0.2 is then used to output a list of association pairs. The generation unit is used to vectorize and embed body descriptions and symbolic semantics in historical documents based on a list of associated pairs using a pre-trained language model. It calculates the basic semantic association degree through cosine similarity and introduces state transition probability to weight and enhance the temporal contribution of symbols in the action chain, generating a fused body language-symbol semantic mapping table. The body language-symbol semantic mapping table includes body category, associated symbol, chi-square value, and enhanced similarity.
10. The image document body language multi-dimensional analysis system according to claim 6, characterized in that, The evaluation drawing module includes: Quantization Unit: Based on the symbol distribution data in the body language-symbol semantic mapping table, it uses the entropy method to calculate the frequency distribution dispersion of each symbol in different sites and ages, outputs a symbol spatiotemporal entropy table, and quantifies the spatiotemporal stability of each symbol's propagation. Construction Unit: Used to combine the symbol spatiotemporal entropy value table and the collected site geographical environment data to calculate the cultural resistance coefficient, construct the edge weights of the propagation network, and obtain the symbol propagation network with cultural resistance. The edge weight matrix reflects the ease or difficulty of symbol propagation across sites. The third calculation unit is used to calculate the propagation influence of each site node using the edge weight matrix of the propagation network and the improved PageRank algorithm. The fourth calculation unit is used to perform spatial interpolation based on the propagation influence, the PageRank value list, and the site's UTM coordinates, using Gaussian kernel density estimation. Its calculation formula is as follows: In the formula, f(x, y) is the estimated propagation intensity at location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) represents the propagation influence of the site nodes, and K is the Gaussian kernel function. Let be the Euclidean distance between the target point (x,y) and the site vi; Diffusion unit: Used to output a symbol propagation heatmap and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color levels.
Citation Information
Patent Citations
Cultural relic image retrieval method and device based on multi-modal representation learning model
CN117473114A
Character action recognition analysis method and system based on infrared laser and deep learning
CN118747911A