Image literature posture language multi-dimensional analysis method and system

By using the mirror symmetry completion algorithm and the spatiotemporal weighted K-means clustering model, combined with the Markov chain and PageRank algorithms, the robustness problem of the association between human body language and symbols in early civilization image documents was solved, high-precision restoration and quantitative analysis were achieved, and a verifiable research method for non-written civilizations was provided.

CN120656174AActive Publication Date: 2025-09-16SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510556271.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-16
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing technologies lack robustness when analyzing the association between human body language and symbols in early civilization image documents. They are unable to repair incomplete images, ignore spatiotemporal dynamic factors, and rely on subjective interpretation, making it difficult to quantify the temporal role of symbols in the ritual action chain.

Method used

A mirror symmetry completion algorithm is used to repair missing joints. A dynamic sequence model is constructed by combining spatiotemporal weighted K-means clustering and Markov chain model. The PageRank algorithm is used to identify propagation nodes, generate a body language-symbol semantic mapping table and a propagation heat map, extract principal component features through the singular value decomposition method, and develop an interactive three-dimensional visualization resource library.

Benefits of technology

It achieves high-precision restoration of incomplete images, quantifies the probability of body movement transfer, builds a verifiable symbol communication network, and provides a quantifiable technical path for reconstructing the logic of non-written civilization and culture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656174A_ABST
    Figure CN120656174A_ABST
Patent Text Reader

Abstract

The invention provides an image literature posture language multi-dimensional analysis method and system, and relates to the technical field of image recognition and analysis, and the method comprises the steps: repairing missing joint points through a mirror symmetry completion algorithm, integrating archaeological stratigraphic age labels and ruins geographic coordinates, and forming a standardized posture language spatio-temporal feature data set; establishing a morphological classification library, and constructing a dynamic sequence model; performing association analysis on the form classification library and the obtained archaeological symbol data, screening significance association pairs by adopting chi-square test, and compiling a posture language-symbol semantic mapping table; evaluating the space-time stability of symbols in the posture language-symbol semantic mapping table through an entropy evaluation method, and drawing a symbol propagation thermodynamic diagram; and integrating the form classification library, the posture language-symbol semantic mapping table and the propagation thermodynamic diagram to complete spatio-temporal evolution analysis of the literature posture language. The beneficial effects of the invention are that deep learning, a complex network and a digital humanity method system are introduced into archaeological image interpretation for the first time, and a quantifiable and verifiable technical path is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition and analysis, and in particular to a method and system for multi-dimensional analysis of body language in image documents. Background Art

[0002] Currently, in the field of archaeology, the analysis of the relationship between body language and symbols in early civilization imagery (such as bronze ornamentation and brick carvings) relies primarily on archaeologists' empirical judgment and manual statistical methods. Existing techniques typically use basic image processing tools (such as edge detection and color segmentation) to extract image features, combined with simple frequency statistics (such as the chi-square test) to analyze the co-occurrence relationship between body language and symbols.

[0003] However, such methods have significant drawbacks: First, they lack robustness against incomplete, low-contrast archaeological images, failing to effectively repair missing joints (e.g., the missing arm of the Sanxingdui bronze statue); second, they ignore spatiotemporal dynamics (e.g., the time decay of cultural transmission and geographic barriers), rendering symbolic association analysis detached from historical context; and third, semantic understanding relies on subjective interpretation, lacking cross-modal validation of text and image data. While previous studies have attempted to incorporate traditional machine learning algorithms (e.g., support vector machine (SVM) classification), these algorithms are poorly adapted to the specificities of archaeological contexts (small sample sizes, multimodal and heterogeneous data) and struggle to quantify the temporal role of symbols in ritual action chains. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for multi-dimensional analysis of body language in image documents to improve the above-mentioned problems. To achieve the above-mentioned purpose, the technical solutions adopted by the present invention are as follows:

[0005] In a first aspect, the present application provides a multi-dimensional analysis method of body language in image documents, comprising:

[0006] The coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period were detected. Missing joints were repaired using a mirror symmetry completion algorithm. The degree of hand opening and closing, trunk tilt angle, and gaze direction vector were calculated. The archaeological stratigraphic age labels and site geographic coordinates were integrated to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 joints in the hands and 4 joints in the trunk. The mirror symmetry completion algorithm involves flipping the coordinates of the healthy side joints along the spinal axis of symmetry.

[0007] A spatiotemporal weighted K-means cluster analysis was performed on the standardized spatiotemporal feature dataset of body language to classify body language features into categories, establish a morphological classification library, and use the Markov chain model to derive the motion transition probability, thereby constructing a dynamic sequence model.

[0008] The morphological classification database was analyzed in association with the acquired archaeological symbol data. The co-occurrence relationship was constrained based on the state transition probability of the dynamic sequence model. The chi-square test was used to screen significant association pairs. The semantic similarity was calculated by combining the literature semantic embedding model, and a body language-symbol semantic mapping table was compiled.

[0009] The entropy method was used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table, a communication network model including the cultural resistance coefficient was established, and the PageRank algorithm was applied to identify core communication nodes and draw a heat map of symbol communication.

[0010] By integrating the morphological classification library, body language-symbol semantic mapping table and propagation heat map, extracting the principal component features through the singular value decomposition method, developing an interactive three-dimensional visual digital resource library, and completing the spatiotemporal evolution analysis of literature body language.

[0011] Preferably, the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period are detected, missing joints are repaired using a mirror symmetry completion algorithm, and the degree of hand gestures, torso tilt angle, and gaze direction vector are calculated. The archaeological stratigraphic age labels and site geographic coordinates are integrated to form a standardized body language spatiotemporal feature dataset, which includes:

[0012] Detect the coordinates of 25 joint points through the improved OpenPose network;

[0013] Based on the joint point coordinate set, the missing joint points on one side are supplemented along the symmetry axis of the spine to obtain a complete joint point coordinate set;

[0014] Based on the complete set of joint coordinates, the ratio of the Euclidean distance of the palm joints to the maximum physiological distance is calculated to obtain a normalized degree of openness and closeness, which is recorded as the first posture parameter. The inclination angle of the trunk relative to the vertical direction is calculated by the vector angle of the line connecting the shoulder and hip joints, which is recorded as the second posture parameter. Based on the coordinate difference between the key points of the head and eyes, a unitized direction vector is generated, which is recorded as the third posture parameter. The first, second, and third posture parameter values ​​are summarized as a posture parameter set.

[0015] The obtained archaeological stratigraphic age labels and site geographic coordinates are fused to output spatiotemporal labels; the body language parameter set and spatiotemporal labels are spliced ​​to obtain the original feature vector, which is then standardized and optimized to obtain a standardized body language spatiotemporal feature dataset.

[0016] Preferably, the standardized body language spatiotemporal feature dataset is subjected to spatiotemporal weighted K-means cluster analysis to classify body language features, establish a morphological classification library, and use a Markov chain model to derive motion transition probabilities, thereby constructing a dynamic sequence model, which includes:

[0017] Based on the standardized spatiotemporal feature dataset, the samples were weighted using a spatiotemporal weight function, where the temporal weight was calculated using an exponential decay form, while the spatial weight was calculated inversely proportional to the distance between sites. Taking the weighted feature data as the target, clustering was performed by minimizing the weighted distance and generating a set of body feature categories, where each category is marked with a typical morphological feature range and a confidence sample.

[0018] The central feature vector of each category in the body feature category set and the top 5% samples with the highest confidence are extracted to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of the trunk tilt angle, and the principal component of the gaze direction vector;

[0019] Extract multiple continuous frames from high-resolution image literature, perform three-dimensional scanning, generate virtual continuous action sequences, and then arrange them into scene image sequences. Based on the structured classification library, count the transition frequencies between adjacent posture categories, and calculate the transition probability through Laplace smoothing. Finally, construct a dynamic sequence model M = (P, S), where is the state transition probability matrix, S = {S1,…,S10} is the body category state set.

[0020] Preferably, the morphological classification library is subjected to association analysis with the acquired archaeological symbol data, the state transition probability constraint co-occurrence relationship based on the dynamic sequence model is used, the chi-square test is used to screen significant association pairs, the semantic similarity is calculated in combination with the document semantic embedding model, and a body language-symbol semantic mapping table is compiled, which includes:

[0021] Based on the morphological classification library and archaeological symbol data set, the dynamic sequence model is used to perform statistical processing on the co-occurrence frequency under the constraints of adjacent state transitions. The co-occurrence frequency is weighted and calculated in combination with the state transition probability to obtain the co-occurrence frequency matrix with dynamic constraints.

[0022] Based on the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, the improved chi-square test method is used to modify the expected frequency dynamically by introducing the transition probability, and the samples that meet the significance threshold x are screened. 2 >3.841 and the transition probability constraint p ij >0.2 association pair set, and then output the association pair list;

[0023] Based on the association pair list, a pre-trained language model is used to vectorize and embed the body language descriptions and symbolic semantics in historical documents. The basic semantic association is calculated through cosine similarity, and the state transition probability is introduced to weightedly enhance the temporal contribution of symbols in the action chain. The fused body language-symbol semantic mapping table is generated, where the body language-symbol semantic mapping table includes body language categories, associated symbols, chi-square values ​​and enhanced similarity.

[0024] Preferably, the entropy method is used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table, a communication network model including a cultural resistance coefficient is established, and the PageRank algorithm is applied to identify core communication nodes and draw a heat map of symbol communication, which includes:

[0025] Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol in different sites and eras, outputting a symbol spatiotemporal entropy table and quantifying the spatiotemporal stability of each symbol's propagation.

[0026] Combining the symbolic spatiotemporal entropy table with the collected site geographical environment data, the cultural resistance coefficient is calculated and the edge weights of the communication network are constructed to obtain a symbolic communication network with cultural resistance. The edge weight matrix reflects the difficulty of symbolic communication across sites.

[0027] The edge weight matrix of the communication network is used to calculate the communication influence of each site node through the improved PageRank algorithm;

[0028] Gaussian kernel density estimation was used for spatial interpolation based on the communication influence, PageRank value list and the UTM coordinates of the site;

[0029] Output symbol propagation heat map, and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color scale.

[0030] In a second aspect, the present application also provides a multi-dimensional analysis system for body language in image documents, including:

[0031] Integration module: This module detects the coordinates of key human joints in high-resolution image documents unearthed from the early Bashu period, repairs missing joints using a mirror symmetry completion algorithm, calculates gesture opening and closing, trunk tilt angle, and gaze direction vector, and integrates archaeological stratigraphic age labels and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. Key human joints include 21 hand joints and 4 trunk joints. The mirror symmetry completion algorithm involves flipping the coordinates of the healthy-side joints along the spinal axis of symmetry.

[0032] Analysis and Construction Module: This module is used to perform spatiotemporal weighted K-means clustering analysis on the standardized spatiotemporal feature dataset of body language, classify body language features into categories, establish a morphological classification library, and use the Markov chain model to derive motion transition probabilities, thereby constructing a dynamic sequence model.

[0033] Association compilation module: used to conduct association analysis between the morphological classification library and the acquired archaeological symbol data, use the chi-square test to screen significant association pairs, calculate semantic similarity in combination with the document semantic embedding model, and compile a body language-symbol semantic mapping table;

[0034] Evaluation and mapping module: This module is used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establish a communication network model that includes a cultural resistance coefficient, apply the PageRank algorithm to identify core communication nodes, and draw a heat map of symbol communication;

[0035] Extraction and Analysis Module: This module is used to integrate the morphological classification library, the body language-symbol semantic mapping table, and the propagation heat map, extract the principal component features through the singular value decomposition method, develop an interactive three-dimensional visual digital resource library, and complete the spatiotemporal evolution analysis of literature body language.

[0036] In a third aspect, the present application also provides a device for multi-dimensional analysis of body language in image documents, comprising:

[0037] memory for storing computer programs;

[0038] A processor is used to implement the steps of the multi-dimensional analysis method of image document body language when executing the computer program.

[0039] In a fourth aspect, the present application further provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned method for multi-dimensional analysis of body language based on image documents.

[0040] The beneficial effects of the present invention are:

[0041] This paper designs a mirror-symmetric completion algorithm and a spatiotemporal weighted OpenPose model to solve the problem of high-precision restoration of incomplete joints in archaeological images. It also combines UTM projections with stratigraphic data to construct standardized spatiotemporal features. Secondly, it develops a dynamic sequence Markov chain model to quantify the probability of body movement transitions and integrates cultural resistance coefficients to construct a symbolic propagation network, enabling spatiotemporal context-aware association analysis. Finally, it integrates BERT semantic embedding and kernel density estimation techniques to generate a body language-symbol semantic mapping table and propagation heat map, forming a comprehensive analytical framework from image parsing to dynamic modeling to semantic verification. This method is the first to systematically integrate deep learning, complex networks, and digital humanities methods into archaeological image interpretation, providing a quantifiable and verifiable technical path for reconstructing the cultural logic of pre-literate civilizations.

[0042] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0044] Figure 1 Schematic diagram of the process of the multi-dimensional analysis method of body language in image documents according to an embodiment of the present invention;

[0045] Figure 2 Schematic diagram of the structure of the multi-dimensional analysis system of body language in image documents according to an embodiment of the present invention;

[0046] Figure 3 Schematic diagram of the structure of the multi-dimensional analysis device for body language in image documents according to an embodiment of the present invention.

[0047] In the figure: 701, integration module; 702, analysis and construction module; 703, association compilation module; 704, evaluation and drawing module; 705, extraction and analysis module; 800, image document body language multi-dimensional analysis device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0049] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.

[0050] Example 1:

[0051] This embodiment provides a multi-dimensional analysis method for body language in image documents.

[0052] See also Figure 1 , the figure shows that the method includes step S100, step S200, step S300, step S400 and step S500.

[0053] S100. Detect the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period, repair missing joints through a mirror symmetry completion algorithm, calculate the degree of hand gestures, trunk tilt angle, and gaze direction vector, integrate archaeological stratigraphic age labels and site geographic coordinates, and form a standardized body language spatiotemporal feature dataset. The key human joints include 21 joints in the hands and 4 joints in the trunk. The mirror symmetry completion algorithm includes flipping the coordinates of the healthy side joints along the spinal symmetry axis.

[0054] It can be understood that this step includes S101, S102, S103 and S104, wherein:

[0055] S101. Detect the coordinates of 25 joint points using the improved OpenPose network. The calculation formula is as follows:

[0056]

[0057] Where J is the set of joint point coordinates, (x i ,y i ) is the image coordinate of the joint point, c i is the detection confidence, W is the width of the input image, and H is the height of the input image;

[0058] S102. Based on the joint point coordinate set, the missing joint points on one side are completed along the symmetry axis of the spine to obtain a complete joint point coordinate set. The calculation formula for completing the missing joint points on one side is as follows:

[0059]

[0060] Where p s is the midpoint of the spine, p missing is the coordinate of the missing joint point to be completed, p mirror is the coordinate of the joint point corresponding to the mirror side, p ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh for the right hip;

[0061] S103. Calculate the ratio of the Euclidean distance of the palm joints to the maximum physiological distance based on the complete set of joint coordinates to obtain a normalized degree of openness and closeness, which is recorded as a first posture parameter. Calculate the inclination angle of the trunk relative to the vertical direction based on the vector angle of the line connecting the shoulder and hip joints, which is recorded as a second posture parameter. Generate a normalized direction vector based on the coordinate difference between the key points of the head and eyes, which is recorded as a third posture parameter. The first, second, and third posture parameter values ​​are aggregated into a posture parameter set.

[0062] It should be noted that the process of calculating posture parameters is as follows: calculate the Euclidean distance between the palm joints (such as the index finger tip and the palm), and divide it by the maximum physiological distance (standardized to 200 pixels based on the size of the bronze statue) to obtain the normalized opening and closing value; calculate the inclination angle of the torso relative to the vertical direction through the vector angle between the shoulder and hip joints; generate a unitized direction vector based on the coordinate difference between the key points of the head and eyes to represent the direction of the person's gaze; then output the posture parameter set including the gesture opening and closing degree, the torso inclination angle, and the gaze direction vector. The calculation formula for the gesture opening and closing degree is as follows:

[0063]

[0064] Where p tip is the coordinate of the index finger tip, p palm is the palm coordinate, S is the degree of gesture opening and closing;

[0065] The torso inclination angle is: And take the average value of the left and right sides, its p h =(x h ,y h ) is the hip joint coordinate; its gaze direction vector is Therefore, in view of the incomplete characteristics of archaeological images, geometric constraint completion based on the spine symmetry axis is used to solve the failure problem of conventional posture estimation models under missing data.

[0066] S104. Fuse the obtained archaeological stratigraphic age labels and the site geographic coordinates to output spatiotemporal labels; concatenate the body language parameter set and the spatiotemporal labels to obtain the original feature vector, and standardize and optimize the original feature vector to obtain a standardized body language spatiotemporal feature dataset.

[0067] Archaeological stratigraphic age tags can be obtained in several ways, including symbolic information extracted from inscriptions, ornamentation, and patterns on artifacts from the same batch of Bashu excavations (such as Sanxingdui bronzes, Jinsha gold artifacts, and Han Dynasty pictorial bricks). These include, but are not limited to, carved symbols on artifact surfaces (such as "feathered figure patterns" and "divine tree patterns"); graphic symbols in stele inscriptions; and decorative patterns on pictorial bricks. These methods include extracting symbolic features from archaeological excavation reports and catalogs, using image recognition technology to extract symbolic features from artifact photographs, or combining archaeologists' professional annotation and classification. During this processing, the extracted symbols can be standardized and coded to establish a symbol classification system, recording contextual information such as the excavation site and age of each symbol. These symbolic data represent archaeological materials excavated contemporaneously and from the same source as the body language imagery, possessing clear temporal and spatial correlations rather than pre-defined theoretical symbols. Site geographic coordinates are derived by converting GPS coordinates to Universal Transverse Mercator projection coordinates, which ensures spatial consistency between geographic coordinates and body language features.

[0068] The original feature vectors are standardized and optimized to obtain the standardized body language spatiotemporal feature dataset, which is calculated column by column based on the feature values ​​of all samples in the training set:

[0069]

[0070] Where N is the total number of samples and i is the feature dimension.

[0071] In summary, the standardized body language spatiotemporal feature dataset finally generated through this process can simultaneously reflect the morphological characteristics of the character's posture and the archaeological spatiotemporal background, laying a data foundation for subsequent multi-dimensional analysis.

[0072] S200. Perform spatiotemporal weighted K-means clustering analysis on the standardized spatiotemporal feature dataset of body language, divide the body feature categories, establish a morphological classification library, and use the Markov chain model to derive the action transition probability, thereby constructing a dynamic sequence model.

[0073] It can be understood that this step includes S201, S202 and S203, wherein:

[0074] S201. Based on the standardized spatiotemporal feature dataset, weight the samples using a spatiotemporal weight function, where the temporal weight is calculated using an exponential decay form, and the spatial weight is calculated inversely proportional to the distance between sites. Using the weighted feature data as the target, clustering is performed by minimizing the weighted distance and generating a set of body feature categories, where each category is labeled with a typical morphological feature range and a confidence sample.

[0075] It should be noted that the time weight is calculated using the exponential decay formula: wt = e, where λ = 0.01 is the time decay coefficient, Tmax The latest age benchmark value.

[0076] S202, extracting the central feature vector of each category in the body feature category set and the top 5% samples with the highest confidence, and establishing a structured classification library, where the category features are the hand gesture opening and closing range, the trunk tilt angle threshold, and the principal component of the gaze direction vector;

[0077] It should be noted that continuous action images are extracted from a series of cultural relics of the same sacrificial scene (such as Han Dynasty pictorial bricks and bronze decorations). For example, a Han Dynasty pictorial brick depicts the continuous actions of a sacrificial dance (kneeling → offering sacrifice → dancing); after three-dimensional scanning of incomplete cultural relics (such as Sanxingdui bronzes), a virtual continuous action sequence is generated. Then, according to the narrative order of the decoration on the surface of the artifact or the scene restoration of archaeologists, the image time sequence is arranged, and the long sequence is divided into short action units (3-5 frames per unit) according to the changes in posture, and then the scene image sequence is output.

[0078] S203. Extract multiple continuous frames from high-resolution image documents, perform three-dimensional scanning, generate a virtual continuous action sequence, and then arrange it into a scene image sequence; based on the structured classification library, count the transition frequencies between adjacent posture categories, and calculate the transition probability through Laplace smoothing, and finally construct a dynamic sequence model M = (P, S), where is the state transition probability matrix, S = {S1,…,S10} is the body category state set.

[0079] It's understandable that in this step, after 3D scanning an incomplete artifact (such as a bronze figure with only one side in motion), the missing parts can be mirrored using tools like Blender or Maya to create a complete 3D model. Alternatively, keyframe interpolation techniques can be used to generate a virtual motion sequence (e.g., a smooth transition from "standing" to "kneeling"). The images are then arranged in temporal order according to archaeological narrative logic (e.g., the order of reading artifact decorations and the stratigraphic relationships) and divided into units based on the motion changes. For example, the "holding an object high" motion of the bronze standing figure at K2②:35 in Sanxingdui is decomposed into three sub-movements: "raising hand → extending arm → stabilizing posture." Each posture category in the morphological classification library (e.g., C1: "both arms raised") directly corresponds to a state S1 in the Markov chain. The feature range provided by the classification library (e.g., S∈[0.7,0.9]) is used as a threshold for state determination. In this step, the constructed Markov chain model not only reflects the observed motion transitions but also, through spatiotemporal weighting and smoothing strategies, restores the underlying logic of cultural communication.

[0080] S300. Perform correlation analysis on the morphological classification library and the acquired archaeological symbol data, constrain the co-occurrence relationship based on the state transition probability of the dynamic sequence model, use the chi-square test to screen significant association pairs, calculate the semantic similarity in combination with the document semantic embedding model, and compile a body language-symbol semantic mapping table.

[0081] It can be understood that step S300 includes S301, S302 and S303, wherein:

[0082] S301, based on the morphological classification library and the archaeological symbol data set, performing co-occurrence frequency statistical processing under the adjacent state transition constraint through a dynamic sequence model, performing weighted calculation on the co-occurrence frequency in combination with the state transition probability, and obtaining a co-occurrence frequency matrix with dynamic constraints;

[0083] S302, using the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, adopting the improved chi-square test method, dynamically modifying the expected frequency by introducing the transition probability, and screening the frequencies that meet the significance threshold χ 2 >3.841 and the transition probability constraint p ij >0.2 association pair set, and then output the association pair list;

[0084] It should be noted that only the posture-symbol co-occurrence of adjacent states in the same action chain is counted. If the symbol s j Appears in state Sk, then only when S k When associated with the predecessor state Si or the successor state Sl with the transition probability pi→k>0.1, the body type Ci / Cl and s are recorded. j The weighted frequency calculation formula is as follows:

[0085]

[0086] Where, Indicates symbol s j In state S k 1 if it appears in the i→k From state Si to S k The transition probability is output as the co-occurrence frequency matrix of dynamic constraints.

[0087] It is understandable that the traditional chi-square test only calculates the expected value based on frequency, but the association between symbols and postures in archaeological scenes must conform to the temporal logic of action. For example, if the symbol s j Only appears in low probability transfer paths (such as p i→k<0.1), its co-occurrence may be accidental mixing (such as the mixing of debris), and needs to be downgraded. In the above steps, the observed value Oij of the traditional chi-square formula needs to be weighted by the transfer probability. If wij→0, even if Oij is high, the χ2 value will be suppressed. Among them, the screening conditions: significance threshold: χ2>3.841 (α=0.05, degree of freedom=1); transition probability lower limit: pi→j>0.2, excluding noise associations of low-probability paths. Therefore, the state transition probability is introduced as the frequency weight to make the statistical test results consistent with the logic of cultural behavior. For example, the "sacred tree pattern" is strengthened in the early high-probability transfer path, while the same symbol in the later low-probability path is weakened. That is, the Bashu model Compared with the Shang and Zhou models, if a symbol (such as "dragon pattern") is only significant in the high-probability path of the Shang and Zhou dynasties, it may reflect the evolution of its function in cultural communication.

[0088] S303. Based on the association pair list, a pre-trained language model is used to vectorize and embed the body language descriptions and symbolic semantics in historical documents. The basic semantic association is calculated through cosine similarity, and the state transition probability is introduced to weightedly enhance the temporal contribution of symbols in the action chain. The fused body language-symbol semantic mapping table is generated, where the body language-symbol semantic mapping table includes body language categories, associated symbols, chi-square values ​​and enhanced similarity.

[0089] It should be noted that the semantic embedding and basic similarity calculation are to extract the body shape C from the literature. i and symbol s j For a paragraph description (such as "raise your hands to pay respect to the sacred tree"), the BERT model is used to generate paragraph-level semantic vectors. The dimension is 768. If a symbol appears frequently in the subsequent action chain, its semantic association needs to be enhanced. Define the temporal contribution factor and enhance the similarity. After that, perform a screening threshold, retain the association pairs with sim′>0.75, and sort them in descending order according to the chi-square value. Therefore, the text semantics (document description) and action timing (transition probability) are combined to break through the limitations of single modal analysis. For example, the meaning of a symbol that is not clearly recorded in the literature (such as "feathered man pattern") can be determined by its frequent appearance in the action chain (p i→k Gao) is inferred to be a "guiding spirit".

[0090] In summary, through the deep coupling of dynamic sequence models and multimodal data analysis, accidental associations that do not conform to ritual logic are eliminated, text descriptions and action timing are integrated, the credibility of symbolic meaning analysis is improved, and a quantitative basis is provided for archaeological controversial issues such as symbolic functional differentiation and ritual stage division.

[0091] S400. Evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table through the entropy method, establish a communication network model including the cultural resistance coefficient, apply the PageRank algorithm to identify core communication nodes, and draw a heat map of symbol communication.

[0092] It can be understood that step S400 includes S401, S402, S403, S404 and S405, wherein:

[0093] Use the PageRank algorithm to identify core communication nodes and draw a heat map of symbol propagation, including:

[0094] S401. Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol at different sites and ages, output a symbol spatiotemporal entropy table, and quantify the spatiotemporal stability of each symbol's propagation;

[0095] S402. Calculate the cultural resistance coefficient by combining the symbolic spatiotemporal entropy table and the collected site geographical environment data, construct the edge weights of the communication network, and obtain a symbolic communication network with cultural resistance, where the edge weight matrix reflects the difficulty of symbolic communication across sites.

[0096] S403, using the edge weight matrix of the communication network, calculate the communication influence of each heritage node through the improved PageRank algorithm;

[0097] S404. Based on the communication influence, PageRank value list and the UTM coordinates of the site, Gaussian kernel density estimation is used for spatial interpolation. The calculation formula is as follows:

[0098]

[0099] Where f(x, y) is the estimated value of the propagation strength at the location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) is the communication influence of the site node, K is the Gaussian kernel function, is the Euclidean distance between the target point (x, y) and the site vi;

[0100] S405: Output a symbol propagation heat map, and visualize the spatiotemporal diffusion pattern of the symbol from the core area to the edge area through a red-blue gradient color scale.

[0101] It should be noted that the Silverman criterion is used for adaptive calculation to select the bandwidth. For example, if the standard deviation of the site distribution is 50 km, then h is approximately 45 km. The study area is divided into a 1000×1000 grid, and f(x, y) is calculated for each grid point (x, y). In this embodiment, the high-density area, such as the Chengdu Plain, is locally encrypted, and then the intensity is normalized and the color scale is designed, such as the red-yellow-blue gradient color scale: red (RGB 255,0,0) f norm >0.8 (core propagation area); Yellow (RGB 255,255,0): 0.3 <f norm ≤0.8; blue (RGB 0,0,255): f norm ≤0.3 (edge ​​area), and overlay the site locations (black dots), mountains (dark gray filling), and rivers (blue lines) on the map.

[0102] S500, comprehensive morphological classification library, body language-symbol semantic mapping table and propagation heat map, extract principal component features through singular value decomposition method, develop interactive three-dimensional visual digital resource library, and complete the spatiotemporal evolution analysis of literature body language.

[0103] It can be understood that in this step, based on the morphological classification library, body language-symbol semantic mapping table and communication heat map data, the singular value decomposition (SVD) method is used to extract the principal component features of the joint feature matrix (such as the first principal component reflects the "sacrifice-power" dimension, and the weight covers the symbol association chi-square value and the communication PageRank value; the second principal component represents the "action-environment" relationship, with a high load on the trunk inclination angle and the geographical vertical coordinate; the third principal component maps the "symbol-semantic" association, and the dominant feature is the enhanced semantic similarity and gesture opening and closing degree), constructs a three-dimensional principal component space (the X / Y / Z axes correspond to the three major principal components respectively, and the time dimension maps the changes in the era through color gradients), and develops an interactive visualization engine based on WebGL and Three.js to implement it. The virtual reconstruction of the current sacrificial action chain (such as the dynamic sequence of "raising arms → kneeling on one knee → turning around"), the spatiotemporal superposition of symbol transmission heat maps (Gaussian kernel density estimation generates red-blue gradient color levels), and multi-dimensional data linkage (clicking on a heat node can highlight the related literature paragraphs); further use of B-spline curves to fit the cultural evolution trajectory, and detect key mutation periods through curvature extremes (such as the decline period of Sanxingdui culture BC1000±50), ultimately forming a digital resource library that supports VR immersive observation, cross-period comparison (such as the difference in action chains between the Shang and Zhou dynasties and the Han Dynasty), and sensitivity simulation (modifying geographical resistance parameters to predict transmission paths), providing an intelligent platform with both quantitative analysis and scene restoration capabilities for the study of non-written civilizations, breaking through the limitations of traditional archaeology that relies on static charts and empirical speculation.

[0104] Example 2:

[0105] like Figure 2 As shown, this embodiment provides a multi-dimensional analysis system for body language in image documents, see Figure 2 The system comprises:

[0106] Integration Module 701: Detects the coordinates of key human joints in high-resolution image documents unearthed from the early Bashu period, repairs missing joints using a mirror symmetry completion algorithm, calculates gesture opening and closing, trunk tilt angle, and gaze direction vector, and integrates archaeological stratigraphic age labels and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. Key human joints include 21 hand joints and 4 trunk joints. The mirror symmetry completion algorithm involves flipping the coordinates of the healthy side joints along the spinal axis of symmetry.

[0107] Analysis and construction module 702: for performing spatiotemporal weighted K-means cluster analysis on the standardized spatiotemporal feature dataset of body language, classifying body language features into categories, establishing a morphological classification library, and using a Markov chain model to derive motion transition probabilities, thereby constructing a dynamic sequence model;

[0108] Association compilation module 703: for performing association analysis between the morphological classification library and the acquired archaeological symbol data, screening significant association pairs using a chi-square test, calculating semantic similarity in combination with a document semantic embedding model, and compiling a body language-symbol semantic mapping table;

[0109] Evaluation and drawing module 704: used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establish a communication network model including the cultural resistance coefficient, apply the PageRank algorithm to identify core communication nodes, and draw a symbol communication heat map;

[0110] Extraction and Analysis Module 705: used to integrate the morphological classification library, body language-symbol semantic mapping table and propagation heat map, extract the principal component features through the singular value decomposition method, develop an interactive three-dimensional visual digital resource library, and complete the spatiotemporal evolution analysis of literature body language.

[0111] Specifically, the integration module 701 includes:

[0112] Detection unit: used to detect the coordinates of 25 joint points through the improved OpenPose network. The calculation formula is as follows:

[0113]

[0114] Where J is the set of joint point coordinates, (x i ,y i ) is the image coordinate of the joint point, c i is the detection confidence, W is the width of the input image, and H is the height of the input image;

[0115] Completion unit: It is used to complete the missing joint points on one side along the symmetry axis of the spine based on the joint point coordinate set to obtain a complete joint point coordinate set. The calculation formula for completing the missing joint points on one side is as follows:

[0116]

[0117] Where p s is the midpoint of the spine, p missing is the coordinate of the missing joint point to be completed, p mirror is the coordinate of the joint point corresponding to the mirror side, p ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh for the right hip;

[0118] The first calculation unit is configured to calculate the ratio of the Euclidean distance of the palm joint points to the maximum physiological distance based on the complete set of joint point coordinates, thereby obtaining a normalized degree of openness and closeness, which is recorded as a first posture parameter value; calculate the inclination angle of the trunk relative to the vertical direction based on the vector angle of the line connecting the shoulder and hip joint points, which is recorded as a second posture parameter value; generate a unitized direction vector based on the coordinate difference between the key points of the head and eyes, which is recorded as a third posture parameter value; and summarize the first posture parameter value, the second posture parameter value, and the third posture parameter value into a posture parameter set;

[0119] Optimization unit: used to fuse the obtained archaeological stratigraphic age labels and the site geographic coordinates to output spatiotemporal labels; splice the body language parameter set and spatiotemporal labels to obtain the original feature vector, and standardize and optimize the original feature vector to obtain a standardized body language spatiotemporal feature dataset.

[0120] Specifically, the analysis construction module 702 includes:

[0121] The second calculation unit is used to weight the samples based on the standardized spatiotemporal feature dataset using a spatiotemporal weight function, where the temporal weight is calculated in an exponential decay form, and the spatial weight is calculated inversely proportional to the distance between sites. Taking the weighted feature data as the target, clustering is performed by minimizing the weighted distance and generating a set of body feature categories, where each category is marked with a typical morphological feature range and a confidence sample;

[0122] Establishment unit: used to extract the central feature vector of each category in the body feature category set and the top 5% samples with the highest confidence, and establish a structured classification library, where the category features are the gesture opening and closing range, the trunk tilt angle threshold, and the principal component of the gaze direction vector;

[0123] Extraction unit: It is used to extract multiple frames of continuous images from high-resolution image documents, perform three-dimensional scanning, generate virtual continuous action sequences, and then arrange them into scene image sequences; based on the structured classification library, it counts the transition frequencies between adjacent posture categories and calculates the transition probability through Laplace smoothing, and finally constructs a dynamic sequence model M = (P, S), where is the state transition probability matrix, S = {S1,…,S10} is the body category state set.

[0124] Specifically, the association compilation module 703 includes:

[0125] Processing unit: used to perform co-occurrence frequency statistics under the constraints of adjacent state transitions based on the morphological classification library and the archaeological symbol data set through a dynamic sequence model, and perform weighted calculation of the co-occurrence frequency in combination with the state transition probability to obtain a co-occurrence frequency matrix with dynamic constraints;

[0126] Correction unit: It is used to use the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, adopt the improved chi-square test method, and dynamically correct the expected frequency by introducing the transition probability to screen the frequencies that meet the significance threshold x. 2 >3.841 and the transition probability constraint p ij >0.2 association pair set, and then output the association pair list;

[0127] Generation unit: Used to vectorize and embed body language descriptions and symbolic semantics in historical documents based on the association pair list using a pre-trained language model, calculate the basic semantic association through cosine similarity, and introduce state transition probability to weightedly enhance the temporal contribution of symbols in the action chain, thereby generating a fused body language-symbol semantic mapping table, where the body language-symbol semantic mapping table includes body language categories, associated symbols, chi-square values, and enhanced similarity.

[0128] Specifically, the evaluation and drawing module 704 includes:

[0129] Quantification unit: Based on the symbol distribution data in the body language-symbol semantic mapping table, it uses the entropy method to calculate the frequency distribution dispersion of each symbol in different sites and eras, outputs a symbol spatiotemporal entropy value table, and quantifies the spatiotemporal stability of each symbol's propagation;

[0130] Construction unit: used to combine the symbolic spatiotemporal entropy table and the collected site geographical environment data to calculate the cultural resistance coefficient, construct the edge weights of the communication network, and obtain the symbolic communication network with cultural resistance, where the edge weight matrix reflects the difficulty of symbol transmission across sites;

[0131] The third calculation unit is used to calculate the communication influence of each site node using the improved PageRank algorithm using the edge weight matrix of the communication network;

[0132] The fourth calculation unit is used to perform spatial interpolation using Gaussian kernel density estimation based on the communication influence, PageRank value list, and the UTM coordinates of the site. The calculation formula is as follows:

[0133]

[0134] Where f(x, y) is the estimated value of the propagation strength at the location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) is the communication influence of the site node, K is the Gaussian kernel function, is the Euclidean distance between the target point (x, y) and the site vi;

[0135] Diffusion unit: It is used to output the symbol propagation heat map and visualize the spatiotemporal diffusion pattern of the symbol from the core area to the edge area through red-blue gradient color scale.

[0136] It should be noted that, regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0137] Example 3:

[0138] Corresponding to the above method embodiment, this embodiment also provides an image document body language multi-dimensional analysis device. The image document body language multi-dimensional analysis device described below and the image document body language multi-dimensional analysis method described above can refer to each other.

[0139] Figure 3 FIG. 8 is a block diagram of a multi-dimensional analysis device 800 for body language in image documents according to an exemplary embodiment. Figure 3 As shown, the apparatus 800 for multi-dimensional analysis of body language in image documents includes: a processor 801 and a memory 802 . The apparatus 800 for multi-dimensional analysis of body language in image documents also includes one or more of a multimedia component 803 , an I / O interface 804 , and a communication component 805 .

[0140] The processor 801 is used to control the overall operation of the multi-dimensional image and document body language analysis device 800 to complete all or part of the steps in the above-mentioned multi-dimensional image and document body language analysis method. The memory 802 is used to store various types of data to support the operation of the multi-dimensional image and document body language analysis device 800. This data may include, for example, instructions for any application or method operating on the multi-dimensional image and document body language analysis device 800, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the above-mentioned other interface modules can be a keyboard, a mouse or buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for wired or wireless communication between the image document body language multidimensional analysis device 800 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 805 may include: a Wi-Fi module, a Bluetooth module or an NFC module.

[0141] In an exemplary embodiment, the image document body language multi-dimensional analysis device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned image document body language multi-dimensional analysis method.

[0142] In another exemplary embodiment, a computer-readable storage medium containing program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described method for multi-dimensional analysis of body language in image documents. For example, the computer-readable storage medium may be the aforementioned memory 802 containing the program instructions. The program instructions may be executed by the processor 801 of the device 800 for multi-dimensional analysis of body language in image documents to implement the above-described method for multi-dimensional analysis of body language in image documents.

[0143] Example 4:

[0144] Corresponding to the above method embodiment, this embodiment further provides a readable storage medium. The readable storage medium described below and the multi-dimensional analysis method of body language in image documents described above can refer to each other.

[0145] The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the multi-dimensional analysis method of image document body language in the above method embodiment.

[0146] The readable storage medium may specifically be any readable storage medium that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0147] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0148] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multi-dimensional analysis method for body language in image documents, characterized by: include: The coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period were detected. Missing joints were repaired using a mirror symmetry completion algorithm. The degree of hand opening and closing, trunk tilt angle, and gaze direction vector were calculated. The archaeological stratigraphic age labels and site geographic coordinates were integrated to form a standardized body language spatiotemporal feature dataset. The key human joints include 21 joints in the hands and 4 joints in the trunk. The mirror symmetry completion algorithm involves flipping the coordinates of the healthy side joints along the spinal axis of symmetry. A spatiotemporal weighted K-means cluster analysis was performed on the standardized spatiotemporal feature dataset of body language to classify body language features into categories, establish a morphological classification library, and use the Markov chain model to derive the motion transition probability, thereby constructing a dynamic sequence model. The morphological classification database was analyzed in association with the acquired archaeological symbol data. The co-occurrence relationship was constrained based on the state transition probability of the dynamic sequence model. The chi-square test was used to screen significant association pairs. The semantic similarity was calculated by combining the literature semantic embedding model, and a body language-symbol semantic mapping table was compiled. The entropy method was used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table, a communication network model including the cultural resistance coefficient was established, and the PageRank algorithm was applied to identify core communication nodes and draw a heat map of symbol communication. By integrating the morphological classification library, body language-symbol semantic mapping table and propagation heat map, extracting the principal component features through the singular value decomposition method, developing an interactive three-dimensional visual digital resource library, and completing the spatiotemporal evolution analysis of literature body language.

2. The multi-dimensional analysis method of body language in image documents according to claim 1, characterized in that: The method detects the coordinates of key human joints in high-resolution image documents unearthed in the early Bashu period, repairs missing joints using a mirror symmetry completion algorithm, calculates the degree of hand gestures, torso tilt angle, and gaze direction vector, and integrates archaeological stratigraphic age labels and site geographic coordinates to form a standardized body language spatiotemporal feature dataset, which includes: The coordinates of 25 joint points are detected through the improved OpenPose network, and the calculation formula is as follows: Where J is the set of joint point coordinates, (x i ,y i ) is the image coordinate of the joint point, c i is the detection confidence, W is the width of the input image, and H is the height of the input image; Based on the joint point coordinate set, the missing joint points on one side are supplemented along the symmetry axis of the spine to obtain a complete joint point coordinate set. The calculation formula for supplementing the missing joint points on one side is as follows: Where p s is the midpoint of the spine, p missing is the coordinate of the missing joint point to be completed, p mirror is the coordinate of the joint point corresponding to the mirror side, p ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh for the right hip; Based on the complete set of joint coordinates, the ratio of the Euclidean distance of the palm joints to the maximum physiological distance is calculated to obtain a normalized degree of openness and closeness, which is recorded as the first posture parameter. The inclination angle of the trunk relative to the vertical direction is calculated by the vector angle of the line connecting the shoulder and hip joints, which is recorded as the second posture parameter. Based on the coordinate difference between the key points of the head and eyes, a unitized direction vector is generated, which is recorded as the third posture parameter. The first, second, and third posture parameter values ​​are summarized as a posture parameter set. The obtained archaeological stratigraphic age labels and site geographic coordinates are fused to output spatiotemporal labels; the body language parameter set and spatiotemporal labels are spliced ​​to obtain the original feature vector, which is then standardized and optimized to obtain a standardized body language spatiotemporal feature dataset.

3. The multi-dimensional analysis method of body language in image documents according to claim 1, characterized in that: The method performs spatiotemporal weighted K-means clustering analysis on the standardized spatiotemporal feature dataset, divides the body language features into categories, establishes a morphological classification library, and uses a Markov chain model to derive action transition probabilities, thereby constructing a dynamic sequence model, which includes: Based on the standardized spatiotemporal feature dataset, the samples were weighted using a spatiotemporal weight function, where the temporal weight was calculated using an exponential decay form, while the spatial weight was calculated inversely proportional to the distance between sites. Taking the weighted feature data as the target, clustering was performed by minimizing the weighted distance and generating a set of body feature categories, where each category is marked with a typical morphological feature range and a confidence sample. The central feature vector of each category in the body feature category set and the top 5% samples with the highest confidence are extracted to establish a structured classification library, where the category features are the range of hand gesture opening and closing, the threshold of the trunk tilt angle, and the principal component of the gaze direction vector; Extract multiple continuous frames from high-resolution image literature, perform three-dimensional scanning, generate virtual continuous action sequences, and then arrange them into scene image sequences. Based on the structured classification library, count the transition frequencies between adjacent posture categories, and calculate the transition probability through Laplace smoothing. Finally, construct a dynamic sequence model M = (P, S), where is the state transition probability matrix, S = {S1,…,S10} is the body category state set.

4. The multi-dimensional analysis method of body language in image documents according to claim 1, characterized in that: The morphological classification library is associated with the acquired archaeological symbol data through analysis. Based on the state transition probability constraint co-occurrence relationship of the dynamic sequence model, the chi-square test is used to screen significant association pairs. The semantic similarity is calculated in combination with the document semantic embedding model to compile a body language-symbol semantic mapping table, which includes: Based on the morphological classification library and archaeological symbol data set, the dynamic sequence model is used to perform statistical processing on the co-occurrence frequency under the constraints of adjacent state transitions. The co-occurrence frequency is weighted and calculated in combination with the state transition probability to obtain the co-occurrence frequency matrix with dynamic constraints. Based on the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, the improved chi-square test method is used to modify the expected frequency dynamically by introducing the transition probability, and the expected frequency is screened to meet the significance threshold χ 2 >3.841 and the transition probability constraint p ij >0.2 association pair set, and then output the association pair list; Based on the association pair list, a pre-trained language model is used to vectorize and embed the body language descriptions and symbolic semantics in historical documents. The basic semantic association is calculated through cosine similarity, and the state transition probability is introduced to weightedly enhance the temporal contribution of symbols in the action chain. The fused body language-symbol semantic mapping table is generated, where the body language-symbol semantic mapping table includes body language categories, associated symbols, chi-square values ​​and enhanced similarity.

5. The multi-dimensional analysis method of body language in image documents according to claim 1, characterized in that: The entropy method is used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table, establish a communication network model including the cultural resistance coefficient, apply the PageRank algorithm to identify core communication nodes, and draw a symbol communication heat map, which includes: Based on the symbol distribution data in the body language-symbol semantic mapping table, the entropy method is used to calculate the frequency distribution dispersion of each symbol in different sites and eras, outputting a symbol spatiotemporal entropy table and quantifying the spatiotemporal stability of each symbol's propagation. Combining the symbolic spatiotemporal entropy table with the collected site geographical environment data, the cultural resistance coefficient is calculated and the edge weights of the communication network are constructed to obtain a symbolic communication network with cultural resistance. The edge weight matrix reflects the difficulty of symbolic communication across sites. The edge weight matrix of the communication network is used to calculate the communication influence of each site node through the improved PageRank algorithm; According to the communication influence, PageRank value list and the UTM coordinates of the site, Gaussian kernel density estimation is used for spatial interpolation. The calculation formula is as follows: Where f(x, y) is the estimated value of the propagation strength at the location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) is the communication influence of the site node, K is the Gaussian kernel function, is the Euclidean distance between the target point (x, y) and the site vi; Output symbol propagation heat map, and visualize the spatiotemporal diffusion pattern of symbols from the core area to the edge area through red-blue gradient color scale.

6. A multi-dimensional analysis system for body language in image documents, based on the multi-dimensional analysis method for body language in image documents according to claim 1, characterized in that: include: Integration module: This module detects the coordinates of key human joints in high-resolution image documents unearthed from the early Bashu period, repairs missing joints using a mirror symmetry completion algorithm, calculates gesture opening and closing, trunk tilt angle, and gaze direction vector, and integrates archaeological stratigraphic age labels and site geographic coordinates to form a standardized body language spatiotemporal feature dataset. Key human joints include 21 hand joints and 4 trunk joints. The mirror symmetry completion algorithm involves flipping the coordinates of the healthy-side joints along the spinal axis of symmetry. Analysis and Construction Module: This module is used to perform spatiotemporal weighted K-means clustering analysis on the standardized spatiotemporal feature dataset of body language, classify body language features into categories, establish a morphological classification library, and use the Markov chain model to derive motion transition probabilities, thereby constructing a dynamic sequence model. Association compilation module: used to conduct association analysis between the morphological classification library and the acquired archaeological symbol data, use the chi-square test to screen significant association pairs, calculate semantic similarity in combination with the document semantic embedding model, and compile a body language-symbol semantic mapping table; Evaluation and mapping module: This module is used to evaluate the spatiotemporal stability of symbols in the body language-symbol semantic mapping table using the entropy method, establish a communication network model that includes a cultural resistance coefficient, apply the PageRank algorithm to identify core communication nodes, and draw a heat map of symbol communication; Extraction and Analysis Module: This module is used to integrate the morphological classification library, the body language-symbol semantic mapping table, and the propagation heat map, extract the principal component features through the singular value decomposition method, develop an interactive three-dimensional visual digital resource library, and complete the spatiotemporal evolution analysis of literature body language.

7. The multi-dimensional analysis system of image document body language according to claim 6, characterized in that: The integration module includes: Detection unit: used to detect the coordinates of 25 joint points through the improved OpenPose network. The calculation formula is as follows: Where J is the set of joint point coordinates, (x i ,y i ) is the image coordinate of the joint point, c i is the detection confidence, W is the width of the input image, and H is the height of the input image; Completion unit: It is used to complete the missing joint points on one side along the symmetry axis of the spine based on the joint point coordinate set to obtain a complete joint point coordinate set. The calculation formula for completing the missing joint points on one side is as follows: Where p s is the midpoint of the spine, p missing is the coordinate of the missing joint point to be completed, p mirror is the coordinate of the joint point corresponding to the mirror side, p ls For the left shoulder, p rs For the right shoulder, p lh For the left hip, p rh for the right hip; The first calculation unit is configured to calculate the ratio of the Euclidean distance of the palm joint points to the maximum physiological distance based on the complete set of joint point coordinates, thereby obtaining a normalized degree of openness and closeness, which is recorded as a first posture parameter value; calculate the inclination angle of the trunk relative to the vertical direction based on the vector angle of the line connecting the shoulder and hip joint points, which is recorded as a second posture parameter value; generate a unitized direction vector based on the coordinate difference between the key points of the head and eyes, which is recorded as a third posture parameter value; and summarize the first posture parameter value, the second posture parameter value, and the third posture parameter value into a posture parameter set; Optimization unit: used to fuse the obtained archaeological stratigraphic age labels and the site geographic coordinates to output spatiotemporal labels; splice the body language parameter set and spatiotemporal labels to obtain the original feature vector, and standardize and optimize the original feature vector to obtain a standardized body language spatiotemporal feature dataset.

8. The multi-dimensional analysis system of image document body language according to claim 6, characterized in that: The analysis building blocks include: The second calculation unit is used to weight the samples based on the standardized spatiotemporal feature dataset using a spatiotemporal weight function, where the temporal weight is calculated in an exponential decay form, and the spatial weight is calculated inversely proportional to the distance between sites. Taking the weighted feature data as the target, clustering is performed by minimizing the weighted distance and generating a set of body feature categories, where each category is marked with a typical morphological feature range and a confidence sample; Establishment unit: used to extract the central feature vector of each category in the body feature category set and the top 5% samples with the highest confidence, and establish a structured classification library, where the category features are the gesture opening and closing range, the trunk tilt angle threshold, and the principal component of the gaze direction vector; Extraction unit: It is used to extract multiple frames of continuous images from high-resolution image documents, perform three-dimensional scanning, generate virtual continuous action sequences, and then arrange them into scene image sequences; based on the structured classification library, it counts the transition frequencies between adjacent posture categories and calculates the transition probability through Laplace smoothing, and finally constructs a dynamic sequence model M = (P, S), where is the state transition probability matrix, S = {S1,…,S10} is the body category state set.

9. The multi-dimensional analysis system of body language in image documents according to claim 6, characterized in that: The association compilation module includes: Processing unit: used to perform co-occurrence frequency statistics under the constraints of adjacent state transitions based on the morphological classification library and the archaeological symbol data set through a dynamic sequence model, and perform weighted calculation of the co-occurrence frequency in combination with the state transition probability to obtain a co-occurrence frequency matrix with dynamic constraints; Correction unit: It is used to use the co-occurrence frequency matrix and the transition probability matrix in the dynamic sequence model as input, adopt the improved chi-square test method, and perform dynamic weight correction on the expected frequency by introducing the transition probability to screen the co-occurrence frequency matrix that meets the significance threshold χ 2 >3.841 and the transition probability constraint p ij >0.2 association pair set, and then output the association pair list; Generation unit: Used to vectorize and embed body language descriptions and symbolic semantics in historical documents based on the association pair list using a pre-trained language model, calculate the basic semantic association through cosine similarity, and introduce state transition probability to weightedly enhance the temporal contribution of symbols in the action chain, thereby generating a fused body language-symbol semantic mapping table, where the body language-symbol semantic mapping table includes body language categories, associated symbols, chi-square values, and enhanced similarity.

10. The multi-dimensional analysis system of image document body language according to claim 6, characterized in that: The evaluation drawing module includes: Quantification unit: Based on the symbol distribution data in the body language-symbol semantic mapping table, it uses the entropy method to calculate the frequency distribution dispersion of each symbol in different sites and eras, outputs a symbol spatiotemporal entropy value table, and quantifies the spatiotemporal stability of each symbol's propagation; Construction unit: used to combine the symbolic spatiotemporal entropy table and the collected site geographical environment data to calculate the cultural resistance coefficient, construct the edge weights of the communication network, and obtain the symbolic communication network with cultural resistance, where the edge weight matrix reflects the difficulty of symbol transmission across sites; The third calculation unit is used to calculate the communication influence of each site node using the improved PageRank algorithm using the edge weight matrix of the communication network; The fourth calculation unit is used to perform spatial interpolation using Gaussian kernel density estimation based on the communication influence, PageRank value list, and the UTM coordinates of the site. The calculation formula is as follows: Where f(x, y) is the estimated value of the propagation strength at the location (x, y), N is the total number of site nodes, h is the bandwidth parameter, and PR(v i ) is the communication influence of the site node, K is the Gaussian kernel function, is the Euclidean distance between the target point (x, y) and the site vi; Diffusion unit: It is used to output the symbol propagation heat map and visualize the spatiotemporal diffusion pattern of the symbol from the core area to the edge area through red-blue gradient color scale.

Citation Information

Patent Citations

  • Cultural relic image retrieval method and device based on multi-modal representation learning model

    CN117473114A

  • Character action recognition analysis method and system based on infrared laser and deep learning

    CN118747911A