A free layout cursive handwriting multi-dimensional feature analysis method

The four-stage cascaded architecture of the free-form cursive handwriting multi-dimensional feature parsing method directly models the features of the original image, decouples the stroke temporal sequence and spatial structure, provides semantic constraints and dynamically integrates multi-dimensional features, solves the parsing accuracy and robustness problems of cursive handwriting in free-form layout scenarios, and achieves efficient and accurate feature parsing.

CN122493474APending Publication Date: 2026-07-31SICHUAN YUNTONG ZHILIAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN YUNTONG ZHILIAN TECHNOLOGY CO LTD
Filing Date
2026-05-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies cannot effectively parse cursive handwriting that spans regions and is not continuously typed. They suffer from low feature parsing accuracy and poor robustness. In particular, in free typesetting scenarios, segmentation error propagation is severe, and the lack of global semantic constraints and low efficiency of multi-dimensional feature fusion are also problems.

Method used

It adopts a four-stage cascaded architecture, including free layout spatial feature modeling, spatiotemporal sequence deep decoding, global semantic prior reasoning and multi-dimensional feature adaptive fusion. Through multi-scale deformable convolution, bidirectional spatiotemporal sequence decoding, global semantic knowledge base and adaptive weighted fusion, it directly models the features of the original image, decouples the stroke temporal sequence and spatial structure, provides semantic constraints and dynamically integrates multi-dimensional features.

Benefits of technology

It achieves high-precision multi-dimensional feature analysis of cursive handwriting, reduces the false judgment rate, and improves robustness and generalization ability. It is suitable for free-form cursive handwriting in different scenarios, taking into account both computational efficiency and engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493474A_ABST
    Figure CN122493474A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-dimensional feature analysis method for free-form cursive handwriting, belonging to the field of computer vision technology. It solves the problem that existing analysis methods cannot effectively analyze cross-regional, non-continuous cursive handwriting. This invention constructs a four-stage cascaded analysis method that integrates free-form spatial feature modeling, spatiotemporal sequence depth decoding, global semantic prior reasoning, and adaptive fusion of multi-dimensional features. It abandons the existing handwriting recognition method that relies on page segmentation preprocessing, directly modeling the spatial distribution features of the original free-form cursive handwriting image, avoiding the transmission of segmentation errors to the feature analysis stage. Simultaneously, it decouples the writing temporal features from the spatial structure features through bidirectional interaction, and provides contextual constraints for local stroke feature analysis through global semantic priors. Finally, it dynamically integrates multi-dimensional features using an adaptive weighted fusion method, achieving multi-dimensional feature analysis of cross-regional, non-continuous cursive handwriting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to a method for analyzing the multidimensional features of freehand cursive handwriting. Background Technology

[0002] Currently, handwritten character feature analysis technology is mainly divided into two categories: traditional machine vision methods and deep learning methods. Traditional machine vision methods achieve handwritten character feature recognition through manual feature engineering such as edge detection, contour extraction, and stroke segmentation, and are suitable for structured handwritten characters such as printed characters and standard regular script / running script. Deep learning methods are based on convolutional neural networks (CNN), recurrent neural networks (RNN), and Transformer, and combine spatiotemporal sequence modeling to achieve handwritten character feature extraction and decoding, which has achieved certain results in semi-structured handwritten character recognition.

[0003] For cursive handwriting, existing technologies mostly involve fine-tuning the network structure based on general handwriting recognition models. Some methods introduce local spatiotemporal sequence modeling or simple semantic context information to try to solve the problems of connected strokes and deformation in cursive handwriting. At the same time, for free layout scenarios, preprocessing steps such as layout detection and text line segmentation are used to assist feature parsing.

[0004] However, free-form cursive script lacks fixed lines of text and character spacing standards. Existing OCR relies on layout analysis, such as using projection methods, CRAFT (Character-Region Awareness For Text detection) or EAST (Efficient and Accuracy Scene Text detection) as a preliminary segmentation step. Segmentation errors directly transmit and amplify recognition errors. Furthermore, it does not model the spatial distribution characteristics of irregular layouts, making it impossible to effectively parse cursive handwriting that spans regions and is not continuous. Summary of the Invention

[0005] To address the aforementioned shortcomings of existing technologies, this invention provides a multi-dimensional feature analysis method for free-form cursive handwriting, which solves the problem that existing analysis methods cannot effectively analyze cursive handwriting that spans regions and is not continuously typed.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for analyzing the multi-dimensional features of free-form cursive handwriting is provided, including the following steps: S1. Obtain the original image of the free-form cursive handwritten font, and perform standardization processing on the original image to obtain the feature sequence of the original image; S2. Input the original image feature sequence into the free layout spatial feature modeling layer. It extracts multi-scale spatial features through multi-scale deformable convolution. It calculates the spatial association weights between each pixel and its surrounding pixels in the multi-scale spatial features based on the spatial distance-based anchor-free spatial distribution perception method. It then performs weighted aggregation of the multi-scale spatial features according to the spatial association weights to obtain a basic spatial feature set containing free layout spatial distribution information. Finally, it enhances the basic spatial feature set based on the spatial attention mechanism to obtain an enhanced spatial distribution feature map. S3. The enhanced spatial distribution feature map is input into the spatiotemporal sequence deep decoding layer, which extracts the writing temporal features of the enhanced spatial distribution feature map through a bidirectional temporal coding branch, and upsamples and parses the enhanced spatial distribution feature map through a spatial decoding branch to obtain a spatial structure feature map. Based on the spatiotemporal cross-attention mechanism, the writing temporal features and the spatial structure feature map are bidirectionally decoupled to obtain a spatiotemporal fusion feature set. S4. Input the spatiotemporal fusion feature set into the global semantic prior reasoning layer. Based on the pre-built cursive script semantic knowledge base, it generates a candidate semantic set through the calligraphy semantic pre-training model, and combines the global semantic context rules to filter and correct the candidate semantic set to generate semantic prior feature vectors. S5. Obtain the morphological features of strokes in the original image feature sequence, and input them together with the enhanced spatial distribution feature map, spatiotemporal fusion feature set and semantic prior feature vector into the multi-dimensional feature adaptive fusion layer. It dynamically weights the input of each dimension through adaptive weighted fusion, and generates a multi-dimensional fusion feature map based on the feature fusion gating mechanism to complete the multi-dimensional feature analysis of free-style cursive handwriting.

[0007] Compared with the prior art, the present invention has the following significant advantages: 1. This invention constructs a four-stage cascaded parsing method that integrates free-form cursive spatial feature modeling, spatiotemporal sequence deep decoding, global semantic prior reasoning, and multi-dimensional feature adaptive fusion. It abandons the existing technical paradigm of handwritten character recognition relying on page segmentation preprocessing, directly modeling the spatial distribution features of the original image of free-form cursive handwritten characters, thus avoiding the transmission of segmentation errors to the feature parsing stage. Compared to the serial processing method of independent spatiotemporal feature modeling in existing technologies, this invention bidirectionally decouples writing temporal features from spatial structural features, provides contextual constraints for local stroke feature parsing through global semantic priors, and finally dynamically integrates multi-dimensional features using an adaptive weighted fusion method. This effectively solves the problems of low feature parsing accuracy and poor robustness caused by the lack of fixed text lines, high stroke density and deformation, and complex semantic context in free-form cursive handwritten characters, achieving high-precision multi-dimensional feature parsing for cross-regional and non-continuous cursive handwritten characters.

[0008] 2. Enhance the robustness of feature parsing under semantic constraints: Construct a global semantic prior reasoning layer, integrate the semantic context of cursive script and the rules of calligraphy art into the feature parsing process, provide accurate semantic guidance for the parsing of local stroke features, solve the parsing error problem caused by the blurring and overlapping of single character strokes, and significantly reduce the misjudgment rate of feature parsing.

[0009] 3. Achieve efficient adaptive fusion of multi-dimensional features: The designed adaptive weighted fusion module can dynamically weight multi-dimensional features such as stroke shape, spatiotemporal sequence, spatial layout, and global semantics. It automatically adjusts the weight of each dimension feature according to the feature distribution of cursive handwriting, which avoids feature redundancy and prevents the loss of key features, thus improving the efficiency of feature fusion while ensuring parsing accuracy.

[0010] 4. Possesses excellent generalization ability and scene adaptability: The system architecture of this invention does not rely on fixed typesetting standards and cursive font styles, and can be directly applied to the free typesetting cursive handwriting analysis of different scenarios such as ancient books, calligraphy works, and handwritten manuscripts. At the same time, it is compatible with the feature analysis of standard handwriting such as regular script and running script, and has strong generalization ability.

[0011] 5. Balancing computational efficiency and engineering applications: The four stages of free-format spatial feature modeling, spatiotemporal sequence deep decoding, global semantic prior reasoning, and multi-dimensional feature adaptive fusion all adopt a fully parallel computing design, eliminating the loop bottleneck of traditional serial feature processing and significantly reducing the inference time of the model; at the same time, the system adopts an end-to-end architecture design, eliminating the need for complex manual feature engineering and preprocessing steps, making it easy to deploy in engineering and meeting the real-time application needs of business-level applications such as ancient book digitization and intelligent text recognition.

[0012] 6. Provides multi-dimensional feature support for calligraphy art analysis: The multi-dimensional fusion feature map obtained by this invention not only includes the semantic features of the characters, but also retains the calligraphy art features such as the writing sequence, stroke shape, and spatial layout of cursive script. It can provide quantitative feature basis for style analysis, author tracing, and art appreciation of cursive script, and expand the application boundaries of handwriting feature analysis technology. Attached Figure Description

[0013] Figure 1 A framework diagram for a multi-dimensional feature analysis method for free-form cursive handwriting; Figure 2 A flowchart for the multi-dimensional feature analysis method of free-form cursive handwriting. Detailed Implementation

[0014] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0015] Example 1

[0016] For cursive handwriting, existing technologies mostly involve fine-tuning the network structure based on general handwriting recognition models. Some methods introduce local spatiotemporal sequence modeling or simple semantic context information to try to solve the problems of connected strokes and deformation in cursive handwriting. At the same time, for free layout scenarios, preprocessing steps such as layout detection and text line segmentation are used to assist feature parsing.

[0017] The existing technology has the following drawbacks: 1. Insufficient decoupling of spatiotemporal features: The strokes of cursive handwriting have a high degree of spatiotemporal continuity. The writing sequence and spatial deformation of the strokes are coupled with each other. Existing methods mostly model spatial features or temporal features separately, without achieving deep fusion and decoding of the two. It is difficult to capture the spatiotemporal evolution of connected strokes, resulting in low completeness and accuracy of feature analysis.

[0018] 2. Lack of global semantic prior constraints: The semantics of cursive handwriting are highly contextual. The deformation of individual character strokes needs to be combined with the global semantics of the whole sentence and paragraph to be accurately interpreted. Existing methods mostly focus on the extraction of individual / local features and do not integrate global semantic priors into the feature interpretation process, which is prone to interpretation errors caused by the blurring of individual character strokes.

[0019] 3. Poor adaptability to free layout: Free layout cursive script has no fixed text lines or character spacing standards. Existing methods rely on the preprocessing of page segmentation in the early stage. The segmentation error will be directly transmitted to the feature analysis stage. Furthermore, it does not model the spatial distribution characteristics of irregular layout, and cannot effectively analyze cursive handwriting with cross-regional and non-continuous layout.

[0020] 4. Inefficient multi-dimensional feature fusion: The analysis of cursive handwriting requires the fusion of multi-dimensional features such as stroke shape, writing sequence, spatial layout, and semantic context. Existing methods use a serial feature fusion approach, and the weight allocation of different dimensional features lacks adaptability, resulting in feature redundancy or loss of key features. It is difficult to balance computational efficiency and analysis accuracy.

[0021] 5. Weak fitting ability for extreme deformed strokes: There are a large number of extreme deformed strokes in cursive script, such as stretching, overlapping, and omission. These features are distributed in a long tail. Existing deep learning models are affected by spectral bias and tend to fit the features of conventional strokes. They are not good at capturing the features of extreme deformed strokes and have a high analytical error rate.

[0022] refer to Figure 1 and Figure 2 To address the aforementioned issues, this embodiment provides a multi-dimensional feature analysis method for free-form cursive handwriting. The method employs a four-stage cascaded architecture: a free-form spatial feature modeling layer, a spatiotemporal sequence deep decoding layer, a global semantic prior reasoning layer, and a multi-dimensional feature adaptive fusion layer. The core principle is to break away from the traditional paradigm of "first page segmentation, then feature analysis," directly modeling the spatial distribution features of the original image of the free-form cursive script; achieving coupling and decoupling of stroke writing temporal sequence and spatial deformation through bidirectional spatiotemporal sequence decoding; constructing a global semantic prior reasoning mechanism based on a semantic knowledge base to provide semantic constraints for local feature analysis; and designing an adaptive feature fusion module to achieve dynamic weighted fusion of multi-dimensional features such as stroke shape, spatiotemporal sequence, spatial layout, and global semantics, ultimately achieving accurate feature analysis of the cursive handwriting. The method includes the following steps: Step S1: Obtain the original image of the free-form cursive handwriting, and perform standardization processing on the original image to obtain the feature sequence of the original image.

[0023] Step S2: Input the original image feature sequence into the free layout spatial feature modeling layer. This layer extracts multi-scale spatial features through multi-scale deformable convolution. Based on a spatial distance-based, anchor-free spatial distribution perception method, it calculates the spatial association weights between each pixel and its surrounding pixels in the multi-scale spatial features. The multi-scale spatial features are then weighted and aggregated according to these spatial association weights to obtain a basic spatial feature set containing free layout spatial distribution information. Finally, the basic spatial feature set is enhanced using a spatial attention mechanism to obtain an enhanced spatial distribution feature map. Specifically, in the spatial association weights between each pixel and its surrounding pixels in the multi-scale spatial features, the surrounding pixels include pixels within a preset range centered on the pixel in the multi-scale spatial feature.

[0024] Step S3: Input the enhanced spatial distribution feature map into the spatiotemporal sequence deep decoding layer. It extracts the writing temporal features of the enhanced spatial distribution feature map through a bidirectional temporal coding branch, and upsamples and parses the enhanced spatial distribution feature map through a spatial decoding branch to obtain a spatial structure feature map. Based on the spatiotemporal cross-attention mechanism, it performs bidirectional interactive decoupling between the writing temporal features and the spatial structure feature map to obtain a spatiotemporal fusion feature set.

[0025] Step S3 also includes completing the spatiotemporal fusion feature set. The completion method is as follows: The system identifies stroke deformation regions (extremely deformed stroke regions) where the feature matching error between writing temporal features and spatial structural features exceeds a preset threshold during the spatiotemporal fusion process. Based on stroke topological constraints and writing temporal constraints, the system performs interpolation to complete the stroke deformation regions. The stroke topological constraints are obtained by extracting the stroke skeleton and tracing the trajectory of the enhanced spatial distribution feature map, while the writing temporal constraints are obtained based on the extraction of writing temporal features.

[0026] Step S4: Input the spatiotemporal fusion feature set into the global semantic prior inference layer. Based on the pre-built cursive script semantic knowledge base, it generates a candidate semantic set through the calligraphy semantic pre-training model, and combines the global semantic context rules to filter and correct the candidate semantic set to generate semantic prior feature vectors.

[0027] Step S5: Obtain the morphological features of strokes in the original image feature sequence, and input them together with the enhanced spatial distribution feature map, spatiotemporal fusion feature set and semantic prior feature vector into the multi-dimensional feature adaptive fusion layer. The layer dynamically weights the input of each dimension through adaptive weighted fusion, and generates a multi-dimensional fusion feature map based on the feature fusion gating mechanism to complete the multi-dimensional feature analysis of free-style cursive handwriting.

[0028] Among them, obtaining the morphological features of strokes in the original image or the feature sequence of the original image is an existing technology. For the original image, it includes binarizing the original image and extracting the skeleton to obtain the skeleton features of each stroke in the original image. Based on the skeleton features, the basic morphological features of the strokes are determined. The basic morphological features of the strokes include at least one of the following: stroke thickness, ink density, and stroke curvature.

[0029] Step S5 includes: S51. Perform feature standardization on morphological features, enhanced spatial distribution feature map, spatiotemporal fusion feature set and semantic prior feature vector respectively, and input them into the adaptive weighted fusion layer for dynamic weighting to obtain weighted fusion features.

[0030] S52. Based on the feature fusion gating mechanism, the weighted fusion features are suppressed and enhanced by the update gate and the reset gate, respectively, to obtain the gated fusion features.

[0031] S53. Perform residual connections on the gated fusion features to enhance the stroke deformation region features in the gated fusion features, and integrate and map the enhanced gated fusion features through a multilayer perceptron to output a multidimensional fusion feature map.

[0032] Specifically, step S1 involves acquiring the original image (with spatial resolution adaptively adapted to 512×512 pixels) and inputting it into an unprocessed input adaptation unit. Through pixel value normalization and resolution bilinear interpolation resampling, cursive images from different sources and in different formats are uniformly mapped into a standardized grayscale / color image sequence in the [0,1] interval. This eliminates the need for traditional layout segmentation and text line detection preprocessing, and directly outputs an unbiased original image feature sequence.

[0033] In step S1, the standardization process abandons the layout segmentation and text line detection preprocessing steps commonly used in traditional handwriting recognition technology. That is, the standardized raw image feature sequence output by step S1 completely preserves the original spatial topology and stroke overlapping relationships of the free-form cursive handwriting, without applying any image cropping or segmentation operations based on rectangular anchor boxes or projected contour analysis. The aim is to block the path of segmentation errors generated in the preprocessing stage from propagating to subsequent spatiotemporal decoding layers, providing a lossless global input for the anchor-free spatial distribution perception modeling in step S2.

[0034] Specifically, in step S2, the free-form layout spatial feature modeling layer uses an anchorless spatial distribution perception module to directly extract multi-scale features from the original cursive handwritten image, abandoning the traditional preprocessing steps of page segmentation. Deformable convolution (DCN) is used to downsample the image at multiple scales, capturing stroke spatial features at different scales. A spatial attention mechanism is introduced to model the spatial distribution and stroke overlapping relationships of characters in free-form layout, generating a spatial distribution feature map. Simultaneously, feature enhancement operations suppress background noise and highlight core stroke features. The free-form layout spatial feature modeling layer includes an anchorless spatial distribution perception module, multi-scale deformable convolutional units, and a spatial attention enhancement module.

[0035] As a further aspect of this embodiment, step S2 specifically includes: inputting the original image feature sequence, firstly performing multi-scale downsampling of the image at 4x, 8x, and 16x using a multi-scale deformable convolutional unit. The offset of the deformable convolution is adaptively learned through learnable parameters, which can accurately capture the spatial features of irregular deformation, overlapping, and stretching of cursive strokes, generating a multi-scale spatial feature map. The multi-scale spatial feature map is then input into an anchor-free spatial distribution perception module. This module abandons the anchor box design of traditional object detection, models the spatial pixel distribution of the feature map using a two-dimensional Gaussian kernel, calculates the spatial correlation between pixels, identifies the spatial distribution pattern of text regions in free-form layout, and the hierarchical relationship of stroke overlapping, generating a basic spatial feature set containing spatial distribution information. A spatial attention enhancement module is introduced, employing a dual attention mechanism of channel attention and spatial attention to enhance the features of the basic spatial feature set: channel attention dynamically adjusts the feature weights of different convolutional channels, highlighting the channels corresponding to stroke features; spatial attention weights the pixel regions of the feature map, suppressing background noise such as paper texture and ink stains, and strengthening the feature signals of the core stroke regions. The final output is an enhanced spatial distribution feature map, which serves as the input to the spatiotemporal sequence deep decoding layer. This enhanced spatial distribution feature map contains information on the spatial distribution patterns of free-form layouts, the spatial morphology of strokes, and the regional associations of irregular layouts.

[0036] In step S2, the multi-scale deformable convolutional unit deformable convolution is performed by introducing an offset. To capture the spatial features of irregular strokes, the convolution output formula is as follows: ; in, For multi-scale deformable convolution in Output features at the location; The total number of pixels in the deformable convolution kernel; The first deformable convolution kernel The weight of each sampling point; These are the pixel values ​​of the original image feature sequence; The first deformable convolution kernel Standard offset of each sampling point; For the first The learnable offset of each sampling point is generated by an offset prediction network, which adaptively fits the spatial deformation of cursive strokes.

[0037] The anchorless spatial distribution sensing module models the spatial relationships between pixels using a two-dimensional Gaussian kernel and calculates pixel... and The spatial association weight is expressed as follows: , ; in,; For pixels With pixels Spatial correlation weights between them; The first in the basic spatial feature set 1 pixel; For local neighborhood centered on The first 1 pixel; For The local neighborhood range centered on; For pixels With pixels The Euclidean distance between them; This refers to the Gaussian kernel bandwidth parameter; This is the normalization factor.

[0038] Basic Spatial Feature Set The calculation expression is: .

[0039] Spatial attention enhancement module combined with channel attention and spatial attention The final output is an enhanced spatial distribution feature map, whose calculation expression is: ; ; ; in, To enhance the spatial distribution feature map; Channel attention weights; Spatial attention weights; This is element-wise multiplication; Use the Sigmoid activation function; It is a multilayer perceptron; For global average pooling; This is for global max pooling; for Convolutional layer; This is for channel splicing operations.

[0040] Specifically, in step S3, the spatiotemporal sequence deep decoding layer constructs a bidirectional spatiotemporal coupled decoding unit (B-STDU) to achieve decoupling and fusion of the spatiotemporal features of cursive script strokes. This unit is divided into a temporal encoding branch and a spatial decoding branch: the temporal encoding branch, based on a gated recurrent unit (GRU) combined with a stroke trajectory detection algorithm, reconstructs the writing sequence of cursive script strokes and generates writing temporal features; the spatial decoding branch uses a deconvolutional network to upsample the enhanced spatial distribution feature map, performs refined analysis of the spatial deformation and connecting stroke features of the strokes, and generates a refined spatial structure feature map; through a spatiotemporal cross-attention mechanism, bidirectional interaction between the writing temporal features and the spatial structure feature map is achieved, decoupling the coupling relationship between the two and generating a spatiotemporal fusion feature set.

[0041] The spatiotemporal sequence deep decoding layer includes a bidirectional spatiotemporal coupled decoding unit (B-STDU), a stroke trajectory detection submodule, a spatiotemporal cross-attention mechanism, and a feature completion submodule.

[0042] The core is a bidirectional spatiotemporal coupled decoding unit (B-STDU), which is divided into a temporal coding branch and a spatial decoding branch. The two branches are computed in parallel and feature interaction is achieved through spatiotemporal cross attention.

[0043] Temporal encoding branch: First, the stroke skeleton is extracted from the enhanced spatial distribution feature map through the stroke trajectory detection submodule. Based on pixel connectivity and gradient changes, the starting point, ending point, and connecting stroke nodes of the stroke are identified. This submodule is a mature existing technology in the field of handwriting recognition. The core adopts the Zhang-Suen thinning algorithm (skeleton extraction) + eight-neighbor pixel tracking algorithm (trajectory generation). It can directly reuse the OpenCV functions cv::ximgproc::thinning() and cv::findContours(). The Hanvon handwriting recognition system and the Wacom digital tablet offline handwriting restoration module both adopt the same technical architecture. Then, the stroke skeleton features are input into the bidirectional gated recurrent unit (Bi-GRU). The forward GRU encodes the writing time sequence from the starting point to the ending point of the stroke, and the reverse GRU completes the temporal features of the connecting strokes from the ending point to the starting point. It adaptively learns the writing order rules of the strokes and generates a temporal feature sequence containing time dimension information. At the same time, it completes the temporal features of broken strokes and omitted strokes.

[0044] Spatial Decoding Branch: A combination of deconvolution network and dilated convolution is used to upsample the enhanced spatial distribution feature map and restore the detailed spatial features of the strokes; dilated convolution captures the long-distance spatial relationship of connected strokes through multi-scale dilation rate, and deconvolution performs fine restoration of the edges and contours of the strokes to generate spatial structure feature maps, solving the problem of stroke detail loss caused by traditional convolution. Spatiotemporal feature fusion: A spatiotemporal cross-attention mechanism is introduced to conduct bidirectional attention interaction between writing sequence features and spatial structure feature maps: writing sequence features provide constraints on the "writing order" of spatial features, assisting in the analysis of spatial segmentation of connected strokes; spatial structure feature maps provide references on the "spatial form" of sequence features, correcting sequence recognition errors caused by stroke overlap; The feature completion submodule completes the features of extremely deformed strokes (such as strokes that are severely stretched or overlapped) that are lost during the spatiotemporal fusion process. It achieves intelligent filling of missing features based on the correlation of neighborhood features and finally outputs a spatiotemporal fusion feature set (including the temporal features of strokes, refined spatial features, and spatiotemporal coupling relationship features).

[0045] In this embodiment, the basic completion algorithm of the feature completion submodule adopts a neighborhood feature interpolation strategy based on the fast moving method. Its core logic is the same as the algorithm enabled by the cv::inpaint() function in the OpenCV library using the cv::INPAINT_TELEA flag, and it is also consistent with the core implementation idea of ​​the handwritten stroke completion unit in the Hanvon OCR recognition engine. However, this embodiment makes the following three key customized improvements on the above basic completion algorithm to take into account the spatiotemporal characteristics of cursive handwriting: First, the completion space is shifted: the completion operation is moved from the pixel space of the original image to the 64-channel feature space of the spatiotemporal fusion feature set, directly interpolating and filling missing features at the feature level, avoiding the interference of texture artifacts that may be generated by pixel-level completion on subsequent feature parsing. This improvement allows the completion operation and subsequent feature extraction, decoding, and fusion processes to be carried out in a unified high-dimensional semantic space, ensuring feature consistency.

[0046] Second, specific constraints are introduced: A dual constraint is applied during feature interpolation: first, a stroke topology constraint, based on the stroke skeleton connectivity and trajectory direction output by the upstream stroke trajectory detection submodule, ensuring the completed features strictly conform to the physical connections of cursive strokes; second, a writing temporal constraint, based on the temporal feature sequence output by the bidirectional temporal coding branch, ensuring the completed features follow the writing order of strokes. These dual constraints ensure that the completed features maintain continuity with surrounding strokes in spatial form and are consistent with the overall temporal sequence in writing logic.

[0047] Third, computational efficiency optimization: adopting a regional parallel computing architecture, feature completion operations are only performed on automatically marked extreme deformation regions, while normal stroke regions are skipped directly, avoiding the computational power consumption caused by dense computation of the entire image, and significantly improving processing efficiency while ensuring the completion effect.

[0048] In this embodiment, extreme deformation regions are identified as follows: the local feature matching error between the writing temporal features and the spatial structural feature map during the spatiotemporal fusion process is calculated, and stroke regions with local feature matching errors exceeding a preset threshold are marked as extreme deformation regions. The local feature matching error can be quantitatively evaluated based on the sparsity of the distribution of attention weights in the spatiotemporal cross-attention mechanism or the cosine similarity of feature vectors.

[0049] As a further embodiment, the calculation expression for the spatiotemporal fusion feature set is as follows: ; ; ; ; ; ; ; ; ; in, For spatiotemporal fusion feature set; The attention weight matrix for writing temporal features on the spatial structure feature map; This is the attention weight matrix for the spatial structure feature map on the writing temporal features; This is element-wise multiplication; This is a spatial structure feature diagram; Characteristics of the writing sequence; This is the attention scaling factor; The length of the writing sequence feature; The total number of spatial pixels in the spatial structure feature map; The first in the spatial structure feature diagram Feature vectors of spatial locations; For the first in the writing sequence characteristics Feature vectors at each time step; For the deconvolution layer at position The output; For the hollow convolutional layer at position The output; To enhance the spatial distribution feature map at location Features; These are the pixel position coordinates; This is the deconvolution stride; The dilation rate of the convolution with holes; This represents the total number of sampling points for the deconvolution kernel; The first in the deconvolution kernel The weight of each sampling point; This represents the total number of sampling points for the dilated convolution kernel; The first in the hollow convolution kernel The weight of each sampling point; For the positive GRU in the first The hidden state vector at each time step. ; , and The positive GRU at the 1st, 2nd and 3rd respectively The hidden state vector at each time step; For the reverse GRU in the first The hidden state vector at each time step. ; , and The reverse GRU at the 1st, 2nd and 3rd respectively The hidden state vector at each time step; This represents the total number of time steps for the stroke trajectory; This is a vector concatenation operation; For the first The stroke skeleton feature vector input at each time step; , The positive GRU at the 1st Update and reset gates for time steps; , The reverse GRU at the 1st Update and reset gates for time steps; For the positive GRU in the first Candidate hidden states at a time step; For the reverse GRU in the first Candidate hidden states at a time step; Here is the weight matrix of the GRU unit; The bias vector of the GRU cell; Use the Sigmoid activation function; It is the hyperbolic tangent activation function; This is a vector concatenation operation.

[0050] Specifically, step S4 includes: S41. Input the spatiotemporal fusion feature set into the calligraphy semantic pre-training model, calculate the similarity between the spatiotemporal fusion feature set and each candidate semantic feature vector in the cursive script semantic knowledge base, and generate a candidate semantic set based on the similarity.

[0051] The method for constructing a semantic knowledge base specifically for cursive script is as follows: integrating the basic character database of cursive script (containing the forms of single characters in cursive script by different calligraphers and in different styles), vocabulary database (common words and fixed collocations in cursive script), syntax rules (grammar and semantic logic of ancient / modern texts), and artistic rules of calligraphy (artistic rules of stroke omission, connection, and deformation in cursive script), while incorporating corpus information from ancient books and calligraphy works to construct a structured and updatable semantic knowledge base for cursive script, which serves as the foundation for global semantic priors.

[0052] The initial semantic matching method is as follows: input the spatiotemporal fusion feature set into the calligraphy semantic pre-training large model (based on the Transformer architecture, pre-trained on the basis of cursive calligraphy corpus and handwritten character recognition corpus), the model performs semantic matching based on local stroke features, generates a candidate semantic set containing multiple candidate semantics, and assigns an initial matching confidence score to each candidate semantic.

[0053] S42. Based on the semantic rules in the cursive semantic knowledge base, the candidate semantic set is modified, and candidate semantics that conflict with the global semantic context are removed to obtain the modified candidate semantic set.

[0054] The semantic prior reasoning method is as follows: the semantic prior reasoning module retrieves global contextual information (such as word collocation, syntax rules, and calligraphy transformation rules) related to the candidate semantic set from the cursive script-specific semantic knowledge base, performs reasoning verification on the candidate semantics, eliminates candidate semantics that conflict with the global semantics, and improves the matching confidence.

[0055] S43. Extract the upper and lower semantic features from the spatiotemporal fusion feature set, perform contextual semantic filtering on the corrected candidate semantic set based on the upper and lower semantic features, determine the optimal semantic, and generate a semantic prior feature vector based on the optimal semantic.

[0056] The context semantic filtering method is as follows: the context semantic filtering submodule identifies the boundaries of sentences and paragraphs in cursive handwriting based on the regional correlation of the spatiotemporal fusion feature set, extracts context features, performs secondary filtering on the candidate semantics after reasoning, determines the optimal global semantic prior constraints, and transforms them into semantic prior feature vectors that can be fused with the feature set (containing semantic labels, semantic association weights, and calligraphy rule constraint information).

[0057] As a further embodiment, the expression for calculating the semantic prior feature vector is as follows: ; ; ; ; in, For semantic prior feature vectors; For candidate semantic set The optimal candidate semantics; the set of candidate semantics ,in , and They are the 1st, 2nd, and 3rd in the candidate semantic set, respectively. One candidate semantic; for The corresponding feature vector; for The corrected confidence level; For contextual semantic features; This is element-wise multiplication; For the first The corrected confidence of each candidate semantic; For the first Candidate semantics The initial confidence level; For the first Candidate semantics The semantic rule correction coefficient, The matching degree between candidate semantics and global semantic rules in the cursive script semantic knowledge base is used to determine the semantics. For spatiotemporal fusion feature set; For the first Candidate semantics The corresponding feature vector in the cursive script semantic knowledge base; Let be the magnitude of the vector.

[0058] Specifically, the multi-dimensional feature adaptive fusion layer in step S5 uses an adaptive weighted fusion module (AWFM) as the core fusion unit of the system. The inputs include spatial distribution features, spatiotemporal fusion features, semantic prior features, and stroke morphological features. This module dynamically weights the features across each dimension using a learnable feature weight matrix. It automatically increases the weights of spatiotemporal and semantic prior features for extremely deformed strokes and highly connected strokes, while simplifying calculations and preserving basic morphological features for regular stroke areas. A feature fusion gating mechanism is introduced to suppress feature redundancy and enhance the transmission of key features, ultimately generating a multi-dimensional fusion feature map to complete the multi-dimensional feature analysis of free-form cursive handwriting. The inputs to this layer are multi-dimensional features: enhanced spatial distribution features from the free-form spatial feature modeling layer, spatiotemporal fusion features from the spatiotemporal sequence deep decoding layer, semantic prior features from the global semantic prior inference layer, and basic stroke morphological features extracted from the original image (such as stroke thickness, ink density, and stroke curvature). The core is to achieve intelligent integration of multi-dimensional features through the adaptive weighted fusion module. Specific steps are as follows: First, the feature standardization submodule normalizes features across all dimensions, mapping features of different dimensions and scales to the same feature space and eliminating fusion bias caused by differences in feature scale. The input is then fed into an adaptive weighted fusion module (AWFM), which dynamically weights features across all dimensions using a learnable feature weight matrix. This weight matrix is ​​adaptively learned through model training. For severely deformed stroke regions (such as those with overlapping, stretching, or severe omission), it automatically increases the weights of spatiotemporal fusion features and semantic prior features, achieving accurate parsing using temporal and semantic constraints. For regular stroke regions, it automatically reduces computational complexity while preserving basic stroke morphology and spatial distribution features, balancing parsing efficiency. For cross-regional related regions in free-form layout, it increases the weights of spatial distribution features, reinforcing the spatial regularity of the layout. Finally, feature fusion is introduced. The system employs a gating mechanism, drawing inspiration from the gating logic of a gated loop unit. It designs update and reset gates to filter the weighted multidimensional features: the update gate determines the retention ratio of preceding features, and the reset gate determines the introduction ratio of new features, effectively suppressing feature redundancy and strengthening the transmission and fusion of key features. An extreme feature enhancement submodule is added to specifically enhance the extreme deformed stroke features with long-tail distribution in cursive script. Through feature residual connections, the information of extreme features is integrated into the fused features, addressing the problem of weak fitting ability of existing models for extreme features. Finally, a multilayer perceptron (MLP) is used to integrate and map the fused features, outputting a multidimensional fused feature map. This map contains all multidimensional feature information, including stroke morphology, writing sequence, spatial layout, semantic context, and calligraphic art rules of cursive handwriting, achieving the core objective of feature parsing.

[0059] The expression for feature standardization is: ,in, For the first dimensional features, ; ; ; ; To analyze the target variable; For the first Dimensional features and parsing objectives Mutual information; To analyze the target marginal entropy; For a given number 3D feature parsing target Conditional entropy; For the first Dimensional standardization features; For the first Dimensional original input features; For the first The mean of the dimensional features; For the first Standard deviation of dimensional features. To enhance the spatial distribution feature map; For spatiotemporal fusion feature set; These are semantic prior feature vectors; This refers to the morphological features of strokes. Feature standardization is a prerequisite for achieving fully objective adaptive fusion. Its core necessity lies in the fact that the four types of features input to this module are completely heterogeneous: spatial distribution features come from the downsampling output of multi-scale deformable convolution, with numerical ranges typically concentrated in the [0, 1] interval; spatiotemporal fusion features come from the encoding and decoding of Bi-GRU and deconvolution, with numerical values ​​potentially exhibiting positive or negative offsets and large variances; semantic prior features come from the vector mapping between the knowledge base and Transformer, with dimensions and scales independent of visual features; and basic stroke morphological features come from manual or basic visual extraction, with numerical ranges varying greatly depending on specific features (such as ink density and stroke curvature). Without standardization, features with large numerical scales will "dominate" weight allocation in subsequent weighted fusion, causing the objective weight logic based on information gain and homoscedastic uncertainty to fail. This step independently performs zero-mean, unit-variance standardization on each dimension feature, eliminating scale differences while fully preserving the relative distribution structure within each feature, ensuring the fairness and accuracy of subsequent objective weight calculations.

[0060] The Adaptive Weighted Fusion (AWFM) method is as follows: First calculate the... Dimensional features and parsing objectives Mutual information: ;in, To analyze the target variable; For the first Dimensional features and parsing objectives Mutual information; To analyze the target marginal entropy; For a given number 3D feature parsing target The conditional entropy.

[0061] Information gain is calculated based on mutual information. And generate normalized weights: ; ;in, They are respectively The corresponding feature weights satisfy ; For the first Information gain of dimensional features.

[0062] Weighted feature fusion: ; ;in, Weighted fusion features The method for local weight adaptive adaptation of extreme deformation regions (stroke deformation regions where the feature matching error between writing temporal features and spatial structural features exceeds a preset threshold during spatiotemporal fusion) is as follows: For dynamic weight adjustment in regions of extreme deformation, an objective weight offset expression based on local feature confidence is designed: ;in, For the pixel-level local region, the first Adaptive weights for dimensional features ; For the first The analytical confidence of a feature in a local region is objectively calculated from the cosine similarity of the feature matching degree; the confidence of spatiotemporal / semantic features in extremely deformable regions is higher, and the weights are automatically increased, without any manually set weight adjustment rules throughout the process. The feature fusion gating mechanism method is as follows: Update Gate With Reset Door : ; ; Gated output: ; in, To update the gate output; To reset the gate output; This is element-wise multiplication; This is a gating fusion feature.

[0063] The feature fusion gating mechanism is a core module for solving multidimensional feature redundancy and strengthening the transmission of key features. Its design addresses two major pain points in the feature fusion of cursive handwriting: First, the information overlap and redundancy between features—spatial distribution features already contain the basic morphological information of stroke outlines, and semantic prior features also implicitly contain some spatiotemporal correlation logic through global constraints. If they are directly spliced ​​and weighted, the redundant information will interfere with the expression of key features. Second, the dynamic requirements of feature fusion—the complementarity of features in different regions varies greatly: regular stroke regions can rely on spatial and morphological features, while extreme deformation regions require strong complementarity of spatiotemporal and semantic features.

[0064] This module borrows the "gated filtering" logic from the Gated Circular Unit (GRU): Reset door : Control the weighted fusion features of new inputs "In the process, which information needs to be re-filtered and introduced?" The closer it is to 1, the more completely the original information of the new feature is preserved; The closer it is to 0, the more likely it is to filter out redundant parts of the new features; Update Gate Controls the retention ratio between the previous fusion state (implicit in the iterative logic of the gating mechanism) and the current new feature. The closer it is to 1, the more it depends on the current new features; The closer it is to 0, the more key features already selected in the previous step are retained; Final gated output By using element-wise multiplication and summation, we achieve dynamic suppression of redundant features and enhanced transmission of key features (such as temporal constraints of extreme deformations and contextual guidance of global semantics), providing a pure feature foundation for subsequent enhancement of extreme features.

[0065] The calculation expression for feature enhancement of stroke deformation region features is as follows: ;in, For multi-dimensional fusion feature maps; is the feature enhancement coefficient for the deformed region of the stroke.

[0066] In this implementation, feature enhancement of stroke deformation regions, specifically extreme feature enhancement, is the key design for solving the problem of "long-tailed distribution of extreme deformed strokes" in cursive script. Its core addresses the spectral bias problem in deep learning models—during training, the model prioritizes fitting "head features" (regular strokes, simple connections) that account for a high proportion in the dataset, while ignoring "tail features" (extremely deformed strokes with large stretching, multi-layer overlapping, and complete character omissions) that have a low proportion but are crucial to analytical accuracy, resulting in a persistently high analytical error rate in extreme regions. This module constructs a "dedicated tail feature transfer channel" through residual connections: a residual network. It employs a lightweight multilayer perceptron (MLP) or low-receptivity convolutional structure to specifically fit extreme deformation regions with high training errors, learning unique patterns in tail features; extreme feature enhancement coefficients In conjunction with the "Adaptive Local Weighting in Extreme Deformation Regions" module, it utilizes local feature confidence levels. Objective mapping (such as) ): Confidence level of extreme deformation regions Low (high difficulty in parsing) The residual network output is automatically approached to 1, thus being fully amplified and directly passed to the final fused feature through the residual path. To avoid tail features being compressed or filtered by gating mechanisms in intermediate layers; confidence level of regular stroke regions. high, The residual network automatically approaches zero, suppressing its output and avoiding overfitting due to excessive enhancement of head features. Finally, the sum of the residual connections and the gated output retains the stable features of the normal regions while specifically enhancing the long-tail features of the extreme deformation regions, perfectly adapting to the feature distribution characteristics of cursive handwriting.

[0067] Example 2

[0068] This embodiment is based on Embodiment 1 and further defines the improvements. Specifically, it provides a complete method for end-to-end joint training of the free layout spatial feature modeling layer, the spatiotemporal sequence deep decoding layer, the global semantic prior reasoning layer, and the multi-dimensional feature adaptive fusion layer. Other parts not mentioned refer to Embodiment 1 or existing technologies.

[0069] refer to Figure 1 and Figure 2 The end-to-end joint training method includes a feature feedback optimization layer, which specifically includes the following steps: S61. Obtain the multidimensional fusion feature map through the forward propagation process of steps S1 to S5 in Example 1. This feature map integrates spatial distribution features, spatiotemporal fusion features, semantic prior features, and stroke morphology features, and forms the basis for subsequent loss calculations.

[0070] S62, Based on Multidimensional Fusion Feature Map Calculate the joint loss function The formula for calculating the joint loss function is: ; in, Joint loss function; Feature matching loss is used to characterize the matching error between the multidimensional fused feature map and the labeled features; Semantic loss is used to characterize the error between the semantic prior feature vector and the labeled semantics; Spatial feature loss is used to characterize the error between the enhanced spatial distribution feature map and the labeled spatial distribution. Temporal feature loss is used to characterize the error between the written temporal features and the labeled temporal features; Information entropy of corresponding dimensional features; The total information entropy of all four dimensions of features ; The corresponding homoscedasticity uncertainty noise parameter is automatically learned by the model during end-to-end training and is used to adaptively balance the contribution weights of each loss term.

[0071] S63, Based on Joint Loss Function The gradients of the learnable parameters in each layer are calculated using the backpropagation algorithm, and the learnable parameters are updated based on these gradients. The learnable parameters specifically include: the hidden layer parameters of the bidirectional temporal coding branch in the spatiotemporal sequence deep decoding layer, the spatiotemporal cross-attention weights in the spatiotemporal sequence deep decoding layer, and the semantic matching threshold of the global semantic prior inference layer.

[0072] S64. Iterate training until convergence. Repeat steps S61 to S63 until the joint loss function is achieved. If the preset convergence conditions are met (such as the loss value being lower than the preset value or no longer decreasing after multiple consecutive iterations), the end-to-end joint training of the free layout spatial feature modeling layer, the spatiotemporal sequence deep decoding layer, the global semantic prior inference layer, and the multi-dimensional feature adaptive fusion layer is completed.

[0073] As a further preferred embodiment, to improve the adaptability to local feature differences in free-form cursive script, this embodiment also provides a pixel-level local adaptive weight field mechanism. This mechanism constructs a local weight field that is completely aligned with the size of the multi-dimensional fused feature map, so that the loss weight at each pixel position is objectively determined by the local feature statistics in the neighborhood of that pixel, thereby achieving refined and fully objective weight allocation.

[0074] The expression for calculating the pixel-level local adaptive weight field is: ; in: Pixel position First Adaptive weights for each loss term; In pixels Within the neighborhood window centered Local information entropy of dimensional features; In pixels The local total information entropy of all four-dimensional features within the neighborhood window centered on the center; Within the local neighborhood The local uncertainty of each loss term is objectively calculated from the variance of the feature matching error within that neighborhood. Correspondingly, the pixel-level joint loss function... The calculation expression is: ; in: Full pixel space of multidimensional fused feature maps; Pixels The local loss value at the location. Through the pixel-level adaptive weight field mechanism described above, the joint loss function can automatically adjust the contribution ratio of each loss term according to the feature distribution and parsing difficulty of different regions of the image, so that the model pays more attention to difficult areas such as extremely deformed strokes and cross-regional strokes during the training process, and further improves the overall parsing accuracy.

[0075] As another preferred embodiment, the training method further includes introducing a feature feedback mechanism between each layer to construct a closed-loop optimization system. The specific implementation steps are as follows: the analysis result of the multi-dimensional fused feature map is converted into a feedback feature signal, which includes information such as the confidence level of feature analysis, error region marking, and semantic matching accuracy.

[0076] Feedback feature signals are injected into the following two levels respectively: (1) Feedback to the spatiotemporal sequence deep decoding layer: For the error region marked by the feedback feature signal, dynamically adjust the hidden layer parameters of the bidirectional temporal coding branch and the spatiotemporal cross attention weight to correct spatiotemporal feature decoding deviations such as continuous parsing errors and temporal restoration deviations.

[0077] (2) Feedback to the global semantic prior reasoning layer: Based on the semantic matching accuracy in the feedback feature signal, dynamically adjust the matching threshold of the calligraphy semantic pre-training large model and the reasoning rules of the semantic prior reasoning module to improve the accuracy of the global semantic prior.

[0078] For the multidimensional feature parsing model composed of the free-form layout spatial feature modeling layer, the spatiotemporal sequence deep decoding layer, the global semantic prior reasoning layer, and the multidimensional feature adaptive fusion layer, the feedback feature signal participates in the backpropagation training of the model. It is updated together with the above-mentioned learnable parameters through the gradient descent algorithm, thus forming a closed-loop iterative mechanism of "parsing-feedback-optimization". This enables the model to continuously optimize the parameters of each layer during the training process, and continuously improve the accuracy and robustness of multidimensional feature parsing of free-form cursive handwriting.

[0079] In summary, the beneficial effects of this plan are as follows: This scheme abandons the traditional technical paradigm that relies on page segmentation preprocessing. Instead, it directly performs multi-scale deformable convolutional modeling on the original image through anchor-free spatial distribution perception, eliminating the problem of segmentation error propagation and amplification. Through bidirectional temporal coding and spatial decoding branches in the spatiotemporal sequence deep decoding layer, and by introducing a spatiotemporal cross-attention mechanism, it achieves bidirectional interactive decoupling of writing temporal features and spatial structural features, capturing the spatiotemporal evolution law of connected strokes. At the same time, it constructs a global semantic prior reasoning layer, which generates a candidate semantic set based on the cursive semantic knowledge base and the calligraphy semantic pre-trained model, and filters and corrects it in combination with global context rules, providing strong semantic constraints for local stroke analysis and significantly reducing the misjudgment rate of blurred and overlapping strokes. Furthermore, through a multi-dimensional feature adaptive fusion layer, it performs dynamic weighted fusion of four types of features—stroke shape, spatial distribution, spatiotemporal fusion, and semantic prior—driven by information gain. It uses a feature fusion gating mechanism to suppress redundancy and strengthen key features, and uses residual connections to specifically strengthen the feature expression of extremely deformed stroke regions, balancing analysis accuracy and computational efficiency. Furthermore, this solution provides an end-to-end joint training method. The joint loss function automatically balances each loss term using information entropy and homoscedastic uncertainty. It can introduce pixel-level local adaptive weight fields and feature feedback mechanisms, enabling the model to adaptively learn different layout and deformation styles. It has good generalization ability and engineering deployment value. Moreover, the multi-dimensional fusion feature map obtained by analysis completely preserves the writing sequence, stroke shape, spatial layout and calligraphy art rules, providing quantitative feature support for the analysis of cursive calligraphy style, author tracing and art appreciation. It solves the problem that existing technologies cannot effectively analyze cross-regional and non-continuous cursive handwriting.

Claims

1. A method for multi-dimensional feature analysis of free-form cursive handwriting, characterized in that, Including the following steps: S1. Obtain the original image of the free-form cursive handwritten font, and perform standardization processing on the original image to obtain the feature sequence of the original image; S2. Input the original image feature sequence into the free layout spatial feature modeling layer, which extracts multi-scale spatial features through multi-scale deformable convolution, calculates the spatial association weights between each pixel and its surrounding pixels in the multi-scale spatial features based on the spatial distance-based anchor-free spatial distribution perception method, and performs weighted aggregation of the multi-scale spatial features according to the spatial association weights to obtain a basic spatial feature set containing free layout spatial distribution information, and enhances the basic spatial feature set based on the spatial attention mechanism to obtain an enhanced spatial distribution feature map; S3. The enhanced spatial distribution feature map is input into the spatiotemporal sequence deep decoding layer, which extracts the writing temporal features of the enhanced spatial distribution feature map through a bidirectional temporal coding branch, and performs upsampling and parsing of the enhanced spatial distribution feature map through a spatial decoding branch to obtain a spatial structure feature map. Based on the spatiotemporal cross-attention mechanism, the writing temporal features and the spatial structure feature map are bidirectionally decoupled to obtain a spatiotemporal fusion feature set. S4. Input the spatiotemporal fusion feature set into the global semantic prior reasoning layer, which is based on the pre-built cursive script semantic knowledge base, generates a candidate semantic set through the calligraphy semantic pre-training model, and filters and corrects the candidate semantic set in combination with the global semantic context rules to generate a semantic prior feature vector. S5. Obtain the morphological features of the strokes in the original image feature sequence, and input them together with the enhanced spatial distribution feature map, the spatiotemporal fusion feature set and the semantic prior feature vector into the multi-dimensional feature adaptive fusion layer. The layer dynamically weights the input of each dimension through adaptive weighted fusion, and generates a multi-dimensional fusion feature map based on the feature fusion gating mechanism to complete the multi-dimensional feature analysis of free-style cursive handwriting.

2. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 1, characterized in that, The calculation expression for the enhanced spatial distribution feature map is as follows: ; ; ; ; , ; ; in, To enhance the spatial distribution feature map; Channel attention weights; Spatial attention weights; The basic spatial feature set; This is element-wise multiplication; Use the Sigmoid activation function; It is a multilayer perceptron; For global average pooling; This is for global max pooling; for Convolutional layer; For channel splicing operations; The first in the basic spatial feature set 1 pixel; For local neighborhood centered on The first 1 pixel; For The local neighborhood range centered on; For pixels With pixels Spatial correlation weights between them; For multi-scale deformable convolution in Output features at the location; For pixels With pixels The Euclidean distance between them; This refers to the Gaussian kernel bandwidth parameter; Normalization factor; The total number of pixels in the deformable convolution kernel; The first deformable convolution kernel The weight of each sampling point; These are the pixel values ​​of the original image feature sequence; The first deformable convolution kernel Standard offset of each sampling point; For the first The learnable offset of each sampling point.

3. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 1, characterized in that, Step S3 also includes completing the spatiotemporal fusion feature set, the completion method being: The system identifies stroke deformation regions where the feature matching error between writing temporal features and spatial structural features exceeds a preset threshold during the spatiotemporal fusion process. Based on stroke topological constraints and writing temporal constraints, the system performs interpolation to complete the stroke deformation regions. The stroke topological constraints are obtained by extracting the stroke skeleton and tracing the trajectory of the enhanced spatial distribution feature map, while the writing temporal constraints are obtained based on the extracted writing temporal features.

4. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 1, characterized in that, The calculation expression for the spatiotemporal fusion feature set is as follows: ; ; ; ; ; ; ; ; ; in, For spatiotemporal fusion feature set; The attention weight matrix for writing temporal features on the spatial structure feature map; This is the attention weight matrix for the spatial structure feature map on the writing temporal features; This is element-wise multiplication; This is a spatial structure feature diagram; Characteristics of the writing sequence; This is the attention scaling factor; The length of the writing sequence feature; The total number of spatial pixels in the spatial structure feature map; The first in the spatial structure feature diagram Feature vectors of spatial locations; For the first in the writing sequence characteristics Feature vectors at each time step; For the deconvolution layer at position The output; For the hollow convolutional layer at position The output; To enhance the spatial distribution feature map at location Features; These are the pixel position coordinates; This is the deconvolution stride; The dilation rate of the convolution with holes; This represents the total number of sampling points for the deconvolution kernel; The first in the deconvolution kernel The weight of each sampling point; This represents the total number of sampling points for the dilated convolution kernel; The first in the hollow convolution kernel The weight of each sampling point; For the positive GRU in the first The hidden state vector at each time step. ; , and The positive GRU at the 1st, 2nd and 3rd respectively The hidden state vector at each time step; For the reverse GRU in the first The hidden state vector at each time step. ; , and The reverse GRU at the 1st, 2nd and 3rd respectively The hidden state vector at each time step; This represents the total number of time steps for the stroke trajectory; This is a vector concatenation operation; For the first The stroke skeleton feature vector input at each time step; , The positive GRU at the 1st Update and reset gates for time steps; , The reverse GRU at the 1st Update and reset gates for time steps; For the positive GRU in the first Candidate hidden states at a time step; For the reverse GRU in the first Candidate hidden states at a time step; Here is the weight matrix of the GRU unit; The bias vector of the GRU cell; Use the Sigmoid activation function; It is the hyperbolic tangent activation function; This is a vector concatenation operation.

5. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 1, characterized in that, Step S4 includes: S41. Input the spatiotemporal fusion feature set into the calligraphy semantic pre-training model, calculate the similarity between the spatiotemporal fusion feature set and each candidate semantic feature vector in the cursive script semantic knowledge base, and generate a candidate semantic set based on the similarity. S42. Based on the semantic rules in the cursive script semantic knowledge base, the candidate semantic set is modified, and candidate semantics that conflict with the global semantic context are eliminated to obtain the modified candidate semantic set. S43. Extract the upper and lower semantic features from the spatiotemporal fusion feature set, perform contextual semantic filtering on the modified candidate semantic set based on the upper and lower semantic features, determine the optimal semantic, and generate a semantic prior feature vector based on the optimal semantic.

6. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 5, characterized in that, The expression for calculating the semantic prior feature vector is as follows: ; ; ; ; in, These are semantic prior feature vectors; The optimal candidate semantics; for The corresponding feature vector; for The corrected confidence level; For contextual semantic features; This is element-wise multiplication; For the first The corrected confidence of each candidate semantic; For the first Candidate semantics The initial confidence level; For the first Candidate semantics Semantic rule correction coefficients; For spatiotemporal fusion feature set; For the first Candidate semantics The corresponding feature vector in the cursive script semantic knowledge base; Let be the magnitude of the vector; This represents the total number of time steps for the stroke trajectory.

7. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 3, characterized in that, Step S5 includes: S51. Perform feature standardization processing on the morphological features, the enhanced spatial distribution feature map, the spatiotemporal fusion feature set, and the semantic prior feature vector, and input them into the adaptive weighted fusion layer for dynamic weighting to obtain weighted fusion features; S52. Based on the feature fusion gating mechanism, the weighted fusion features are suppressed and enhanced by the update gate and the reset gate, respectively, to obtain the gated fusion features. S53. Perform residual connections on the gated fusion features to enhance the stroke deformation region features in the gated fusion features, and integrate and map the enhanced gated fusion features through a multilayer perceptron to output a multidimensional fusion feature map.

8. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 7, characterized in that, The calculation expression for the multidimensional fused feature map is as follows: ; ; ; ; ; ; ; ; ; ; in, For multi-dimensional fusion feature maps; This is a gating fusion feature; The feature enhancement coefficient for the stroke deformation region; For residual network functions; To update the gate output; To reset the gate output; This is element-wise multiplication; For weighted fusion features; This is the weight matrix in the gating mechanism; This is the bias vector in the gating mechanism; Use the Sigmoid activation function; They are respectively The corresponding feature weights; To enhance the spatial distribution feature map; For spatiotemporal fusion feature set; These are semantic prior feature vectors; The morphological characteristics of strokes; For the pixel-level local region, the first Adaptive weights for dimensional features ; For the first The analytical confidence of the dimensional feature in the local region; For the first Information gain of 3D features; For the first Dimensional features, ; ; ; ; To analyze the target variable; For the first Dimensional features and parsing objectives Mutual information; To analyze the target marginal entropy; For a given number 3D feature parsing target Conditional entropy; For the first Dimensional standardization features; For the first Dimensional original input features; For the first The mean of the dimensional features; For the first Standard deviation of dimensional features.

9. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 1, characterized in that, It also includes an end-to-end training method for the free-form layout spatial feature modeling layer, the spatiotemporal sequence deep decoding layer, the global semantic prior inference layer, and the multi-dimensional feature adaptive fusion layer, the training method comprising: S61. Obtain the multidimensional fusion feature map through steps S1 to S5; S62. Calculate the joint loss function based on the multi-dimensional fused feature map; S63. Based on the joint loss function, the gradient of the learnable parameters is calculated using the backpropagation algorithm, and the learnable parameters are updated according to the gradient; wherein, the learnable parameters include the learnable parameters in the spatiotemporal sequence deep decoding layer and the global semantic prior inference layer; S64. Repeat steps S61 to S63 until the joint loss function meets the preset convergence condition, and complete the end-to-end joint training.

10. The method for multi-dimensional feature analysis of free-form cursive handwriting according to claim 9, characterized in that, The formula for calculating the joint loss function is: ; in, For the joint loss function; For the first One loss item; The feature matching loss is used to characterize the matching error between the multidimensional fused feature map and the labeled features; Semantic loss is used to characterize the error between the semantic prior feature vector and the labeled semantics; Spatial feature loss is used to characterize the error between the enhanced spatial distribution feature map and the labeled spatial distribution; The temporal feature loss is used to characterize the error between the writing temporal features and the annotation temporal features; for Information entropy of corresponding dimensional features; The total information entropy of all four dimensions of features; for The corresponding homoscedasticity uncertainty noise parameter.