Ink painting brush stroke technique deep learning evaluation system and method

By using multimodal data acquisition and deep learning networks, we have achieved real-time and objective quantitative evaluation of ink painting brushwork techniques, solving the problem of insufficient feedback in traditional teaching and improving learning efficiency.

CN122365149APending Publication Date: 2026-07-10ZHEJIANG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG NORMAL UNIV
Filing Date
2026-04-15
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies struggle to provide real-time, objective, and quantifiable feedback on brushwork techniques in ink painting, particularly in the dynamic capture and evaluation of form, force, and ink, resulting in low learning efficiency.

Method used

A multimodal data acquisition module is used to simultaneously capture the three-dimensional spatial coordinates of the pen tip, pen pressure, and ink diffusion images. Data fusion and feature extraction are performed through a deep learning network to construct an authoritative quantitative evaluation system that provides real-time feedback and guidance.

Benefits of technology

It achieves high-precision identification and quantitative evaluation of brushstroke techniques, provides personalized and real-time learning guidance, and improves learning efficiency and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365149A_ABST
    Figure CN122365149A_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning evaluation system and method for ink painting brushwork techniques. The system includes: a multimodal data acquisition module, a data fusion and preprocessing module, a brushwork deep feature extraction module based on a temporal convolutional network, a brushwork technique recognition and quantitative evaluation module, and a real-time feedback and correction guidance generation module. The multimodal data acquisition module synchronously collects brushwork process data; performs temporal synchronization and feature extraction on the data to generate a multimodal feature matrix; extracts deep temporal feature vectors through a dilated causal temporal convolutional network; outputs technique classification probabilities and quantitative scores in parallel; and generates real-time visual correction guidance. Through the above technical solutions, this invention can significantly improve the recognition accuracy of complex brushwork techniques and provide scientific quantitative evaluation and personalized teaching feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence application technology in the art field, specifically to a deep learning evaluation system and method for ink painting brushwork techniques. Background Technology

[0002] The exquisite beauty of traditional Chinese ink painting lies in its unique brushwork techniques. Its artistic essence and expressiveness rely on a series of complex brushwork techniques, such as the central brushstroke, side brushstroke, texturing strokes, dry brush, and splashed ink. These brushwork techniques, through the interaction of the brush tip's angle, strength, and speed with ink, water, and Xuan paper, create ever-changing ink and brush effects. However, the teaching and inheritance of ink painting has long faced a challenge: its core evaluation criteria rely on the oral and hands-on transmission of experience between master and apprentice. Learners struggle to obtain immediate, objective, and quantifiable feedback on their own brushwork, resulting in a lengthy learning cycle and slow skill improvement. Existing technologies have significant limitations. For example, digital input devices (such as graphics tablets) can capture the two-dimensional coordinates and pressure of a pen, but they cannot perceive the three-dimensional dynamics of the pen tip, the actual contact mechanics between the pen tip and the Xuan paper, or capture the core "ink diffusion" process of ink painting, resulting in insufficient information dimensions. Computer vision methods that track pen or hand movements using cameras are easily interfered with and typically analyze "form" in isolation, without integrating "force" and "ink" data. Traditional machine learning methods rely on manually designed features (such as average speed and inflection points) and use shallow models for classification. These methods struggle to depict the complex dynamic relationships between multimodal signals during brushwork and cannot effectively model multi-scale temporal dependencies ranging from millisecond-level micro-movements to second-level rhythms, resulting in limited recognition accuracy and generalization ability. Moreover, most existing methods can only achieve simple technique classification and lack the ability to continuously and finely quantify and evaluate brushwork techniques, leading to a one-sided evaluation system. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies by providing a deep learning evaluation system for ink painting brushwork techniques, simultaneously capturing the complete dynamics of "form, force, and ink." Through in-depth analysis and the establishment of an authoritative quantitative evaluation system, it forms personalized real-time guidance, transforming experience-based teaching into a standardized and personalized teaching process based on data and models.

[0004] To achieve the above objectives, this invention provides a deep learning evaluation system for ink painting brushwork techniques, which includes the following components:

[0005] Multimodal data acquisition module: used to synchronously acquire data on the brushstroke process of ink painting on paper. The data includes the time sequence of the three-dimensional spatial coordinates of the brush tip, the time sequence of the pressure of the brush handle, and the time sequence of the ink diffusion image of the brushstroke area.

[0006] Data fusion and preprocessing module: used to perform time stamp alignment and interpolation synchronization on the pen tip three-dimensional spatial coordinate time series, pen barrel pressure time series and ink diffusion image time series acquired by the multimodal data acquisition module, and extract pen movement trajectory features, pressure dynamic features and ink color morphology features respectively to generate a multimodal feature matrix with a unified time dimension;

[0007] The deep feature extraction module for brushstrokes based on temporal convolutional networks is used to receive the multimodal feature matrix output by the data fusion and preprocessing module. Through a multi-layer dilated causal temporal convolutional network architecture, the multimodal feature matrix is ​​processed. By stacking dilated convolutional layers, the receptive field is expanded exponentially to capture multi-scale temporal dependencies from milliseconds to seconds, and a deep temporal feature vector that fuses trajectory, pressure and ink color dynamic coupling information is output.

[0008] The brushwork technique recognition and quantitative evaluation module receives a deep temporal feature vector from the brushwork deep feature extraction module based on a temporal convolutional network. This vector includes a pre-trained technique classification sub-network and a regression evaluation sub-network. The technique classification sub-network outputs the recognition probability for 36 predefined ink painting brushwork techniques based on the deep temporal feature vector. The regression evaluation sub-network, based on the same deep temporal feature vector, generates a comprehensive quantitative score for brushwork techniques aligned with an expert scoring system through a fully connected layer mapping. (The expert scoring system consists of 5 national-level painters with over 10 years of ink painting teaching experience; the scoring dimensions and weights are: strength control 30%, speed and rhythm 25%, ink color control 25%, and technique standardization 20%; the alignment error threshold between the model output score and the expert score is ≤ ±3 points.)

[0009] Real-time feedback and correction guidance generation module: This module receives the recognition probability and quantitative score of ink painting brushwork techniques output by the brushwork technique recognition and quantitative evaluation module. Based on the recognition probability and quantitative score, and combined with the pre-set expert knowledge rule base and historical learning data, it generates real-time visual correction guidance information that includes specific deviation descriptions, standard technique demonstration animations, and targeted practice suggestions.

[0010] To optimize the above technical solution, the specific measures also include:

[0011] Furthermore, the multimodal data acquisition module includes: a high frame rate optical tracking unit, a pressure sensor and a high-speed microscopic imaging unit. The high frame rate optical tracking unit uses a binocular infrared high-speed camera to reconstruct the three-dimensional motion trajectory of the pen tip with sub-millimeter precision by capturing active infrared markers fixed on the pen tip.

[0012] The pressure sensor includes a miniature piezoresistive sensor array integrated below the pen grip area, used to measure the axial pressure on the pen and its distribution.

[0013] The high-speed microscopic imaging unit includes an industrial camera and a supplementary light source mounted perpendicular to the paper surface, which is focused on the brushstroke area to capture the dynamic process of ink diffusion in the Xuan paper fibers.

[0014] Furthermore, the timestamp alignment operation performed by the data fusion and preprocessing module employs a method combining hardware trigger signals and cubic spline interpolation algorithms to synchronize sensor data with different sampling rates onto the time axis of the highest sampling rate. The hardware trigger signal is issued by the main control unit that controls the operation of the three sensing units at the beginning of each acquisition cycle, ensuring that the start time of all sensor data acquisition is synchronized. For data points that are not one-to-one due to different sensor sampling rates, a cubic spline interpolation algorithm is used to interpolate the low-frequency data sequence onto the time axis of the highest sampling rate.

[0015] Furthermore, the pen stroke trajectory features include the three-dimensional coordinates, velocity components, acceleration components, curvature, inflection point density, and pause interval duration of the pen tip trajectory; the pressure dynamic features include the mean, variance, short-time energy, zero-crossing rate, and correlation coefficient between the pressure signal and the trajectory velocity; the ink color morphology feature extraction includes threshold segmentation and contour extraction of each frame of microscopic image (using the Otsu threshold segmentation algorithm with an adaptive window size of 5×5 pixels (the window is dynamically adjusted according to the ink area)), calculating the equivalent diameter, circularity, and grayscale histogram entropy of the ink area (using a 3×3 pixel local window, with the entropy value normalized to the [0,10] interval).

[0016] Furthermore, in the pen stroke depth feature extraction module based on the temporal convolutional network, the dilated causal temporal convolutional network contains multiple dilated convolutional blocks, with the dilation rate of each block increasing exponentially. Each dilated convolutional layer is followed by a batch normalization layer, a ReLU activation function layer, and a Dropout layer. The network finally aggregates the temporal dimension through a global average pooling layer, outputting a fixed-length deep temporal feature vector. The calculation formula for the dilated causal convolutional network is as follows:

[0017]

[0018] Where t is the time step, For each time step, x is the output, and x is the input sequence. Here, K represents the kernel weights, k represents the kernel size, k represents the output channels, and d represents the dilation rate of the layer. , where l is the number of layers.

[0019] Furthermore, in the brushstroke technique recognition and quantitative evaluation module, the technique classification sub-network and the regression evaluation sub-network are trained based on the cross-entropy loss function and the mean squared error loss function, respectively, and share the deep temporal feature vector extracted by the dilated causal temporal convolutional network. The technique classification sub-network adopts a fully connected network with two hidden layers and finally outputs a 36-dimensional probability vector through the Softmax function; the regression evaluation sub-network adopts a fully connected network with one hidden layer and directly outputs a scalar score, with the alignment error threshold between the output score and the expert score ≤ ±3 points.

[0020] Furthermore, the real-time feedback and correction guidance generation module has a built-in expert knowledge rule base that encodes the diagnostic logic of various technique errors in a condition-conclusion format. It matches the quantitative score and the identified technique type with the expert knowledge rule base, and combines the comparison results of the dimension in the deep feature vector with its corresponding threshold to diagnose the specific deviations in pen strokes in terms of force control, speed rhythm or angle transformation.

[0021] Furthermore, the real-time feedback and correction guidance generation module also integrates a learner model, which is used to record the technique scoring sequence and common error types in an individual's historical practice. When generating real-time guidance, it recommends practice content that best matches the weak points in the learner model.

[0022] This invention also provides an evaluation method for a deep learning evaluation system of ink painting brushwork techniques, comprising the following steps:

[0023] The multimodal data acquisition module synchronously acquires time-series images of the pen tip's three-dimensional spatial motion, pen pressure, and ink diffusion in the pen stroke area. The data fusion and preprocessing module performs time synchronization and feature extraction on the acquired time-series images, converting the pen tip's three-dimensional spatial coordinate time-series into pen trajectory features including positional three-dimensional coordinates, velocity components, acceleration components, curvature, inflection point density, and pause interval duration; converting the pen pressure time-series into pressure dynamic features including mean, variance, short-time energy, zero-crossing rate, and the correlation coefficient between pressure signal and trajectory velocity; and converting the ink diffusion image time-series into ink morphology features including equivalent diameter, roundness, and grayscale histogram entropy. These three features are then unified along a time dimension to form an initial multimodal feature matrix. This initial multimodal feature matrix is ​​input into a pre-trained dilated causal temporal convolutional network in the pen stroke depth feature extraction module. This network uses dilated convolutional layers with exponentially increasing dilation rates for multi-level abstraction, outputting a deeply fused temporal feature vector. The deep fusion temporal feature vectors are input in parallel into the technique classification subnetwork and evaluation regression subnetwork in the brushstroke technique recognition and quantitative evaluation module, respectively, to obtain the probability distribution vector and comprehensive brushstroke technique quantitative score, with the score range between 1 and 100. Based on the obtained probability distribution vector and brushstroke technique quantitative score, the real-time feedback and correction guidance generation module combines the obtained probability distribution vector and brushstroke technique quantitative score with the pre-set expert evaluation rules and historical learning data comparison analysis to generate and present real-time visual correction guidance information that points out specific deviations, provides standard demonstrations and practice suggestions.

[0024] Furthermore, the dilated causal temporal convolutional network contains five dilated convolutional blocks with dilation rates of 1, 2, 4, 8, and 16, respectively, to capture multi-scale temporal dependencies ranging from millisecond-level micro-movements of the pen tip to second-level pen stroke rhythms.

[0025] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the above-mentioned deep learning evaluation method for ink painting brushwork techniques.

[0026] The present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the above-described deep learning evaluation method for ink painting brushwork techniques.

[0027] The beneficial effects of this invention are as follows: This invention achieves holographic digitization and coupled analysis of the "form, force, and ink" of brushstrokes through multimodal perception; it utilizes a multi-layer dilated causal temporal convolutional network for deep mining, efficiently capturing multi-scale temporal dependencies ranging from millisecond-level micro-movements to second-level rhythms with a small number of parameters, resulting in high recognition and evaluation accuracy; the constructed parallel dual-branch evaluation architecture with shared deep feature input simultaneously achieves high-precision technique classification and continuous quantitative scoring, and by aligning with expert scoring, it endows machine evaluation with artistic authority; it provides real-time, interpretable, and personalized feedback and guidance, greatly improving learning efficiency and experience, and breaking through the bottlenecks of traditional teaching. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the structure of the deep learning evaluation system for ink painting brushwork techniques of the present invention;

[0029] Figure 2 This is a schematic diagram of the network architecture of the deep learning evaluation system for brushwork techniques in ink painting based on temporal convolutional networks.

[0030] Figure 3 This is a flowchart illustrating the deep learning evaluation method for ink painting brushwork techniques of the present invention. Detailed Implementation

[0031] The invention will now be described in further detail with reference to the accompanying drawings.

[0032] The embodiments described in this invention are merely some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0033] Example 1: In a Chinese painting classroom or individual practice setting, learners use a specially designed intelligent ink brush to practice brushwork techniques on standard raw Xuan paper. The intelligent ink brush is equipped with the system described in this invention.

[0034] like Figure 1 As shown, this invention discloses a deep learning evaluation system for ink painting brushwork techniques. The system includes a multimodal data acquisition module, a data fusion and preprocessing module, a brushwork deep feature extraction module based on temporal convolutional networks, a brushwork technique recognition and quantitative evaluation module, and a real-time feedback and correction guidance generation module. The entire system's operation flow is as follows: Figure 3 As shown.

[0035] The multimodal data acquisition module is the front end of the system for acquiring raw physical signals. In one implementation, this module consists of three highly integrated sensing units, ensuring the simultaneous capture of key information in three dimensions during pen movement. A high-frame-rate optical tracking unit is used to acquire the time-series sequence of the pen tip's three-dimensional spatial coordinates. It employs a pair of precisely calibrated binocular infrared high-speed cameras with an operating frequency set to 1000 Hz. By capturing two pre-fixed active infrared emitting markers on the pen tip, and utilizing stereo vision principles, the system reconstructs the pen tip's trajectory in the three-dimensional coordinate system in real time with sub-millimeter precision. The raw data output at each moment includes the X, Y, and Z coordinates of the pen tip markers in the coordinate system, as well as the velocity and acceleration vectors calculated from coordinate differences.

[0036] The pressure sensor unit, embedded beneath the grip area of ​​the smart pen, is used to collect the timing sequence of pressure applied to the pen. It employs a miniature piezoresistive sensor array. This array, sampling at 500 Hz, simultaneously measures the axial pressure on the pen and senses the distribution differences of pressure along the circumferential direction of the pen's cross-section. The pressure distribution data is crucial for indirectly inferring the contact area between the pen tip and the Xuan paper, and whether the pen tip is clustered or dispersed.

[0037] The high-speed microscopic imaging unit is used to acquire a time-series image of ink diffusion in the area where the brushstrokes touch the paper. Mounted perpendicular to the painting plane, it includes an industrial-grade global shutter CMOS camera equipped with a coaxial ring-shaped LED shadowless illumination source. The camera has a resolution of 1920×1080 pixels and a frame rate of 120 frames per second. Its optical lens is specially calibrated to always focus on a fixed area of ​​approximately 2 cm by 2 cm around the point of contact between the brush tip and the paper surface, continuously capturing a complete dynamic image sequence of ink penetration, diffusion, and spreading within the Xuan paper fibers until initial drying.

[0038] The acquisition actions of the three sensing units are strictly synchronized by a central control unit through hardware trigger signals to ensure that the start time of each acquisition cycle is perfectly aligned, laying the foundation for subsequent multi-source data fusion.

[0039] The data fusion and preprocessing module is responsible for receiving the raw asynchronous data streams from the three aforementioned sensing units and transforming them into a unified, well-organized, and information-rich initial feature matrix. In one implementation, this module performs timestamp alignment and data interpolation synchronization operations on the pen tip 3D spatial coordinate time series, pen barrel pressure time series, and ink diffusion image time series. Since the acquisition start time is guaranteed by a hardware trigger signal, but due to the different inherent sampling rates of each sensor, the number of data points generated within the same time length is not consistent, a method combining hardware trigger signals and software interpolation is adopted. The module uses a cubic spline interpolation algorithm to interpolate the pen barrel pressure data series and ink diffusion image series with lower sampling rates onto the pen tip trajectory data time axis with the highest sampling rate, i.e., a 1000 Hz time reference. After completing time synchronization, the module executes the three feature extraction tasks in parallel.

[0040] For the three-dimensional spatial coordinate temporal sequence of the pen tip, the velocity components of the pen tip in the X, Y, and Z directions are first calculated using first-order difference, and then the acceleration components are calculated using second-order difference. Combining the coordinates and velocity vectors, the curvature of the pen tip's trajectory, the frequency density of trajectory inflection points (inflection point density), and the duration of pauses where the velocity approaches zero are further calculated. These characteristics reflect the smoothness of pen movement, the strength of transitions, and rhythmic control.

[0041] Specifically, the trajectory characteristics of the pen tip movement are calculated as follows: The velocity component is calculated as follows: Let the sampling frequency of the multimodal data acquisition module be... (In this invention) =1000Hz, corresponding to the sampling time interval =1 / =1ms); the time sequence of the pen tip's coordinates in three-dimensional space is P(t)=(X(t),Y(t),Z(t)), where t is the time step index (t=0, 1, 2, ..., n-1, n is the total number of sampling points), and X(t), Y(t), and Z(t) are the spatial coordinates of the pen tip in the X, Y, and Z directions at the t-th time step (unit: mm). The velocity components in each direction are calculated using the first-order forward difference method, as shown in the following formula:

[0042] Velocity in the X direction: (t)= Where t = 0, 1, 2, ..., n-2

[0043] Velocity in the Y direction: (t)= Where t = 0, 1, 2, ..., n-2

[0044] Z-direction velocity: (t)= Where t = 0, 1, 2, ..., n-2

[0045] For the last time step t=n-1, the backward difference method is used to supplement the calculation: = , , The calculation method is the same. The unit for the velocity component is mm / s.

[0046] Acceleration component calculation method: Based on the velocity component sequences in each direction obtained from the above calculations. (t), (t), (t), the acceleration components in each direction are calculated using the second-order forward difference method, as shown in the following formula:

[0047] X-direction acceleration: (t)= Where t = 0, 1, 2, ..., n-3

[0048] Acceleration in the Y direction: (t)= Where t = 0, 1, 2, ..., n-3

[0049] Z-direction acceleration: (t)= Where t = 0, 1, 2, ..., n-3

[0050] For the final time step t=n-2, the backward difference method is used to supplement the calculation: (n-2)= , t=n-1, (n-1)= (An extension of the second-order central difference, using the first three velocity points to fit the final acceleration). (t), The calculation method for the final time step (t) is similar. The unit for acceleration components is mm / s. .

[0051] The method for calculating the curvature of the pen tip's trajectory, based on the derivation and calculation of the curvature formula for three-dimensional spatial curves, involves the following steps:

[0052] 1. Calculate the tangent vector of the trajectory. The tangent vector is composed of velocity components in each direction. =( (t), (t), (t)), its modulus | ∣= (i.e., the speed of the pen tip);

[0053] 2. Calculate the normal vector of the trajectory. Tangent vector Taking the time derivative yields =( (t), (t), (t) (i.e., acceleration vector), will Decomposed into tangential acceleration and normal acceleration, normal vector for Perpendicular to The directional component, its modulus = ,in It is the dot product of the tangent vector and its derivative;

[0054] 3. Calculate the curvature k(t): According to the three-dimensional curve curvature formula k(t) = The unit is 1 / mm.

[0055] Method for calculating the frequency density of trajectory inflection points (inflection point density): 1. Calculate the first-order difference of curvature. k(t): k(t) = k(t+1) − k(t), where t = 0, 1, 2, ..., n-2; 2. Identify inflection points: When the time step t+1 is determined to be the inflection point of the trajectory (i.e., the point of sudden curvature change); if If k(t)=0, then proceed with the judgment. k(t+1) and 3. Count the total number of inflection points. 4. Calculate the frequency density of inflection points: Traverse all time steps and count the total number of inflection points that satisfy the above conditions; : = Where L is the total length of the pen tip's trajectory (unit: mm), the total trajectory length L = Frequency density is measured in units per mm.

[0056] Method for calculating the duration of pauses with near-zero speed: The pause interval refers to the combined speed of the pen tip. Continuously below the preset threshold (In this invention) =0.5mm / s (based on the industry common sense of "pause" in ink painting brushstrokes) is set within a continuous time interval. The calculation steps are as follows: 1. Set the speed threshold. By statistically analyzing the pause characteristics of 100 sets of standard penmanship samples, the following was determined: =0.5mm / s, below this threshold is judged as "approximately stationary" state; 2. Identify pause intervals: Traverse the combined velocity sequence When the resultant velocity of m consecutive time steps all satisfy ≤ When a pause interval is defined, it is considered a pause interval; where m≥3 (i.e., duration≥3ms, to avoid misjudgment due to sensor noise); 3. Calculate the duration of a single pause interval: Let the starting time step of a pause interval be... The end time step is Then the duration of that interval =( - +1)× 4. Total pause duration: The sum of the durations of all pause intervals is the total duration of the pause intervals where the speed approaches zero. = , where M is the total number of identified pause intervals, and the duration is in ms.

[0057] For the pen pressure time series, the extracted features include the short-time energy of the pressure signal, the zero-crossing rate, the mean and variance of the pressure, and the Pearson correlation coefficient between the pressure signal and the synthesis speed of the pen tip trajectory. This coefficient is a key indicator for judging whether the pressure is "penetrating" or "superficial and weak".

[0058] The calculation method is as follows: Short-time energy calculation method for pressure signals: 1. Set the short-time analysis window: Use a length of... Sliding window (in this invention) =50, corresponding to a time length of 50× =50ms, adapting to the dynamic change frequency of pen pressure), with an overlap rate of 50% between windows; 2. Frame the pressure timing sequence: Let the pen pressure timing sequence be... (t=0, 1, ..., n-1, (where t is the pressure value at time step t, in N), and the pressure data for the i-th frame is... = ( × / 2+ ), where m = 0, 1, ..., -1, 3. Calculate the short-time energy of each frame: (This is the initial step in the text, but the full context is = 4. Generate short-time energy feature sequences: generate short-time energy feature sequences for all frames. By splicing the frames sequentially, short-time energy features aligned with the time dimension of the original pressure sequence are obtained.

[0059] The zero-crossing rate calculation method for pressure signals: The zero-crossing rate reflects the number of times the pressure signal crosses the "reference pressure" within a local window (reflecting the "fluctuation" of the writing force). Calculation steps: 1. Set the reference pressure: Take the mean of the pressure sequence. = As a baseline (i.e., a "stress-free" reference value); the stress sequence is framed: using the same sliding window as the short-time energy ( =50, overlap rate 50%), to obtain the pressure data of the i-th frame. 3. Calculate the number of zero-crossings in each frame: count the number of frames that satisfy the condition. Number of times < 0 (i.e., the pressure value crosses from below the baseline to above it, or vice versa); 4. Generate a zero-crossing rate feature sequence: the zero-crossing rate of each frame is... = (Normalized to window length), the zero-crossing rate feature is obtained by splicing the frames in order.

[0060] Methods for calculating the mean and variance of pressure: The mean and variance of pressure reflect the "overall level" and "stability" of force during the stroke. The calculation steps are as follows: 1. Mean pressure 2. Pressure variance .

[0061] Method for calculating the Pearson correlation coefficient between pressure signal and pen tip trajectory synthesis speed: Determine the synchronization sequence: Take the pressure sequence Synthesized speed sequence with pen tip Ensure that the time steps of both are perfectly aligned (through synchronization processing by the data fusion module); calculate the covariance. : = ,in = The mean of the synthesis rate; calculate the Pearson correlation coefficient. : = ,in = The standard deviation of the synthesis rate; The value range is [-1, 1]: >0.7: Force and speed are positively correlated (the force of the brushstroke increases with the speed, which conforms to the standard of "the force penetrates the back of the paper"). <0.3: Weak or negative correlation between force and speed (the force and speed of the brushstroke are disconnected, manifested as "superficial and weak").

[0062] For the temporal sequence of ink diffusion images, each frame is first subjected to adaptive threshold segmentation (Otsu) based on the Otsu algorithm (maximum inter-class variance method) to separate the ink area from the Xuan paper background. Then, the contour of the ink area is extracted, and the equivalent diameter, circularity, and gray-level histogram entropy of the pixels within the area are calculated. The equivalent diameter and circularity describe the macroscopic shape of the ink, while the gray-level entropy shows the uniformity and layering of the ink diffusion from the center to the edge. The algorithm parameters for ink feature extraction have been tested on 50 sets of samples with different Xuan paper and different ink concentrations, demonstrating good adaptability. Finally, all extracted trajectory features, pressure features, and ink features are concatenated in chronological order to form an initial multimodal feature matrix of dimension F multiplied by T, where F is the total number of features and T is the total number of time steps.

[0063] The feature extraction methods are as follows: 1. Adaptive threshold segmentation and contour extraction based on Otsu's algorithm:

[0064] The Otsu algorithm automatically determines the optimal segmentation threshold by maximizing the inter-class variance between the "ink blot region" and the "rice paper background region". The steps are: Image grayscale conversion – converting each frame of the color microscopic image (RGB channels) to a grayscale image with grayscale values ​​ranging from [0, 255] (0 for black, 255 for white); Calculating the inter-class variance: iterating through all possible grayscale thresholds k ∈ [0, 255], dividing the image pixels into two classes: Class 1 (ink blot region): pixels with grayscale values ​​≤ k, accounting for [percentage missing]. The mean is Category 2 (Xuan paper background): Pixels with grayscale values ​​> k, accounting for [percentage missing]. =1- The mean is The variance between classes is: (k)= Determine the optimal threshold: Select a threshold that makes... (k) largest As a segmentation threshold, the grayscale image is binarized (pixel grayscale ≤ Set to 1, corresponding to ink marks; pixel grayscale > (Set to 0, corresponding to the background); Adaptive window adjustment: For images with different Xuan paper / ink concentrations, if the ink area has "breaks" or "over-connectivity" after binarization, the image is automatically divided into 5×5 pixel local windows, Otsu segmentation is performed independently on each window, and then the binary results of all windows are merged (to ensure segmentation accuracy in complex scenes).

[0065] Ink Blot Contour Extraction: Based on the binarized image, the outer contour of the ink blot region (reflecting the boundary shape of the ink blot) is extracted. The steps are: Connected Component Labeling – Using 8-neighbor connectivity (adjacent pixels sharing edges or corners are considered the same region), all connected components in the binary image are labeled, and the connected component with the largest area (i.e., the effective ink blot region, excluding noise points) is selected; Contour Extraction: The Sobel edge detection operator is used to extract the edges of the labeled ink blot region, obtaining the closed outer contour of the ink blot (in pixel coordinate sequence {( ),( ), ...( )}express).

[0066] Method for calculating the equivalent diameter of the ink blot region: The equivalent diameter refers to "the diameter of a circle with the same area as the ink blot region," used to describe the macroscopic size of the ink blot. The area A of the ink blot region is calculated by: counting the total number of pixels (pixel value = 1) in the binary image of the ink blot region, multiplying by the actual area corresponding to a single pixel (in this invention, the pixel resolution of the microscopic image is pixels, so the area of ​​a single pixel is 0.0001). );

[0067] Calculate the equivalent diameter : =2× The unit is mm, and the larger the value, the wider the ink spread.

[0068] Method for calculating the roundness of an ink blot area: Roundness refers to "the degree to which the ink blot area resembles a circle," used to describe the regularity of the ink blot's shape. The calculation of the contour perimeter L of the ink blot area is as follows: For the contour pixel coordinate sequence {( ),( ), ...( )}, calculate the Euclidean distance between adjacent pixels in sequence and sum them, that is: L= (The last pixel is connected to the first pixel to close the contour), unit is mm; calculate the roundness C: C= The value of C ranges from [0,1]; C=1: the ink area is a perfect circle; the closer C is to 0: the more irregular the shape of the ink area (such as ink marks in the rubbing technique).

[0069] Method for calculating the grayscale histogram entropy of the ink blot region: The grayscale histogram entropy reflects the uniformity of ink distribution within the ink blot region (the higher the entropy value, the richer the ink color levels). Extracting grayscale values ​​of the ink blot region: From the original grayscale image, extract the grayscale values ​​of all pixels belonging to the ink blot region (pixel = 1 after binarization), obtaining the grayscale sequence G = {g1, g2, ..., gk} (k is the number of pixels in the ink blot region);

[0070] Calculate the grayscale histogram: Count the frequency of each grayscale value (0-255) in G. = (count(g=i) is the number of times grayscale value i appears);

[0071] Calculate the grayscale histogram entropy H: H = - (like =0, then this term is not included in the summation), the value range of H is [0,8] (corresponding to the maximum entropy of 256 gray levels), H>6: the ink color distribution is uneven and the sense of layering is strong (such as ink splashing and ink accumulation techniques); H<3: the ink color distribution is uniform and the sense of layering is weak (such as flat wash techniques).

[0072] In one implementation, the network architecture principle of the pen stroke depth feature extraction module based on a temporal convolutional network is described in [link to relevant documentation]. Figure 2 This module receives the initial multimodal feature matrix output from the preprocessing module as input. The main network is a five-layer dilated causal temporal convolutional network. The first layer is a standard one-dimensional convolutional layer with the number of input channels equal to the dimension F of the initial feature matrix, the number of output channels set to 64, a dilation rate of 1, and a kernel size of 3. This layer is responsible for capturing the local correlation between features at adjacent time points. The dilation rate of each subsequent layer increases exponentially, successively to 2, 4, 8, and 16, allowing the receptive field to expand exponentially with the number of layers while only increasing the number of parameters and computational cost. Specifically, the receptive field of the fifth layer convolutional neuron can cover feature relationships hundreds of time steps apart in the input sequence, enabling the modeling of long-range dependencies from millisecond-level pen stroke tremors to second-level complete pen stroke rhythms. The dilated causal convolutional network is constructed using the formula... The calculation shows that, where t is the time step, For each time step, x represents the input sequence, and k represents the output channel. Here, K represents the kernel weights, K is the kernel size (constantly 3 in this embodiment), and d is the dilation rate of the layer. Here, l represents the layer number. Each dilated convolutional layer is followed sequentially by a batch normalization layer, a ReLU (Reduced Activation Unit) layer, and a Dropout layer with a random deactivation rate of 0.2. The batch normalization layer accelerates network training and improves stability, the ReLU layer introduces non-linearity, and the random deactivation layer effectively prevents the model from overfitting the training data. After five layers of progressive abstraction and transformation through dilated convolutional blocks, the network aggregates the time dimension T through a global average pooling layer, outputting a fixed-length 512-dimensional deep temporal feature vector. This vector highly condenses the information about the coupling and dynamic evolution of trajectory, pressure, and ink color during the brushstroke process.

[0073] In one implementation, the brushstroke technique recognition and quantitative evaluation module receives the aforementioned 512-dimensional deep temporal feature vector and simultaneously performs two tasks—technique classification and comprehensive scoring—through a parallel dual-branch network structure. The two branch networks share the same deep temporal feature vector as input but have independent trainable parameters and optimization objectives.

[0074] The technique classification subnetwork consists of two fully connected layers. The first fully connected layer maps the 512-dimensional input to 256 dimensions, and the second fully connected layer further maps it to 36 dimensions, corresponding to the 36 core ink painting brushwork techniques predefined by the system (central brushstroke, side brushstroke, reverse brushstroke, dragging brushstroke, dry brushstroke, moist brushstroke, raindrop texture stroke, hemp fiber texture stroke, folded ribbon texture stroke, unraveling rope texture stroke, ox hair texture stroke, rubbing brushstroke, round dot, square dot, horizontal dot, vertical dot, dyeing brushstroke, splashing ink, accumulating ink, breaking ink, burnt ink, light ink, dark ink, heavy ink, trembling brushstroke, lifting and pressing, turning, moving brushstroke, pausing brushstroke, returning brushstroke, exiting brushstroke, concealing brushstroke, following brushstroke, sweeping brushstroke, pecking brushstroke, lifting brushstroke). Finally, the Softmax activation function is used to convert the 36-dimensional output into a probability distribution vector. The value of each element in the vector represents the probability that the input brushwork sequence belongs to the corresponding technique, and the sum of all elements is 1. The subnetwork is optimized using the cross-entropy loss function during the training phase to make its output probability distribution as close as possible to the real technique labels annotated by experts.

[0075] The regression evaluation subnetwork directly maps 512-dimensional deep temporal features into a single scalar value through a fully connected layer. This scalar value is the comprehensive quantitative score for brushwork techniques, with the score range normalized to between 0 and 100. The training labels for this subnetwork are not manually set rules, but rather the average score given by multiple national-level artists to the same batch of brushwork samples. During the training phase, this subnetwork is optimized using a mean squared error loss function. Through extensive training on large datasets, the model's output score achieves mathematical alignment with the collective subjective evaluation standards of human experts (the alignment error threshold between the output score and the expert score is ≤ ±3 points), thus giving the machine's score artistic authority. Alignment means that the quantitative score output by the model meets the preset "error control standard" and "statistical consistency standard" between the actual scores labeled by the collective of human experts.

[0076] The error control standards are as follows: 1. The difference between the model's predicted value and the expert's true value should be ≤3 points in Mean Absolute Error (MAE). Formula: MAE = ,in, Let be the mean expert rating (true label) for the i-th sample. Output a score for the model, where N is the total number of samples. The root mean square error (RMSE) should be ≤ 4 points. The formula is: RMSE = Maximum permissible deviation ≤ 10 points. Statistical consistency criteria: The statistical distribution (mean, variance, quantiles) of all sample scores output by the model must match the expert score distribution: the difference between the expert score mean (e.g., the training set expert score mean of 72) and the model score mean ≤ 2 points; the difference between the expert score variance (e.g., expert score variance of 120) and the model score variance ≤ 10 points; quantile alignment: if the 75th quantile of the expert score is 80, the 75th quantile of the model score must be between 78 and 82. A Pearson correlation coefficient ≥ 0.9 means that the model score and the expert score are strongly linearly correlated. For samples that experts consider "excellent" (above 85 points) or "needs improvement" (below 60 points), the model will also provide a score within the corresponding range. Consistency of rating (accuracy ≥ 90%): The model is divided into 4 levels (Excellent [90, 100], Good [75, 90], Pass [60, 75], and Needs Improvement [0, 60]) based on the scores. The consistency rate between the model's rating of the samples and the expert rating is ≥ 90%: For samples rated "Excellent" by experts, the proportion of samples rated "Excellent" or "Good" by the model is ≥ 95%; For samples rated "Needs Improvement" by experts, the proportion of samples rated "Needs Improvement" or "Pass" by the model is ≥ 95%.

[0077] Alignment standard training: Training phase: Minimize the squared error between the model score and the expert score using the mean squared error loss function to ensure that MAE and RMSE meet the above quantitative standards; Validation phase: Use an independent test set (containing 1000+ expert-annotated samples) to validate all alignment metrics. Model training is considered complete only when all metrics meet the standards; Application phase: During real-time evaluation, if the deviation between the model score and the historical expert score for a sample exceeds 10 points, the system will automatically trigger "secondary verification" (correcting the score by combining the expert knowledge rule base) to ensure that the alignment standards are continuously met.

[0078] The real-time feedback and correction guidance generation module serves as the terminal for interaction between the system and the learner. In one implementation, this module presets multiple scoring threshold ranges, such as: Excellent [90, 100], Good [75, 90], Pass [60, 75], and Needs Improvement [0, 60]. Only when the comprehensive brushwork technique quantitative score falls into the "Needs Improvement" or "Pass" range will the system forcibly initiate the detailed correction guidance generation process. For results of "Excellent" and "Good," a simple affirmation can be given, or higher-level improvement suggestions can be provided.

[0079] This module incorporates a structured expert knowledge rule base. This base encodes the diagnostic logic of calligraphy and painting education experts for common errors in various techniques using an "IF-THEN" format. Many rules in the expert rule base include a scoring range as a prerequisite. For example, a typical rule states: "If the overall score is between 60 and 75 points and the probability of identifying it as the 'central stroke' technique is highest, and the pressure variance extracted from the deep feature vector is greater than the preset threshold P, while the trajectory curvature feature is less than the preset threshold C, then the current brushstroke is judged to have the error of 'unstable force leading to a slippery line, lacking roundness and three-dimensionality'." Each rule in the expert knowledge rule base is associated with a text description, a standard technique demonstration animation (or 3D reconstructed handwriting), and targeted practice statements. The real-time feedback and correction guidance generation module also integrates a learner model, which continuously records all scores and error types in the learner's historical practice. The threshold parameters of the expert knowledge rule base are determined through statistical analysis of 1000 sets of expert-annotated brushstroke samples, ensuring the scientific validity and accuracy of the diagnostic logic.

[0080] When the module detects that the overall score is lower than the preset passing threshold, or when a condition of an expert rule is triggered, it will start the guidance generation process.

[0081] Real-time feedback and correction guidance generation process: After receiving the real-time recognition probability and score, the system first determines whether the score is below the passing threshold (75 points in this embodiment). If it is below the passing threshold, it continues to match the real-time recognition probability with the rule base, combining the abnormal dimension in the deep temporal feature vector, to generate a text description that clearly points out the deviation. Subsequently, the system retrieves the corresponding correct pen-writing animation from the standard technique demonstration database. This animation can display the trajectory, rhythm, and force changes of the standard movement from multiple angles. Finally, combined with the built-in personalized learner model, the system prioritizes and recommends the most suitable targeted training program for the current weak point from the practice library. Among them, the learner model is a digital archive that records individual learning progress and supports personalized feedback, including four types of information:

[0082] Basic information: Learner ID, learning stage (beginner / intermediate / professional), feedback format preference (e.g., preference for text instructions / animated demonstrations / voice guidance);

[0083] Ability Profile: Mastery rating of 36 predefined techniques (1-100), weak points in core dimensions such as strength control / speed and rhythm, and trend of ability changes;

[0084] Learning behaviors: Technique types, overall scores, error types and frequencies, practice duration and frequency in historical practice;

[0085] Feedback adaptation: Prioritization of weaknesses (calculated by "error frequency × room for improvement", such as "the dry brush technique has the highest error frequency (15 times) and the lowest current mastery (62 points), so it has the highest priority; the central brush technique has the highest error frequency but the lowest mastery (85 points), so it has the lowest priority"), configuration of personalized feedback formats (determine the presentation method of guidance according to the learning stage and preference settings, such as "beginner learners are given priority to 'standard animation + step-by-step text instructions', while professional learners are given 'expert demonstration comparison video + advanced practice suggestions'"), and weighting of targeted practice recommendations (the recommended number of practice sessions and practice time for each technique are dynamically adjusted based on the degree of weakness, such as "the dry brush technique is recommended to practice 2 sets (5 times per set) every day for 1 week; the central brush technique is recommended to practice 1 set (3 times per set) every week to maintain the level").

[0086] All text, animations, and suggestions are integrated into a real-time feedback report, presented to learners via a teaching tablet or augmented reality glasses, completing a full intelligent teaching loop from "perception-analysis-evaluation-feedback." If the score is not lower than the passing threshold, a brief positive feedback is generated directly.

[0087] The expert knowledge rule base contains at least 5 typical diagnostic rules (examples are shown below):

[0088] Rule 1: If the overall score is 60-75 points, the probability of identifying it as a "center forward" technique is ≥0.8, the pressure variance is >0.3 (after normalization), and the trajectory curvature is <0.5, then it is judged as "force control deviation: unstable force leads to slippery lines, lacking roundness and three-dimensionality". The corresponding practice suggestion is "to use 'suspended wrist straight line practice', keep the pressure variance ≤0.2, practice 10 sets per day, 5 straight lines per set".

[0089] Rule 2: If the overall score is <60 points, the probability of being identified as "dry brush" technique is ≥0.7, the ink area growth rate is <0.1mm² / s, and the grayscale histogram entropy is <5.0, then it is judged as "Ink color control deviation: insufficient ink and uneven diffusion, dry and layered brushstrokes". The corresponding practice suggestion is "adjust the ink to water ratio to 1:2, adopt 'dotting and overlay practice', control the amount of ink applied to each stroke, practice 8 sets per day, 3 brushstrokes per set".

[0090] Rule 3: If the overall score is 60-75, the probability of identifying it as "hemp fiber texture" technique is ≥0.75, the density of trajectory inflection points is >0.8 per centimeter, and the speed standard deviation is >2cm / s, then it is judged as "speed rhythm deviation: the brushstrokes are too fast and the transitions are abrupt, and the texture is messy and disorderly". The corresponding practice suggestion is "to use 'slow brush uniform speed practice', control the brushstroke speed to ≤1cm / s, pause for 0.2 seconds at the inflection points, practice 6 sets per day, 2 textures per set".

[0091] Rule 4: If the overall score is 75-90 points, the probability of being identified as "splashing ink" technique is ≥0.8, the average pressure is <0.4 (after normalization), and the equivalent diameter of the ink mark is <3cm, then it is judged as "deviation of strength and range: insufficient force leads to a small splashing range and lack of momentum". The corresponding practice suggestion is "to adopt 'brush swing force practice' to enhance wrist strength, control the average pressure to ≥0.6, practice 5 sets a day, and splash ink 2 times per set".

[0092] Rule 5: If the overall score is <60 points, the probability of being identified as a "return stroke" technique is ≥0.7, the pause interval duration is >0.5 seconds, and the absolute value of acceleration is >5cm / s², then it is judged as "deviation in movement continuity: the pause during the return stroke is too long and the acceleration is too sudden, resulting in disjointed lines". The corresponding practice suggestion is to "use 'continuous return stroke practice', shorten the pause time to ≤0.1 seconds, smoothly control the acceleration to ≤2cm / s², practice 10 sets per day, with 3 return strokes per set".

[0093] Example 2, as Figure 3 As shown, the evaluation method of the deep learning evaluation system for ink painting brushwork techniques disclosed in this invention includes the following steps:

[0094] Step S1: Through a multimodal sensor array deployed on the smart ink painting brush and painting platform, the original sensor data stream of the learner creating ink painting on rice paper is collected synchronously. The original sensor data stream includes a time sequence of the spatial motion of the brush tip obtained at a sampling rate of 1000Hz, a time sequence of the axial pressure of the brush handle obtained at a sampling rate of 500Hz, and a time sequence of the ink diffusion image of the brush stroke area obtained at a frame rate of 120fps.

[0095] Step S2 involves time synchronization and initial feature extraction of the aforementioned time series. The pen tip spatial motion data is converted into pen trajectory features including three-dimensional position coordinates, velocity components, acceleration components, curvature, inflection point density, and pause interval duration. The pen shaft axial pressure data is converted into pressure dynamic features including mean, variance, short-time energy, zero-crossing rate, and correlation coefficient between pressure signal and trajectory velocity. The ink diffusion image time series is converted into ink morphology features including equivalent diameter, roundness, and grayscale histogram entropy. All features are aligned on a unified time axis to form an initial multimodal feature matrix.

[0096] Step S3: The initial multimodal feature matrix is ​​input into a pre-constructed and trained dilated causal temporal convolutional network. The network contains 5 dilated convolutional blocks, with dilation rates of 1, 2, 4, 8, and 16 for each block, and a kernel size of 3. Through the forward propagation of the network layer by layer, the input features are subjected to nonlinear transformation and high-level abstraction, and finally converged into a 512-dimensional deep fusion temporal feature vector at the last layer of the network.

[0097] Step S4: The deep fusion temporal feature vector is input in parallel to two branch networks. The first branch is a technique classification sub-network trained based on the cross-entropy loss function, which outputs the probability distribution vector of 36 techniques such as central stroke, side stroke, dry brush, moist brush, and texturing. The second branch is an evaluation regression sub-network trained based on the mean square error loss function, which outputs a comprehensive technique quantitative score ranging from 0 to 100.

[0098] Step S5: Based on the technique probability distribution vector and comprehensive technique quantitative score, combined with the preset expert evaluation rules and historical learning data comparison and analysis, a technique execution feedback report with pictures and text is generated and presented in real time. The report clearly points out the specific deviations in the current pen stroke in terms of strength control, speed rhythm or angle transformation, and provides dynamic standard pen stroke demonstrations and step-by-step practice suggestions.

[0099] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the above-mentioned deep learning evaluation method for ink painting brushwork techniques.

[0100] The present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the above-described deep learning evaluation method for ink painting brushwork techniques.

[0101] The above embodiments demonstrate that the system can objectively, precisely, and in real time evaluate whether the learner's penmanship conforms to traditional techniques and standards, and can provide specific and actionable improvement suggestions to overcome the limitations of subjective evaluation, difficulty in quantification, and delayed feedback in the traditional "master-apprentice" model.

[0102] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A deep learning evaluation system for ink painting brushwork techniques, characterized in that, include: The multimodal data acquisition module is used to synchronously acquire multimodal temporal data during the brushstroke process of ink painting. The multimodal temporal data includes the time sequence of the three-dimensional spatial coordinates of the brush tip, the time sequence of the pressure of the brush handle, and the time sequence of the ink diffusion image of the brushstroke area. The data fusion and preprocessing module is used to perform time synchronization and feature extraction on the multimodal time series data to generate a multimodal feature matrix; The pen technique deep feature extraction module based on temporal convolutional network is used to receive the multimodal feature matrix and output a deep temporal feature vector through dilated causal temporal convolutional network. The penmanship technique recognition and quantitative evaluation module includes a technique classification subnetwork and a regression evaluation subnetwork, which are used to output the recognition probability and quantitative score of the penmanship technique based on the deep temporal feature vector. The real-time feedback and correction guidance generation module is used to generate real-time visual correction guidance based on the recognition probability and quantitative score.

2. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, The multimodal data acquisition module includes: a high frame rate optical tracking unit, a pressure sensor and a high-speed microscopic imaging unit. The high frame rate optical tracking unit uses a binocular infrared high-speed camera to reconstruct the three-dimensional motion trajectory of the pen tip with sub-millimeter precision by capturing active infrared markers fixed on the pen tip. The pressure sensor includes a miniature piezoresistive sensor array integrated below the pen grip area, used to measure the axial pressure on the pen and its distribution. The high-speed microscopic imaging unit includes an industrial camera and a supplementary light source mounted perpendicular to the paper surface, which is focused on the brushstroke area to capture the dynamic process of ink diffusion in the Xuan paper fibers.

3. The deep learning evaluation system for ink painting brushwork techniques according to claim 2, characterized in that, When performing timestamp alignment, the data fusion and preprocessing module uses a combination of hardware trigger signals and cubic spline interpolation algorithms to synchronize sensor data with different sampling rates to the time axis with the highest sampling rate. The hardware trigger signal is issued by the main control unit that controls the operation of the three sensing units at the beginning of each acquisition cycle to ensure that the start time of all sensor data acquisition is synchronized. For data points that are not one-to-one due to different sensor sampling rates, a cubic spline interpolation algorithm is used to interpolate the low-frequency data sequence onto the time axis with the highest sampling rate.

4. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, The data fusion and preprocessing module extracts pen movement trajectory features, pressure dynamic features, and ink color morphology features from the multimodal time-series data. The pen movement trajectory features include the velocity component, acceleration component, curvature, inflection point density, and pause interval duration of the pen tip movement trajectory. The pressure dynamic features include the mean, variance, short-time energy, zero-crossing rate, and correlation coefficient between the pressure signal and the trajectory velocity. The ink color morphology features include the equivalent diameter, roundness, and grayscale histogram entropy of the ink stain region calculated by threshold segmentation and contour extraction of each frame of microscopic image.

5. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, In the pen stroke depth feature extraction module based on the temporal convolutional network, the dilated causal temporal convolutional network contains multiple dilated convolutional blocks, with the dilation rate of each block increasing exponentially. Each dilated convolutional layer is followed by a batch normalization layer, a ReLU activation function layer, and a Dropout layer. The calculation formula for the dilated causal convolutional network is as follows: Where t is the time step, For each time step, x is the output, and x is the input sequence. Here, K represents the kernel weights, k represents the kernel size, k represents the output channels, and d represents the dilation rate of the layer. , where l is the number of layers.

6. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, In the brushstroke technique recognition and quantitative evaluation module, the technique classification subnetwork and the regression evaluation subnetwork are trained based on the cross-entropy loss function and the mean square error loss function, respectively, and share the deep temporal feature vector extracted by the dilated causal temporal convolutional network. The alignment error threshold between the output score and the expert score is ≤ ±3 points.

7. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, The real-time feedback and correction guidance generation module has a built-in expert knowledge rule base, which encodes the diagnostic logic of various technique errors in a condition-conclusion format. By matching the technique features corresponding to the recognition probability and the quantitative score with the expert knowledge rule base, and combining the comparison results of each dimension in the deep temporal feature vector with the corresponding threshold, deviations in pen strokes in terms of force control, speed rhythm or angle transformation are diagnosed.

8. The deep learning evaluation system for ink painting brushwork techniques according to claim 1, characterized in that, The real-time feedback and correction guidance generation module also integrates a learner model, which records the technique scoring sequence and common error types in an individual's historical practice. When generating real-time guidance, it recommends practice content that best matches the weak points in the learner model.

9. An evaluation method based on the deep learning evaluation system for ink painting brushwork techniques according to any one of claims 1 to 8, characterized in that, Includes the following steps: Simultaneously collect multimodal temporal data during the brushstroke process of ink painting; The multimodal time-series data is time-synchronized and feature-extracted to generate a multimodal feature matrix; The multimodal feature matrix is ​​input into the dilated causal temporal convolutional network, which outputs a deep temporal feature vector. The deep temporal feature vectors are input in parallel into the technique classification subnetwork and the regression evaluation subnetwork to obtain the recognition probability and quantitative score of the brushstroke technique, respectively. Based on the identification probability and quantitative score, combined with expert evaluation rules and historical learning data, real-time visual correction guidance is generated and presented.

10. The evaluation method according to claim 9, characterized in that, The dilated causal temporal convolutional network contains five dilated convolutional blocks with dilation rates of 1, 2, 4, 8, and 16, respectively, used to capture multi-scale temporal dependencies ranging from millisecond-level micro-movements of the pen tip to second-level pen stroke rhythms.