Transfomer-based calligraphy robot control method and system

Through the Transfomer-based font image generation model and quadratic spline curve interpolation algorithm, the shortcomings of calligraphy robots in writing effects and style simulation are solved, high-quality calligraphy robot control is achieved, and writing accuracy and artistic expression are improved.

CN120612701APending Publication Date: 2025-09-09XIAMEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510686413.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing calligraphy robots still have room for improvement in writing effects, artistic expression and intelligence. It is difficult to precisely control calligraphy strokes and simulate the styles of different calligraphers. Font generation lacks artistic sense and diversity, and mapping font outlines to robot trajectories is difficult.

Method used

A Transformer-based font image generation model was adopted. A dataset was constructed by collecting and preprocessing historical calligraphy images. A dual-channel encoder and output layer were combined, and a loss function was set for training. Real-time writing control instructions were generated, and the outline data was converted using the GetGlyphOutline function and the quadratic spline curve interpolation algorithm to control the robotic arm to perform writing actions.

Benefits of technology

It effectively improves the writing quality of calligraphy robots, enhances artistic sense and diversity, realizes precise calligraphy stroke control and accurate mapping of font outlines to robot trajectories, and simulates the styles of different calligraphers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612701A_ABST
    Figure CN120612701A_ABST
Patent Text Reader

Abstract

The invention provides a calligraphy robot control method and system based on Transfomer in the technical field of calligraphy robots. The method comprises the steps that S1, a large number of historical calligraphy images are collected to construct a data set; s2, creating a font image generation model based on the input layer, the two-way encoder, the decoder and the output layer, and setting a loss function of the font image generation model; s3, training the font image generation model through the data set and the loss function, and deploying the trained font image generation model to the calligraphy robot; s4, acquiring a real-time calligraphy image and calligraphy style data by the calligraphy robot, and inputting the real-time calligraphy image and the calligraphy style data into the deployed font image generation model to obtain a real-time font image; and S5, the calligraphy robot analyzes the real-time font image to obtain a writing control instruction so as to execute a writing action. The calligraphy robot has the advantage that the writing quality of the calligraphy robot is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of calligraphy robots, and in particular to a calligraphy robot control method and system based on Transformer. Background Art

[0002] The calligraphy robot is a product that combines modern robotics technology with traditional calligraphy art. It simulates the behavior of human calligraphy writing through high-precision robotic arm control, advanced sensor technology and artificial intelligence algorithms, and can complete the creation of various calligraphy works.

[0003] However, calligraphy robots currently have significant room for improvement in writing quality, artistic expression, and intelligence. In terms of writing control, existing methods struggle to achieve precise control of calligraphy strokes, are unable to effectively simulate the styles of different calligraphers, and the natural fluency of writing needs to be improved. In terms of font generation, traditional models suffer from insufficient style decoupling accuracy and poor multi-font compatibility, resulting in a lack of artistic quality and diversity in generated fonts. In terms of robot motion control, the transition from digital generation to physical execution presents difficulties, making it difficult to accurately map font outlines to robot trajectories, and the dynamic adaptability of multiple font styles is also weak.

[0004] Therefore, how to provide a calligraphy robot control method and system based on Transfomer to improve the writing quality of the calligraphy robot has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a calligraphy robot control method and system based on Transfomer, so as to improve the writing quality of the calligraphy robot.

[0006] In a first aspect, the present invention provides a calligraphy robot control method based on Transformer, comprising the following steps:

[0007] Step S1: collecting a large number of historical calligraphy images, pre-processing and annotating each of the historical calligraphy images, and then constructing a data set;

[0008] Step S2: creating a font image generation model based on the input layer, the dual-path encoder, the decoder, and the output layer, and setting a loss function of the font image generation model;

[0009] Step S3: training the font image generation model using the data set and the loss function, and deploying the trained font image generation model to the calligraphy robot;

[0010] Step S4: The calligraphy robot obtains a real-time calligraphy image and calligraphy style data, and inputs the real-time calligraphy image and calligraphy style data into a deployed font image generation model to obtain a real-time font image;

[0011] Step S5: The calligraphy robot analyzes the real-time font image to obtain a writing control instruction, and controls the robotic arm to perform a writing action based on the writing control instruction.

[0012] Furthermore, the step S1 is specifically as follows:

[0013] A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

[0014] Furthermore, in step S2, the input layer is constructed based on a content input module and a style input module; the content input module is used to input calligraphy images; and the style input module is used to input calligraphy style data;

[0015] The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features.

[0016] The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result;

[0017] The output layer is used to render the decoding result into a font image for output;

[0018] The loss function adopts the mean square error function.

[0019] Furthermore, the step S3 is specifically as follows:

[0020] Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold;

[0021] The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then:

[0022] Gazebo and Rviz are used to build a robot simulation environment. The stroke accuracy, writing speed and ink uniformity are calculated through the test set in the robot simulation environment to test the verified font image generation model. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated and the font image generation model that has passed the test is deployed to the calligraphy robot.

[0023] Furthermore, the step S5 is specifically as follows:

[0024] The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. It generates writing control instructions based on the motion trajectory coordinates and the instruction array, and controls the robotic arm to perform writing actions based on the writing control instructions.

[0025] In a second aspect, the present invention provides a calligraphy robot control system based on Transformer, comprising the following modules:

[0026] A data set construction module is used to collect a large number of historical calligraphy images, pre-process and annotate each of the historical calligraphy images, and then construct a data set;

[0027] A font image generation model creation module is used to create a font image generation model based on the input layer, the two-way encoder, the decoder, and the output layer, and set a loss function of the font image generation model;

[0028] A font image generation model training module, configured to train the font image generation model using the data set and the loss function, and deploy the trained font image generation model to the calligraphy robot;

[0029] A real-time font image generation module is used by the calligraphy robot to obtain a real-time calligraphy image and calligraphy style data, and input the real-time calligraphy image and calligraphy style data into the deployed font image generation model to obtain a real-time font image;

[0030] The writing control instruction execution module is used for the calligraphy robot to analyze the real-time font image to obtain a writing control instruction, and control the robot arm to perform a writing action based on the writing control instruction.

[0031] Furthermore, the dataset construction module is specifically used to:

[0032] A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

[0033] Furthermore, in the font image generation model creation module, the input layer is constructed based on the content input module and the style input module; the content input module is used to input calligraphy images; the style input module is used to input calligraphy style data;

[0034] The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features.

[0035] The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result;

[0036] The output layer is used to render the decoding result into a font image for output;

[0037] The loss function adopts the mean square error function.

[0038] Furthermore, the font image generation model training module is specifically used to:

[0039] Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold;

[0040] The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then:

[0041] Gazebo and Rviz are used to build a robot simulation environment. The stroke accuracy, writing speed and ink uniformity are calculated through the test set in the robot simulation environment to test the verified font image generation model. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated and the font image generation model that has passed the test is deployed to the calligraphy robot.

[0042] Furthermore, the writing control instruction execution module is specifically used to:

[0043] The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. It generates writing control instructions based on the motion trajectory coordinates and the instruction array, and controls the robotic arm to perform writing actions based on the writing control instructions.

[0044] The advantages of the present invention are:

[0045] 1. Collect a large number of historical calligraphy images, pre-process and annotate each historical calligraphy image, and then build a data set; create a font image generation model based on the input layer, dual-channel encoder, decoder, and output layer, set the loss function of the font image generation model, train the font image generation model through the data set and loss function, and deploy the trained font image generation model to the calligraphy robot; then the calligraphy robot obtains real-time calligraphy images and calligraphy style data, inputs the real-time calligraphy images and calligraphy style data into the deployed font image generation model, obtains real-time font images, analyzes the real-time font images to obtain writing control instructions, and controls the robotic arm to execute based on the writing control instructions. Line writing action; that is, the input real-time calligraphy image and calligraphy style data are converted into a real-time font image through a pre-trained font image generation model. Since the font image generation model is built based on Transfomer, it can effectively improve the artistic sense and diversity of the generated fonts, and effectively simulate the styles of different calligraphers. In the process of converting the real-time font image into writing control instructions, the outline data is extracted through the GetGlyphOutline function, and the outline data is converted using a straight line and quadratic spline curve interpolation algorithm, thereby effectively improving the accuracy of calligraphy stroke control, realizing accurate mapping of font outline to robot trajectory, and ultimately greatly improving the writing quality of the calligraphy robot.

[0046] 2. Through multi-level preprocessing such as grayscale conversion, noise reduction, normalization, and image segmentation, the clarity and standardization of calligraphy images are ensured; the annotation content covers font type, calligraphy style, and stroke order, and a multi-dimensional feature labeling system is constructed to provide structured data support for the model to learn complex calligraphy features.

[0047] 3. The content encoder combines ResNet18 (local feature extraction) with Transformer Encoder (global attention mechanism) to effectively capture the stroke order and spatial layout of Chinese characters. The style encoder reuses the underlying features of the content encoder through a shared feature extraction module to reduce computational redundancy. At the same time, it improves the precision of style control by separating and extracting global and local styles (such as ink thickness and brush stroke details).

[0048] 4. By adopting the Adam optimizer combined with a gradual decay strategy of the learning rate, we balance the convergence speed and model stability. By introducing dual verification of the validation set (accuracy, precision, recall rate, and F1 score) and the test set (stroke accuracy, writing speed, and ink uniformity), we ensure the generalization ability of the model in real scenarios.

[0049] 5. Use Gazebo and Rviz to build a robot simulation environment to simulate the robot arm's motion trajectory and physical interaction, reducing hardware debugging costs; directly link test indicators (such as writing speed and ink uniformity) to actual writing effects to ensure reliable performance after model deployment.

[0050] 6. Format conversion from PNG to SVG / TTF preserves the vector outline data of calligraphy strokes to avoid bitmap scaling distortion; a quadratic spline interpolation algorithm is used to generate continuous and smooth motion trajectories, imitating the natural brush movements of human calligraphers.

[0051] 7. The instruction array integrates position, pressure, and speed parameters to achieve dynamic adjustment of brush strokes (such as stopping and lifting the pen), enhancing the expressiveness of the brush strokes in calligraphy works; the robotic arm control logic is decoupled from the font generation model to improve the modular scalability of the system (such as adapting to robotic arms of different brands).

[0052] 8. By integrating the Transformer model with a dual-channel encoder architecture, combined with multimodal data preprocessing, dynamic learning rate optimization, and simulation environment testing, an efficient integrated calligraphy generation and control system was constructed: its innovation lies in the use of ResNet18 to extract multi-scale visual features, and the precise capture of calligraphy art details (such as brush strokes and ink marks) through global and local style separation encoding; and based on SVG / TTF vector conversion and quadratic spline interpolation algorithm, high-precision robotic arm instructions are generated to achieve end-to-end control from image to physical writing, taking into account both writing efficiency and artistic expression. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0054] Figure 1 This is a flow chart of a calligraphy robot control method based on Transfomer of the present invention.

[0055] Figure 2 It is a structural schematic diagram of a calligraphy robot control system based on Transfomer of the present invention. DETAILED DESCRIPTION

[0056] The technical solution in the embodiments of the present application has the following overall idea: the input real-time calligraphy image and calligraphy style data are converted into a real-time font image through a pre-trained font image generation model. Since the font image generation model is built based on Transfomer, it can effectively improve the artistic sense and diversity of the generated fonts, effectively simulate the styles of different calligraphers, and in the process of converting the real-time font image into writing control instructions, the outline data is extracted through the GetGlyphOutline function, and the outline data is converted using a straight line and quadratic spline curve interpolation algorithm, thereby effectively improving the accuracy of calligraphy stroke control, and realizing accurate mapping of font outlines to robot trajectories, thereby improving the writing quality of the calligraphy robot.

[0057] Please refer to Figures 1 to 2 As shown, a preferred embodiment of the calligraphy robot control method based on Transformer of the present invention includes the following steps:

[0058] Step S1: collecting a large number of historical calligraphy images, pre-processing and annotating each of the historical calligraphy images, and then constructing a data set;

[0059] Step S2: creating a font image generation model based on the input layer, the dual-path encoder, the decoder, and the output layer, and setting a loss function of the font image generation model;

[0060] Step S3: training the font image generation model using the data set and the loss function, and deploying the trained font image generation model to the calligraphy robot;

[0061] Step S4: The calligraphy robot obtains a real-time calligraphy image and calligraphy style data, and inputs the real-time calligraphy image and calligraphy style data into a deployed font image generation model to obtain a real-time font image;

[0062] Step S5: The calligraphy robot analyzes the real-time font image to obtain a writing control instruction, and controls the robotic arm to perform a writing action based on the writing control instruction.

[0063] The step S1 is specifically as follows:

[0064] A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

[0065] Through multi-level preprocessing such as grayscale, noise reduction, normalization, and image segmentation, the clarity and standardization of calligraphy images are ensured; the annotation content covers font type, calligraphy style, and stroke order, and a multi-dimensional feature label system is constructed to provide structured data support for the model to learn complex calligraphy features.

[0066] In step S2, the input layer is constructed based on a content input module and a style input module; the content input module is used to input calligraphy images; the style input module is used to input calligraphy style data;

[0067] The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features.

[0068] The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result;

[0069] The output layer is used to render the decoding result into a font image for output;

[0070] The loss function adopts the mean square error function.

[0071] The content encoder combines ResNet18 (local feature extraction) and Transformer Encoder (global attention mechanism) to effectively capture the stroke order and spatial layout of Chinese characters; the style encoder reuses the underlying features of the content encoder through a shared feature extraction module to reduce computational redundancy, while also improving the precision of style control by separately extracting global and local styles (such as ink thickness and brush stroke details).

[0072] The step S3 is specifically as follows:

[0073] Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold;

[0074] By adopting the Adam optimizer combined with a gradual learning rate decay strategy, a balance is achieved between convergence speed and model stability. By introducing dual verification of the validation set (accuracy, precision, recall rate, and F1 score) and the test set (stroke accuracy, writing speed, and ink uniformity), the model's generalization ability in real scenarios is ensured.

[0075] The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then:

[0076] Gazebo and Rviz are used to build a robot simulation environment. The stroke accuracy, writing speed and ink uniformity are calculated through the test set in the robot simulation environment to test the verified font image generation model. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated and the font image generation model that has passed the test is deployed to the calligraphy robot.

[0077] A robot simulation environment is built using Gazebo and Rviz to simulate the robot arm's motion trajectory and physical interactions, reducing hardware debugging costs. Test indicators (such as writing speed and ink uniformity) are directly linked to actual writing effects to ensure reliable performance after model deployment.

[0078] The step S5 is specifically as follows:

[0079] The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. Based on the motion trajectory coordinates and the instruction array, it generates a writing control instruction, and based on the writing control instruction, it controls the robotic arm to perform the writing action.

[0080] In the specific implementation, an anisotropic Gaussian kernel is used to smooth the font image to suppress noise; the six-degree-of-freedom serial robotic arm is kinematically modeled based on the DH parameters of the calligraphy robot. Through forward kinematics calculation and inverse kinematics solution, the two-dimensional motion trajectory coordinates generated based on the contour data are mapped into executable actions in the robotic arm joint space. Fifth-order polynomial interpolation is used for path planning to generate a smooth joint motion trajectory.

[0081] In specific implementation, a dynamic speed planning method based on curvature-sensitive speed and acceleration continuity constraints can be used to establish a speed-curvature coupling model based on the curvature characteristics of the strokes. Different speed parameters can be set for different types of strokes (straight lines, curves, and pauses), and a fifth-order polynomial interpolation method is used to generate a smooth speed curve to avoid uneven ink and mechanical vibration caused by sudden changes in speed. At the same time, through parameter definition and mathematical modeling, the calligraphy brush angle and pen lift height are optimized. Different pen lift heights are set according to different action types (continuous strokes, pauses), and the ink width is defined according to the angle between the brush axis and the surface normal.

[0082] In practice, TTF files and path optimization parameter files for different fonts can be pre-stored. These parameter files contain key information such as font identification, baseline speed, pen pressure, and pen lift height. When switching fonts, the target font is selected through the GUI interface, and the corresponding configuration is automatically matched and fully loaded. This avoids the performance overhead of real-time parameter interpolation and dynamic calculation, enabling rapid switching between multiple font styles while maintaining the continuity of writing movements and the consistency of ink.

[0083] The format conversion from PNG to SVG / TTF preserves the vector outline data of calligraphy strokes, avoiding bitmap scaling distortion; a continuous and smooth motion trajectory is generated through a quadratic spline curve interpolation algorithm, imitating the natural brush movements of human calligraphers.

[0084] The instruction array integrates position, pressure, and speed parameters to achieve dynamic adjustment of brush strokes (such as stopping and lifting the pen), enhancing the expressiveness of the brush strokes in calligraphy works; the robotic arm control logic is decoupled from the font generation model to improve the modular scalability of the system (such as adapting to robotic arms of different brands).

[0085] By integrating the Transformer model with a dual-channel encoder architecture, combined with multimodal data preprocessing, dynamic learning rate optimization, and simulation environment testing, an efficient integrated calligraphy generation and control system was constructed: its innovation lies in the use of ResNet18 to extract multi-scale visual features, and the precise capture of calligraphy art details (such as brush strokes and ink marks) through global and local style separation encoding; and based on SVG / TTF vector conversion and quadratic spline interpolation algorithm, high-precision robotic arm instructions are generated to achieve end-to-end control from image to physical writing, taking into account both writing efficiency and artistic expression.

[0086] A preferred embodiment of the calligraphy robot control system based on Transformer of the present invention includes the following modules:

[0087] A data set construction module is used to collect a large number of historical calligraphy images, pre-process and annotate each of the historical calligraphy images, and then construct a data set;

[0088] A font image generation model creation module is used to create a font image generation model based on the input layer, the two-way encoder, the decoder, and the output layer, and set a loss function of the font image generation model;

[0089] A font image generation model training module, configured to train the font image generation model using the data set and the loss function, and deploy the trained font image generation model to the calligraphy robot;

[0090] A real-time font image generation module is used by the calligraphy robot to obtain a real-time calligraphy image and calligraphy style data, and input the real-time calligraphy image and calligraphy style data into the deployed font image generation model to obtain a real-time font image;

[0091] The writing control instruction execution module is used for the calligraphy robot to analyze the real-time font image to obtain a writing control instruction, and control the robot arm to perform a writing action based on the writing control instruction.

[0092] The dataset construction module is specifically used for:

[0093] A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

[0094] Through multi-level preprocessing such as grayscale, noise reduction, normalization, and image segmentation, the clarity and standardization of calligraphy images are ensured; the annotation content covers font type, calligraphy style, and stroke order, and a multi-dimensional feature label system is constructed to provide structured data support for the model to learn complex calligraphy features.

[0095] In the font image generation model creation module, the input layer is constructed based on the content input module and the style input module; the content input module is used to input calligraphy images; the style input module is used to input calligraphy style data;

[0096] The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features.

[0097] The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result;

[0098] The output layer is used to render the decoding result into a font image for output;

[0099] The loss function adopts the mean square error function.

[0100] The content encoder combines ResNet18 (local feature extraction) and Transformer Encoder (global attention mechanism) to effectively capture the stroke order and spatial layout of Chinese characters; the style encoder reuses the underlying features of the content encoder through a shared feature extraction module to reduce computational redundancy, while also improving the precision of style control by separately extracting global and local styles (such as ink thickness and brush stroke details).

[0101] The font image generation model training module is specifically used for:

[0102] Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold;

[0103] By adopting the Adam optimizer combined with a gradual learning rate decay strategy, a balance is achieved between convergence speed and model stability. By introducing dual verification of the validation set (accuracy, precision, recall rate, and F1 score) and the test set (stroke accuracy, writing speed, and ink uniformity), the model's generalization ability in real scenarios is ensured.

[0104] The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then:

[0105] Gazebo and Rviz are used to build a robot simulation environment. The stroke accuracy, writing speed and ink uniformity are calculated through the test set in the robot simulation environment to test the verified font image generation model. If the test fails, the training set is expanded to continue training; if the test passes, the training is terminated and the font image generation model that has passed the test is deployed to the calligraphy robot.

[0106] A robot simulation environment is built using Gazebo and Rviz to simulate the robot arm's motion trajectory and physical interactions, reducing hardware debugging costs. Test indicators (such as writing speed and ink uniformity) are directly linked to actual writing effects to ensure reliable performance after model deployment.

[0107] The writing control instruction execution module is specifically used for:

[0108] The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. Based on the motion trajectory coordinates and the instruction array, it generates a writing control instruction, and based on the writing control instruction, it controls the robotic arm to perform the writing action.

[0109] In the specific implementation, an anisotropic Gaussian kernel is used to smooth the font image to suppress noise; the six-degree-of-freedom serial robotic arm is kinematically modeled based on the DH parameters of the calligraphy robot. Through forward kinematics calculation and inverse kinematics solution, the two-dimensional motion trajectory coordinates generated based on the contour data are mapped into executable actions in the robotic arm joint space. Fifth-order polynomial interpolation is used for path planning to generate a smooth joint motion trajectory.

[0110] In specific implementation, a dynamic speed planning method based on curvature-sensitive speed and acceleration continuity constraints can be used to establish a speed-curvature coupling model based on the curvature characteristics of the strokes. Different speed parameters can be set for different types of strokes (straight lines, curves, and pauses), and a fifth-order polynomial interpolation method is used to generate a smooth speed curve to avoid uneven ink and mechanical vibration caused by sudden changes in speed. At the same time, through parameter definition and mathematical modeling, the calligraphy brush angle and pen lift height are optimized. Different pen lift heights are set according to different action types (continuous strokes, pauses), and the ink width is defined according to the angle between the brush axis and the surface normal.

[0111] In practice, TTF files and path optimization parameter files for different fonts can be pre-stored. These parameter files contain key information such as font identification, baseline speed, pen pressure, and pen lift height. When switching fonts, the target font is selected through the GUI interface, and the corresponding configuration is automatically matched and fully loaded. This avoids the performance overhead of real-time parameter interpolation and dynamic calculation, enabling rapid switching between multiple font styles while maintaining the continuity of writing movements and the consistency of ink.

[0112] The format conversion from PNG to SVG / TTF preserves the vector outline data of calligraphy strokes, avoiding bitmap scaling distortion; a continuous and smooth motion trajectory is generated through a quadratic spline curve interpolation algorithm, imitating the natural brush movements of human calligraphers.

[0113] The instruction array integrates position, pressure, and speed parameters to achieve dynamic adjustment of brush strokes (such as stopping and lifting the pen), enhancing the expressiveness of the brush strokes in calligraphy works; the robotic arm control logic is decoupled from the font generation model to improve the modular scalability of the system (such as adapting to robotic arms of different brands).

[0114] By integrating the Transformer model with a dual-channel encoder architecture, combined with multimodal data preprocessing, dynamic learning rate optimization, and simulation environment testing, an efficient integrated calligraphy generation and control system was constructed: its innovation lies in the use of ResNet18 to extract multi-scale visual features, and the precise capture of calligraphy art details (such as brush strokes and ink marks) through global and local style separation encoding; and based on SVG / TTF vector conversion and quadratic spline interpolation algorithm, high-precision robotic arm instructions are generated to achieve end-to-end control from image to physical writing, taking into account both writing efficiency and artistic expression.

[0115] In summary, the advantages of the present invention are:

[0116] 1. Collect a large number of historical calligraphy images, pre-process and annotate each historical calligraphy image, and then build a data set; create a font image generation model based on the input layer, dual-channel encoder, decoder, and output layer, set the loss function of the font image generation model, train the font image generation model through the data set and loss function, and deploy the trained font image generation model to the calligraphy robot; then the calligraphy robot obtains real-time calligraphy images and calligraphy style data, inputs the real-time calligraphy images and calligraphy style data into the deployed font image generation model, obtains real-time font images, analyzes the real-time font images to obtain writing control instructions, and controls the robotic arm to execute based on the writing control instructions. Line writing action; that is, the input real-time calligraphy image and calligraphy style data are converted into a real-time font image through a pre-trained font image generation model. Since the font image generation model is built based on Transfomer, it can effectively improve the artistic sense and diversity of the generated fonts, and effectively simulate the styles of different calligraphers. In the process of converting the real-time font image into writing control instructions, the outline data is extracted through the GetGlyphOutline function, and the outline data is converted using a straight line and quadratic spline curve interpolation algorithm, thereby effectively improving the accuracy of calligraphy stroke control, realizing accurate mapping of font outline to robot trajectory, and ultimately greatly improving the writing quality of the calligraphy robot.

[0117] 2. Through multi-level preprocessing such as grayscale conversion, noise reduction, normalization, and image segmentation, the clarity and standardization of calligraphy images are ensured; the annotation content covers font type, calligraphy style, and stroke order, and a multi-dimensional feature labeling system is constructed to provide structured data support for the model to learn complex calligraphy features.

[0118] 3. The content encoder combines ResNet18 (local feature extraction) with Transformer Encoder (global attention mechanism) to effectively capture the stroke order and spatial layout of Chinese characters. The style encoder reuses the underlying features of the content encoder through a shared feature extraction module to reduce computational redundancy. At the same time, it improves the precision of style control by separating and extracting global and local styles (such as ink thickness and brush stroke details).

[0119] 4. By adopting the Adam optimizer combined with a gradual decay strategy of the learning rate, we balance the convergence speed and model stability. By introducing dual verification of the validation set (accuracy, precision, recall rate, and F1 score) and the test set (stroke accuracy, writing speed, and ink uniformity), we ensure the generalization ability of the model in real scenarios.

[0120] 5. Use Gazebo and Rviz to build a robot simulation environment to simulate the robot arm's motion trajectory and physical interaction, reducing hardware debugging costs; directly link test indicators (such as writing speed and ink uniformity) to actual writing effects to ensure reliable performance after model deployment.

[0121] 6. Format conversion from PNG to SVG / TTF preserves the vector outline data of calligraphy strokes to avoid bitmap scaling distortion; a quadratic spline interpolation algorithm is used to generate continuous and smooth motion trajectories, imitating the natural brush movements of human calligraphers.

[0122] 7. The instruction array integrates position, pressure, and speed parameters to achieve dynamic adjustment of brush strokes (such as stopping and lifting the pen), enhancing the expressiveness of the brush strokes in calligraphy works; the robotic arm control logic is decoupled from the font generation model to improve the modular scalability of the system (such as adapting to robotic arms of different brands).

[0123] 8. By integrating the Transformer model with a dual-channel encoder architecture, combined with multimodal data preprocessing, dynamic learning rate optimization, and simulation environment testing, an efficient integrated calligraphy generation and control system was constructed: its innovation lies in the use of ResNet18 to extract multi-scale visual features, and the precise capture of calligraphy art details (such as brush strokes and ink marks) through global and local style separation encoding; and based on SVG / TTF vector conversion and quadratic spline interpolation algorithm, high-precision robotic arm instructions are generated to achieve end-to-end control from image to physical writing, taking into account both writing efficiency and artistic expression.

[0124] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A calligraphy robot control method based on Transformer, characterized by: The steps include: Step S1: collecting a large number of historical calligraphy images, pre-processing and annotating each of the historical calligraphy images, and then constructing a data set; Step S2: creating a font image generation model based on the input layer, the dual-path encoder, the decoder, and the output layer, and setting a loss function of the font image generation model; Step S3: training the font image generation model using the data set and the loss function, and deploying the trained font image generation model to the calligraphy robot; Step S4: The calligraphy robot obtains a real-time calligraphy image and calligraphy style data, and inputs the real-time calligraphy image and calligraphy style data into a deployed font image generation model to obtain a real-time font image; Step S5: The calligraphy robot analyzes the real-time font image to obtain a writing control instruction, and controls the robotic arm to perform a writing action based on the writing control instruction.

2. The calligraphy robot control method based on Transformer according to claim 1, characterized in that: The step S1 is specifically as follows: A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

3. The calligraphy robot control method based on Transformer according to claim 1, characterized in that: In step S2, the input layer is constructed based on a content input module and a style input module; the content input module is used to input calligraphy images; the style input module is used to input calligraphy style data; The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features. The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result; The output layer is used to render the decoding result into a font image for output; The loss function adopts the mean square error function.

4. The calligraphy robot control method based on Transformer according to claim 1, characterized in that: The step S3 is specifically as follows: Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold; The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then: Use Gazebo and Rviz to build a robot simulation environment, and calculate the stroke accuracy, writing speed, and ink uniformity using a test set in the robot simulation environment to test the font image generation model that has passed the verification. If the test fails, expand the training set and continue training; If the test passes, the training is terminated and the font image generation model that passes the test is deployed to the calligraphy robot.

5. The calligraphy robot control method based on Transformer according to claim 1, characterized in that: The step S5 is specifically as follows: The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. Based on the motion trajectory coordinates and the instruction array, it generates a writing control instruction, and based on the writing control instruction, it controls the robotic arm to perform the writing action.

6. A calligraphy robot control system based on Transformer, characterized by: Includes the following modules: A data set construction module is used to collect a large number of historical calligraphy images, pre-process and annotate each of the historical calligraphy images, and then construct a data set; A font image generation model creation module is used to create a font image generation model based on the input layer, the two-way encoder, the decoder, and the output layer, and set a loss function of the font image generation model; A font image generation model training module, configured to train the font image generation model using the data set and the loss function, and deploy the trained font image generation model to the calligraphy robot; A real-time font image generation module is used by the calligraphy robot to obtain a real-time calligraphy image and calligraphy style data, and input the real-time calligraphy image and calligraphy style data into the deployed font image generation model to obtain a real-time font image; The writing control instruction execution module is used for the calligraphy robot to analyze the real-time font image to obtain a writing control instruction, and control the robot arm to perform a writing action based on the writing control instruction.

7. The calligraphy robot control system based on Transformer according to claim 6, characterized in that: The dataset construction module is specifically used for: A large number of historical calligraphy images are collected, and each of the historical calligraphy images is preprocessed by at least grayscale conversion, noise reduction, normalization, and image segmentation. Each of the preprocessed historical calligraphy images is annotated with at least font type, calligraphy style, and stroke order, and a data set is constructed based on the annotated historical calligraphy images.

8. The calligraphy robot control system based on Transformer according to claim 6, characterized in that: In the font image generation model creation module, the input layer is constructed based on the content input module and the style input module; the content input module is used to input calligraphy images; the style input module is used to input calligraphy style data; The dual-path encoder is constructed based on a content encoder and a style encoder; the content encoder is used to extract a content feature vector including at least the stroke order and spatial layout of Chinese characters from an input calligraphy image through ResNet18 and TransformerEncoder; the style encoder is constructed based on a shared feature extraction module, a global style extraction module, and a local style extraction module; the shared feature extraction module is used to extract multi-scale visual features from the ResNet18 of the content encoder; the global style extraction module is used to extract a global style feature vector from the multi-scale visual features; and the local style extraction module is used to extract a local style feature vector from the multi-scale visual features. The decoder is used to decode the content feature vector, the global style feature vector and the local style feature vector to obtain a decoding result; The output layer is used to render the decoding result into a font image for output; The loss function adopts the mean square error function.

9. The calligraphy robot control system based on Transformer according to claim 6, characterized in that: The font image generation model training module is specifically used for: Dividing the dataset into a training set, a validation set, and a test set based on a preset ratio, training a font image generation model using the training set in combination with an Adam optimizer, continuously optimizing hyperparameters of the font image generation model including at least a learning rate, a batch size, and a number of training rounds during the training process, and gradually reducing the learning rate as the number of training rounds increases until the loss value of the loss function is less than a preset loss threshold; The accuracy, precision, recall and F1 score of the validation set are calculated to verify the trained font image generation model. If the validation fails, the training set is expanded to continue training. If the validation passes, then: Use Gazebo and Rviz to build a robot simulation environment, and calculate the stroke accuracy, writing speed, and ink uniformity using a test set in the robot simulation environment to test the font image generation model that has passed the verification. If the test fails, expand the training set and continue training; If the test passes, the training is terminated and the font image generation model that passes the test is deployed to the calligraphy robot.

10. The calligraphy robot control system based on Transformer according to claim 6, characterized in that: The writing control instruction execution module is specifically used for: The calligraphy robot converts the real-time font image from PNG format to SVG format, and then converts it into a TTF font file. It extracts the outline data of the TTF font file through the GetGlyphOutline function, parses the outline data, and converts the outline data into motion trajectory coordinates using a straight line and quadratic spline curve interpolation algorithm. It also generates an instruction array containing position, pressure, and speed. Based on the motion trajectory coordinates and the instruction array, it generates a writing control instruction, and based on the writing control instruction, it controls the robotic arm to perform the writing action.

Citation Information

Cited By

  • A geometric perception curved calligraphy drawing method based on a diffusion model

    CN122683701A