Method for realizing cultural relic image restoration and reconstruction based on large model and fine tuning technology

Through methods based on large-scale models and fine-tuning technology, the problems of insufficient accuracy and poor generalization capabilities in cultural relics image restoration and reconstruction are solved, and the repair and reconstruction of cultural relics image with high accuracy and generalization capabilities are achieved, providing more effective cultural relics protection and research technical support.

CN120107116AInactive Publication Date: 2025-06-06NANTONG SHUOYUANYIKE DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164901.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and poor generalization capabilities in the restoration and reconstruction of cultural relics images, resulting in unsatisfactory repair results.

Method used

Using methods based on large-scale models and fine-tuning technology, the accuracy and generalization capabilities of cultural relics image restoration and reconstruction are improved through the steps of data collection and sorting, model selection and initialization, LoRA fine-tuning settings, model training, model verification and selection, and cultural relics image restoration and reconstruction.

Benefits of technology

It has achieved high accuracy and generalization capabilities for restoring and reconstruction of cultural relics images, can adapt to cultural relics images of different types and degrees of damage, and has improved technical support for cultural relics protection and research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107116A_ABST
    Figure CN120107116A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing and artificial intelligence, in particular to a method for achieving cultural relic image restoration and reconstruction based on a large model and a fine tuning technology, and the method comprises the steps of data collection and arrangement, model selection and initialization, LoRA fine tuning setting, model training, model verification and selection, and cultural relic image restoration and reconstruction. The method has data richness: a unique large-scale Chinese cultural relic information database provides sufficient and high-quality data for model training, and helps to improve the ability of the model to understand and restore cultural relic images; efficient fine tuning: a LoRA fine tuning strategy is adopted, the number of trainable parameters is reduced, the model performance is maintained, the calculation cost and the training time are reduced, and the training efficiency is improved; and accuracy and generalization ability: through training and verification on a specific cultural relic image data set, the model can accurately repair and reconstruct the cultural relic image, and the model has good generalization ability and can adapt to cultural relic images of different types and damage degrees.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing and artificial intelligence technology, and in particular to a method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology. Background Art

[0002] Cultural relics are important carriers of human history and culture. However, due to various factors such as natural erosion and human destruction, many cultural relics are facing the problem of damage and information loss. The restoration and reconstruction of cultural relic images are of vital importance for the protection, research and inheritance of cultural relics. Traditional methods of cultural relic image restoration often rely on manual operations, which are inefficient and the restoration effect is limited by the professional level and experience of the restorers. With the development of artificial intelligence technology, the use of machine learning and deep learning algorithms for cultural relic image restoration has become a research hotspot, but the current methods still have problems such as insufficient accuracy and poor generalization ability when processing complex cultural relic images.

[0003] Therefore, those skilled in the art provide a method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology to solve the problems raised in the above background technology. Summary of the invention

[0004] To solve the above technical problems, the present invention provides a method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology, so as to improve the accuracy, efficiency and generalization ability of cultural relics image restoration and reconstruction, and provide more effective technical support for cultural relics protection and research.

[0005] The following steps are involved:

[0006] Step 1: Data collection and organization: continuously collect and update the Chinese cultural relics information database to ensure the integrity and accuracy of the data, and classify, annotate and pre-process the collected graphic, text and video data to make them suitable for model training;

[0007] Step 2: Model selection and initialization: According to the task requirements of cultural relic image restoration and reconstruction, select appropriate general large models (LLMs) and initialize the model;

[0008] Step 3: LoRA fine-tuning settings, add low-rank matrices A and B in specific layers of the selected general large model, and determine the range and initial values ​​of trainable parameters;

[0009] Step 4: Model training;

[0010] Step 5: Model verification and selection;

[0011] Step 6: Restoration and reconstruction of cultural relic images. Input the cultural relic image to be restored into the selected model, and the model outputs the restored and reconstructed image; post-process and evaluate the quality of the restored image, and perform manual intervention and adjustment when necessary.

[0012] Preferably, the model training process in step 4 includes data loading, parameter adjustment and model saving.

[0013] Preferably, the data loading is loading the preprocessed cultural relic image data into a training environment and grouping them according to a set parallel quantity.

[0014] Preferably, the parameter adjustment is to dynamically adjust the learning rate and other parameters of the model during the training process according to the set parameters, and observe the training effect and loss value changes of the model.

[0015] Preferably, the model is saved as the LORA model generated during the training process according to the set total number of steps and LOSS value standards.

[0016] Preferably, the model verification and selection in step five includes effect evaluation and model determination.

[0017] Preferably, the effect evaluation is to use all the generated LORA models, generate them on the basis of the base model according to the set weight range, and evaluate the restoration effect of the model in terms of color, outline, image, clothing, background decoration, etc. by comparing the generated image with the original cultural relic image.

[0018] Preferably, the model is determined by selecting the best LORA model as the final model for restoration and reconstruction of cultural relics images based on the evaluation score and the generated image effect.

[0019] Preferably, the general large models (LLMs) in step 2 are LLMA, GPT, Gemini, etc.

[0020] Technical effects and advantages of the present invention:

[0021] Data richness: The unique large-scale Chinese cultural relics information database provides sufficient and high-quality data for model training, which helps to improve the model's understanding and restoration capabilities of cultural relic images.

[0022] Efficient fine-tuning: The LoRA fine-tuning strategy is adopted to reduce the number of trainable parameters while maintaining model performance, reducing computing costs and training time, and improving training efficiency.

[0023] Accuracy and generalization ability: Through training and verification on a specific cultural relic image dataset, the model can accurately repair and reconstruct cultural relic images, and has good generalization ability and can adapt to cultural relic images of different types and degrees of damage. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a picture of the Yulin Grottoes training set of the method for realizing restoration and reconstruction of cultural relics images based on a large model and fine-tuning technology provided in an embodiment of the present application, showing the selected pictures of Yulin Grottoes 25 mainly featuring figures, which are used to illustrate the data selection in the training preprocessing stage;

[0025] Figure 2 It is a learning rate variation diagram of the method for realizing restoration and reconstruction of cultural relic images based on a large model and fine-tuning technology provided in an embodiment of the present application, which intuitively presents the floating change of the learning rate during the training process and helps to understand the dynamic adjustment of the training parameters;

[0026] Figure 3 It is a graph showing the variation of the learning rate of the method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology provided in an embodiment of the present application;

[0027] Figure 4 This is a partial model training loss change diagram of the method for realizing cultural relic image restoration and reconstruction based on a large model and fine-tuning technology provided in an embodiment of the present application, which shows the change trend of the loss value during the model training process and indicates the stability and convergence of the model training;

[0028] Figure 5 This is an xy verification diagram of the method for realizing restoration and reconstruction of cultural relics images based on a large model and fine-tuning technology provided in an embodiment of the present application, with the x-axis being the model number and the y-axis being the weight, showing the generation effects of different LORA models under different weights, which is used for model verification and selection;

[0029] Figure 6 This is a mural example (left) and a generated image (right) comparison image of the method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology provided in an embodiment of the present application. The mural figures in Cave 25 in Yulin are compared with the generated image to intuitively demonstrate the evaluation basis of the model generation effect;

[0030] Figure 7 This is a comparison diagram of the original image (left) and the restored image (right) of the method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology provided in an embodiment of the present application. The original image and the restored image of a blurred and damaged mural from the mid-Tang Dynasty are compared to show the restoration and high-definition effect of the model on real murals;

[0031] Figure 8 It is the relationship between the linear layer and the low-rank matrix in the method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology provided in an embodiment of the present application;

[0032] Fig. 9 This is a demonstration process of fine-tuning in the method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for the purpose of illustration and description, and are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.

[0034] Example 1

[0035] See also Figures 1 to 9 In this embodiment, a method for restoring and reconstructing a cultural relic image based on a large model and fine-tuning technology is provided, including:

[0036] The following steps are involved:

[0037] Step 1: Data collection and organization: continuously collect and update the Chinese cultural relics information database to ensure the integrity and accuracy of the data, and classify, annotate and pre-process the collected graphic, text and video data to make them suitable for model training;

[0038] Step 2: Model selection and initialization: According to the task requirements of cultural relic image restoration and reconstruction, select appropriate general large models (LLMs) and initialize the model;

[0039] Step 3: LoRA fine-tuning settings. After training on a specific data set, a sub-model for image restoration and reconstruction of cultural relics of a specific vertical category can be generated;

[0040] Step 4: Model training;

[0041] Step 5: Model verification and selection;

[0042] Step 6: Restoration and reconstruction of cultural relic images. Input the cultural relic image to be restored into the selected model, and the model outputs the restored and reconstructed image; post-process and evaluate the quality of the restored image, and perform manual intervention and adjustment when necessary.

[0043] The Chinese Cultural Relics Information Database has a total data volume of 31T. The database covers graphic and text data related to ancient Chinese cultural relics from mainstream auction houses at home and abroad in the past 25 years, as well as graphic, text and video data related to Chinese art collections in major museums at home and abroad. These exclusive Chinese cultural relics information corpora will serve as key training preprocessing inputs for the subsequent training of AI vertical models for specific cultural relic image restoration and reconstruction, serving as key training preprocessing inputs.

[0044] The general large models (LLMs) in step 2 are LLMA, GPT, Gemini, etc.

[0045] Preferably, the specific method in step three is as follows:

[0046] a) Fine-tuning refers to adjusting the weights of the pre-trained model on a new dataset to improve the performance of the model in a specific field or task.

[0047] b) The core concept of LoRA is to train only a few parameters compared to full fine-tuning while maintaining the performance that can be achieved by full fine-tuning;

[0048] More specifically, two low-rank matrices A and B are added in certain layers. These low-rank matrices contain trainable parameters:

[0049] The linear layer is a layer that performs linear transformations in a neural network, mapping input data to the output space through matrix multiplication and bias terms; specifically:

[0050] The basic function of the linear layer is to implement linear transformation. This transformation can be expressed by the formula y=xW^T+b, where x is the input vector, W is the weight matrix, and b is the bias vector;

[0051] The matrix W and the bias vector b are both trainable parameters that determine how the linear layer transforms input data into output data;

[0052] By adjusting these parameters, the model can learn the linear relationship between input data and output data;

[0053] Figure 8 The letters in are as follows:

[0054] X is the input vector (row vector); its value range belongs to the matrix domain R with 1 row and d columns;

[0055] Left: The weight matrix W belongs to the matrix domain R with k rows and d columns;

[0056] Right side: The weight matrix W can also be transformed into the form of low-rank matrix A*low-rank matrix B, where matrix A belongs to the matrix domain R of d columns*r rows, and matrix B belongs to the matrix domain R of r columns*k rows; r is the intermediate value between dk;

[0057] Top: y is the output vector (column vector); its value range belongs to the matrix with 1 column and k rows, domain R;

[0058] Mathematically, this is achieved by modifying the weight matrix ΔW in the linear layer:

[0059] Where (W+ΔW)x=Wx+ABx

[0060] Where W is the original weight matrix, ΔW is the update or adjustment of the weight, and in LoRA, ΔW is decomposed into the product of two low-rank matrices A and B;

[0061] Since the dimensions of matrices A and B are much smaller compared to ΔW, the number of trainable parameters is significantly reduced.

[0062] Example:

[0063] 1. Training preprocessing: Taking the restoration of the murals in the Yulin Caves in Dunhuang as an example, we selected 50 pictures of people from the 25 caves in Yulin as the training set. Figure 1 shown.

[0064] 2. Parameter training:

[0065] 2.1. Parameter overview: including

[0066] Image (number of images)

[0067] Repeat (number of learning times, the number of times each picture is learned. Too few Repeats will result in insufficient fitting, while too high Repeats will lead to overfitting)

[0068] Epoch (number of cycles, the more cycles, the better the learning effect)

[0069] Batch size (the number of parallel operations, the number of epochs learned at the same time, the larger the batch size, the slower the convergence, and the smaller the batch size, the fewer the batch size)

[0070] Unet lr (learning rate, which can be understood as the step size of deep learning to find the optimal solution)

[0071] Text encord ir (text learning rate)

[0072] Optimizer (Optimizer - adamW8bit, DAdaptation, Lion)

[0073] Dim (the dimension of the neural network. The larger the dimension, the stronger the model's expressive power)

[0074] 2.2 Parameter selection: Settings

[0075] Repeat is 12

[0076] Epoch is 20

[0077] Batch size is 8

[0078] Unet ir is 1e-5

[0079] Text ir is 1e-6

[0080] Dim is 32

[0081] Optimizer: adamW8bit

[0082] The total number of steps is between 5000 and 12000, and the LORA model is output when the LOSS value is low;

[0083] In order to be more effective during training, the learning rate is kept in a floating state, such as Figure 2 shown.

[0084] 3. Model verification and selection: Figure 5 As shown,

[0085] 3.1. Generate 2 models in total at the loss value valley in the previous step

[0086] 3.2. Use all the generated LORA models and generate them on the basis of the base model in turn according to the weight of 0-1 to observe the effect of their influence. The x-axis is the model number and the y-axis is the weight;

[0087] 3.3, the leftmost is the control group without LORA model intervention (weight is 0);

[0088] The keywords are unified as 1girl, white hair, which is used to observe whether there is a fitting problem when different weights are used; finally, according to the generated graph, a model with better effect is selected, such as Figure 5 The third model from the right in ;

[0089] 3.4. This selection Figure 5 The third model from the right in .

[0090] Select one of the generated graphs for analysis.

[0091] like Figure 6 As shown, on the left are the mural figures in Cave 25 in Yulin, and on the right are the generated figures for this time; as shown by the red circle line, the image meets the requirements in many aspects.

[0092] 4. Image restoration and high definition based on real murals: Figure 7 As shown in the figure, a blurred and damaged mural from the middle Tang Dynasty was restored using the selected model; the model destroyed and redrawn some details of the characters and scenes, but it can achieve good restoration and high-definition effects.

[0093] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention are implemented according to the conventional means in the field unless otherwise specified and limited.

Claims

1. A method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology, characterized in that: The following steps are involved: Step 1: Data collection and organization: continuously collect and update the Chinese cultural relics information database, classify, annotate and pre-process the collected pictures, texts and video data to make them suitable for model training; Step 2: Model selection and initialization: According to the task requirements of cultural relic image restoration and reconstruction, select a suitable general large model and initialize the model; Step 3: LoRA fine-tuning settings, add low-rank matrices A and B in specific layers of the selected general large model, and determine the range and initial values ​​of trainable parameters; Step 4: Model training; Step 5: Model verification and selection; Step 6: Restoration and reconstruction of cultural relic images. Input the cultural relic image to be restored into the selected model, and the model outputs the restored and reconstructed image; post-process and evaluate the quality of the restored image, and perform manual intervention and adjustment when necessary.

2. The method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology according to claim 1 is characterized in that: The model training process in step 4 includes data loading, parameter adjustment and model saving.

3. The method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology according to claim 2 is characterized in that: The data loading is to load the pre-processed cultural relic image data into the training environment and group them according to the set parallel quantity.

4. The method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology according to claim 3 is characterized in that: The parameter adjustment is to dynamically adjust the learning rate and other parameters of the model during the training process according to the set parameters, and observe the training effect and loss value changes of the model.

5. The method for restoring and reconstructing cultural relics images based on a large model and fine-tuning technology according to claim 4 is characterized in that: The model is saved as the LORA model generated during the training process according to the set total number of steps and LOSS value standards.

6. The method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology according to claim 1 is characterized in that: The model verification and selection in step five include effect evaluation and model determination.

7. The method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology according to claim 6 is characterized in that: The effect evaluation is to use all the generated LORA models, generate them based on the base model according to the set weight range, and evaluate the restoration effect of the model in terms of color, outline, image, clothing, and background decoration by comparing the generated image with the original cultural relic image.

8. The method for restoring and reconstructing cultural relic images based on a large model and fine-tuning technology according to claim 7 is characterized in that: The model is determined by selecting the best LORA model as the final model for restoration and reconstruction of cultural relics images based on the evaluation score and the generated image effect.