A method, apparatus, device, and medium for digital reconstruction of stele images
By extracting corner maps and character content description information of the stele image, combined with deep feature extraction and prediction of characters, the problem of low accuracy in the reconstruction of the stele image is solved, and higher reconstruction accuracy and information integrity are achieved.
Patent Information
- Application Number
- CN202411057102.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-08-02
AI Technical Summary
The existing method of reconstruction of stele image has the problem of low accuracy in reconstruction of stele image, especially when dealing with the relationship between characters and image background, it is difficult to maintain the integrity of the structural information inside the characters and the global features of the image.
By acquiring multiple stone tablet images, extracting corner maps and character content description information, extracting shallow features and deep feature extraction, predicting characters and obtaining image prior knowledge, the fusion model integrates features, prior knowledge and description information, using the reconstruction model for image reconstruction, and optimizing the model through loss function to improve reconstruction accuracy.
It improves the accuracy and information richness of the stone tablet image reconstruction, ensures the integrity of the internal structure information of characters and the global features of the image, and improves the clarity and accuracy of the reconstruction image.
Smart Images

Figure CN118941668B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image reconstruction, and particularly relates to a method, device, equipment and medium for digital reconstruction of stone tablet images. Background Art
[0002] Stone tablet cultural relics have experienced erosion by various natural factors such as weathering and earthquakes over the long history, resulting in the gradual loss of their structure and surface details, and different degrees of degradation problems such as blurring and fracture. With the development of scanning technology and artificial intelligence technology, researchers use digital reconstruction technology to digitally model and reconstruct cultural relic images, which is not only beneficial to the protection of the stone tablet cultural relics themselves, but also provides valuable historical materials for academic research. Scholars can conduct detailed text and image analysis through digital cultural relics, so as to deeply study the culture, economy and technological level of ancient society.
[0003] During the process of digital reconstruction of stone tablet images, due to the existing blurring and degradation problems in the stone tablet images, there is a certain connection and overlap between the text characters and the background patterns in the images, the shape contours of the characters are complex and not clear enough. Many studies have not been able to effectively maintain the integrity of the internal structure information of the characters and the global features of the images when dealing with the relationship between the characters and the image background. It can be seen that the existing stone tablet image reconstruction methods have the problem of low reconstruction accuracy of stone tablet images. Summary of the Invention
[0004] The present application provides a method, device, equipment and medium for digital reconstruction of stone tablet images, which can solve the problem of low reconstruction accuracy of stone tablet images.
[0005] In a first aspect, an embodiment of the present application provides a method for digital reconstruction of stone tablet images, and the digital reconstruction method includes:
[0006] Obtain a plurality of stone tablet images, and extract the corner point map and character content description information of each stone tablet image; the corner point map is an image in which the corner points of all characters in the stone tablet image are marked;
[0007] For each stone tablet image respectively, obtain the shallow features of the stone tablet image according to the stone tablet image and the corner point map of the stone tablet image, and perform deep feature extraction on the shallow features to obtain the deep features of the stone tablet image;
[0008] For each deep feature respectively, perform prediction on the deep feature to obtain the predicted characters of the stone tablet image corresponding to the deep feature, and obtain the image prior knowledge of the predicted characters; the image prior knowledge is used to describe the image information generated according to the predicted characters;
[0009] The shallow features, image prior knowledge, and character content description information corresponding to each stone tablet image are fused using a fusion model to obtain the final features of each stone tablet image;
[0010] Each final feature is reconstructed using a reconstruction model to obtain the reconstructed image of the stone tablet image corresponding to each final feature;
[0011] A reconstruction loss function and a pixel loss function are constructed based on all the reconstructed images, a boundary loss function is constructed based on all the final features, and the fusion model and the reconstruction model are optimized using the reconstruction loss function, the pixel loss function, and the boundary loss function to obtain an optimized fusion model and an optimized reconstruction model; the reconstruction loss function and the pixel loss function are used to describe the accuracy of all the reconstructed images, and the boundary loss function is used to describe the accuracy of all the final features;
[0012] The optimized fusion model and the optimized reconstruction model are used to reconstruct the stone tablet image to be reconstructed to obtain the final reconstructed image of the stone tablet image to be reconstructed.
[0013] Optionally, according to the stone tablet image and the corner point map of the stone tablet image, the shallow features of the stone tablet image are obtained, including:
[0014] Through the formula:
[0015] R m,2 ,C 2 =UNet(R m,1 ,C 1 )
[0016] R m,1 ,C 1 =ResNet(R m,0 ,C 0 )
[0017] R m,0 ,C 0 =Init(I stone,m ,Cor stone,m )
[0018] Obtain the shallow feature R of the m-th stone tablet image m,2 ;
[0019] Among them, C 2 , C 1 , C 0 all represent the number of feature map channels, UNet( ) represents the UNet network operation, R m,1 represents the secondary feature, ResNet( ) represents the ResNet network operation, R m,0 represents the initial feature, Init( ) represents the convolution layer and pooling layer operations, I stone,mDenote the m-th stele image, Cor stone,m Denote the corner point map of the m-th stele image, where m = 1, 2,..., M, and M represents the total number of stele images.
[0020] Optionally, use a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stele image to obtain the final features of each stele image, including:
[0021] For each stele image, perform the following steps:
[0022] Fuse the shallow features corresponding to the stele image and the corner point map to obtain initial fusion features, and fuse the initial fusion features and the image prior knowledge corresponding to the stele image to obtain secondary fusion features;
[0023] Extract features from the character content description information corresponding to the stele image to obtain text features, and fuse the text features and the secondary fusion features to obtain the final features of the stele image.
[0024] Optionally, use a reconstruction model to reconstruct each final feature to obtain a reconstructed image of the stele image corresponding to each final feature, including:
[0025] Through the formula:
[0026] R m,SR = Pix(R m,2 + RL m,cross )
[0027] Obtain the reconstructed image R of the m-th stele image m,SR ;
[0028] where Pix( ) represents upsampling, R m,2 represents the shallow features of the m-th stele image, and RL m,cross represents the final features of the m-th stele image, where m = 1, 2,..., M, and M represents the total number of stele images.
[0029] Optionally, construct a reconstruction loss function and a pixel loss function based on all the reconstructed images, including:
[0030] Perform text recognition on each reconstructed image to obtain the text recognition results of each reconstructed image, and construct a reconstruction loss function based on all the text recognition results;
[0031] Construct a pixel loss function based on all the reconstructed images.
[0032] Optionally, the reconstruction loss function is:
[0033] l rec = 1 / m ∑(presr,m -pre hr,m ) 2
[0034] where l rec represents the value of the reconstruction loss function, and pre sr,m represents the character recognition result of the m-th reconstructed image, and pre hr,m represents the true character label of the m-th stone tablet image;
[0035] The pixel loss function is:
[0036]
[0037] where l 2 represents the value of the pixel loss function, represents the pixel value of the i-th pixel point of the m-th reconstructed image, represents the pixel value of the i-th pixel point of the m-th stone tablet image.
[0038] Optionally, the boundary loss function is:
[0039]
[0040] where l box represents the value of the boundary loss function, and t m represents the bounding box parameters obtained according to the final feature corresponding to the m-th stone tablet image, represents the true bounding box parameters of the m-th stone tablet image.
[0041] In a second aspect, an embodiment of the present application provides a digital reconstruction device for a stone tablet image, including:
[0042] An acquisition module that acquires a plurality of stone tablet images and extracts the corner point map and character content description information of each stone tablet image; the corner point map is an image in which the corner points of all characters in the stone tablet image are marked;
[0043] A feature extraction module that, for each stone tablet image, obtains the shallow feature of the stone tablet image according to the stone tablet image and the corner point map of the stone tablet image, and performs deep feature extraction on the shallow feature to obtain the deep feature of the stone tablet image;
[0044] A prediction module that, for each deep feature, predicts the deep feature to obtain the predicted character of the stone tablet image corresponding to the deep feature, and obtains the image prior knowledge of the predicted character; the image prior knowledge is used to describe the image information generated according to the predicted character;
[0045] A fusion module that uses a fusion model to fuse the shallow feature, image prior knowledge, and character content description information corresponding to each stone tablet image to obtain the final feature of each stone tablet image;
[0046] The first reconstruction module uses a reconstruction model to reconstruct each final feature to obtain a reconstructed image of the stone tablet image corresponding to each final feature;
[0047] The optimization module constructs a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructs a boundary loss function based on all the final features, and uses the reconstruction loss function, the pixel loss function, and the boundary loss function to optimize the fusion model and the reconstruction model to obtain an optimized fusion model and an optimized reconstruction model; the reconstruction loss function and the pixel loss function are used to describe the accuracy of all the reconstructed images, and the boundary loss function is used to describe the accuracy of all the final features;
[0048] The second reconstruction module uses the optimized fusion model and the optimized reconstruction model to reconstruct the stone tablet image to be reconstructed to obtain a final reconstructed image of the stone tablet image to be reconstructed.
[0049] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the digital reconstruction method of the above-mentioned stone tablet image is implemented.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the digital reconstruction method of the above-mentioned stone tablet image is implemented.
[0051] The above solution of the present application has the following beneficial effects:
[0052] In some embodiments of the present application, by obtaining a plurality of stele images, extracting the corner point map and character content description information of each stele image, then respectively for each stele image, obtaining the shallow features of the stele image according to the stele image and the corner point map of the stele image, performing deep feature extraction on the shallow features to obtain the deep features of the stele image, then respectively for each deep feature, predicting the deep feature to obtain the predicted characters of the stele image corresponding to the deep feature, and obtaining the image prior knowledge of the predicted characters, then using a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stele image to obtain the final features of each stele image, then using a reconstruction model to reconstruct each final feature to obtain the reconstructed image of the stele image corresponding to each final feature, then constructing a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructing a boundary loss function based on all the final features, and using the reconstruction loss function, pixel loss function, and boundary loss function to optimize the fusion model and the reconstruction model to obtain an optimized fusion model and an optimized reconstruction model, and finally using the optimized fusion model and the optimized reconstruction model to reconstruct the stele image to be reconstructed to obtain the final reconstructed image of the stele image to be reconstructed. Among them, obtaining the corner point map of the stele image can provide the corner point information of the stele image, improve the information richness of the deep features obtained based on the corner point information, and further improve the accuracy of the obtained image prior knowledge. The accuracy and information richness of the final features fused according to the accurate image prior knowledge, character content description information, and shallow features are improved, and the accuracy of the reconstructed image obtained by reconstructing based on the final features with high accuracy and richness is improved.
[0053] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0055] Figure 1 It is a flowchart of a digital reconstruction method for a stele image provided by an embodiment of the present application;
[0056] Figure 2 It is a schematic structural diagram of a digital reconstruction device for a stele image provided by an embodiment of the present application;
[0057] Figure 3 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed Implementation Manner
[0058] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present application.
[0059] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0060] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0061] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0062] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0063] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0064] In view of the problem of low accuracy in the reconstruction of stone tablet images in the prior art, the embodiments of the present application provide a digital reconstruction method for stone tablet images. The digital reconstruction method obtains multiple stone tablet images, extracts the corner point map and character content description information of each stone tablet image, and then for each stone tablet image, according to the stone tablet image and the corner point map of the stone tablet image, obtains the shallow features of the stone tablet image, and performs deep feature extraction on the shallow features to obtain the deep features of the stone tablet image. Then, for each deep feature, predicts the deep feature to obtain the predicted characters of the stone tablet image corresponding to the deep feature, and obtains the image prior knowledge of the predicted characters. Then, uses a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stone tablet image to obtain the final features of each stone tablet image. Then, uses a reconstruction model to reconstruct each final feature to obtain the reconstructed image of the stone tablet image corresponding to each final feature. Then, constructs a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructs a boundary loss function based on all the final features, and uses the reconstruction loss function, the pixel loss function, and the boundary loss function to optimize the fusion model and the reconstruction model to obtain an optimized fusion model and an optimized reconstruction model. Finally, uses the optimized fusion model and the optimized reconstruction model to reconstruct the stone tablet image to be reconstructed to obtain the final reconstructed image of the stone tablet image to be reconstructed. Among them, obtaining the corner point map of the stone tablet image can provide the corner point information of the stone tablet image, improve the information richness of the deep features obtained based on the corner point information, and further improve the accuracy of the obtained image prior knowledge. The accuracy and information richness of the final features fused according to the accurate image prior knowledge, character content description information, and shallow features are improved, and the accuracy of the reconstructed image obtained by reconstructing based on the final features with high accuracy and richness is improved.
[0065] Next, an exemplary description will be given of the digital reconstruction method for stone tablet images provided by the present application.
[0066] As Figure 1 shown, the digital reconstruction method for stone tablet images provided by the present application includes the following steps:
[0067] Step 11, obtain multiple stone tablet images, and extract the corner point map and character content description information of each stone tablet image.
[0068] The above-mentioned corner point map is an image in which the corner points of all characters in the stone tablet image are marked, that is, a stone tablet image with corner point markings. The above-mentioned stone tablet image is an image of a stone tablet, such as an image of a historical commemorative stone tablet, an architectural commemorative stone tablet, a calligraphy stone tablet, etc. The character content description information is used to describe the content recorded by all characters in the stone tablet image. For example, for the stone tablet image of an architectural commemorative stone tablet, the character content description information may be "This stone tablet records the completion time of the building, the name of the investor, the time and cost consumed."
[0069] In some embodiments of the present application, devices such as cameras can be used to obtain stone tablet images. To improve the diversity of stone tablet images, data augmentation techniques such as rotation, scaling, and cropping can be employed to process the directly obtained stone tablet images and generate multiple stone tablet images. Algorithms for detecting corner points, such as the Harris corner detection algorithm, can be used to obtain the corner point map of each stone tablet image. The character content description information of the stone tablet image can be extracted manually.
[0070] It should be noted that after obtaining the character content description information, preprocessing can be performed on the character content description information. For example, the jieba word segmentation tool can be used for word segmentation to split the text description into a sequence of words. Then, according to the Harbin Institute of Technology stop word list, the words in the word sequence are screened, and the words in the stop word list are deleted and filtered. Finally, based on the part-of-speech tags marked by the jieba tool during word segmentation and compared with the custom reserved part-of-speech table, the words whose part-of-speech is not in the reserved part-of-speech table are deleted and filtered.
[0071] It is worth mentioning that by obtaining the corner point map and character content description information of the stone tablet image, the corner point information of the stone tablet image and the information of the recorded content can be provided, improving the richness of the information.
[0072] Step 12: For each stone tablet image, based on the stone tablet image and the corner point map of the stone tablet image, obtain the shallow features of the stone tablet image, and perform deep feature extraction on the shallow features to obtain the deep features of the stone tablet image.
[0073] In some embodiments of the present application, the steps of obtaining the shallow features of the stone tablet image based on the stone tablet image and the corner point map of the stone tablet image, and performing deep feature extraction on the shallow features to obtain the deep features of the stone tablet image include:
[0074] The first step: Based on the stone tablet image and the corner point map of the stone tablet image, obtain the shallow features of the stone tablet image.
[0075] Specifically, through the formula:
[0076] R m,2 ,C 2 =UNet(R m,1 ,C 1 )
[0077] R m,1 ,C 1 =ResNet(R m,0 ,C 0 )
[0078] R m,0 ,C 0 =Init(I stone,m ,Cor stone,m)
[0079] Obtain the shallow feature R of the m-th stele image m,2 。
[0080] Among them, C 2 、C 1 、C 0 all represent the number of channels of the feature map, UNet( ) represents the UNet network operation, R m,1 represents the secondary feature, ResNet( ) represents the ResNet network operation, R m,0 represents the initial feature, Init( ) represents the operations of the convolutional layer and the pooling layer, I stone,m represents the m-th stele image, Cor stone,m represents the corner point map of the m-th stele image, m = 1, 2,..., M, and M represents the total number of stele images.
[0081] Exemplarily, the above formula can be run using computer software such as matlab and python to obtain the shallow feature.
[0082] Step 2: Extract deep features from the shallow features to obtain the deep features of the stele image.
[0083] Exemplarily, Transformer can be used to extract deep features from the shallow features to obtain the deep features.
[0084] It is worth mentioning that the shallow features and deep features obtained based on the corner point map incorporate the corner point information of the stele image, which can improve the accuracy and information richness of the shallow features and deep features.
[0085] Step 13: For each deep feature, predict the deep feature to obtain the predicted character of the stele image corresponding to the deep feature, and obtain the image prior knowledge of the predicted character.
[0086] The above image prior knowledge is used to describe the image information generated according to the predicted character.
[0087] In some embodiments of the present application, a linear layer and a softmax layer can be used to gradually process the deep features to obtain the predicted character. The predicted character is the character in the stele image corresponding to the deep feature obtained by predicting the deep feature. The steps of obtaining the image prior knowledge of the predicted character include:
[0088] The contrastive language-image pre-training model (CLIP) can be used to vectorize the description information of each character content in existing literature. The text is decomposed into sub-word units through a tokenizer and mapped to their corresponding word embedding vectors to form a text input sequence. Since the CLIP model has powerful visual and semantic understanding capabilities and can learn the relationship between text and images, it can generate an image corresponding to each text input vector according to the information described by each text input sequence. The combination of all text input vectors and their corresponding images is saved in the same database to obtain a prior knowledge base.
[0089] Then, search for an image matching the predicted character in the prior knowledge base. For example, calculate the cosine similarity value between each text input vector in the prior knowledge base and the predicted character, select the text input vector corresponding to the smallest cosine similarity value, and use the image corresponding to this text input vector as the image matching the predicted character. Then, use a convolutional neural network, etc. to obtain the image feature vector of this image, and use this image feature vector as the image prior knowledge of the corresponding predicted character.
[0090] Step 14: Use a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stele image to obtain the final feature of each stele image.
[0091] In some embodiments of the present application, the step of using a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stele image to obtain the final feature of each stele image includes:
[0092] The fusion model performs the following steps for each stele image respectively:
[0093] The first step: fuse the shallow features corresponding to the stele image and the corner point map to obtain an initial fusion feature, and fuse the initial fusion feature and the image prior knowledge corresponding to the stele image to obtain a secondary fusion feature.
[0094] Exemplarily, a cross-attention mechanism can be used for fusion. That is, the corner point map is used as the query vector in the cross-attention mechanism, and the shallow features are used as the key vector and value vector. After cross-attention calculation, an initial fusion feature is obtained. Then, use the self-attention mechanism to calculate the attention score of the image prior knowledge, use this attention score as the query vector in the cross-attention mechanism, and use the initial fusion feature as the key vector and value vector. After cross-attention calculation, a secondary fusion feature is obtained.
[0095] In the second step, feature extraction is performed on the character content description information corresponding to the stele image to obtain text features, and the text features and the secondary fusion features are fused to obtain the final features of the stele image.
[0096] Exemplarily, an encoder can be used to extract features from the character content description information to obtain text features, and a cross-attention mechanism can be used to fuse the text features and the secondary fusion features. That is, the secondary fusion features are used as the query vector of the cross-attention mechanism, and the text features are used as the key vector and the value vector, and the final features are obtained after cross-attention calculation.
[0097] The above cross-attention calculation is as follows:
[0098]
[0099] Among them, Q i represents the query vector, K t represents the key vector, V t represents the value vector, I represents the input used as the query vector, T represents the input used as the key vector and the value vector, and both represent weights, and d k represents the dimension.
[0100] It can be understood that the structure of the fusion model is a self-attention module, a feature extraction module, and three sequentially connected fusion modules. The input end of the self-attention module inputs the image prior knowledge, the output end of the self-attention module is connected to the input end of the second fusion module, the input end of the feature extraction module inputs the character content description information, the output end of the feature extraction module is connected to the input end of the third fusion module, the input of the first fusion module is the corner point map and the shallow features, and the output is the initial fusion features. The input of the second fusion module is the initial fusion features and the attention scores obtained by using the self-attention module, and the output is the secondary fusion features. The input of the third fusion module is the text features and the secondary fusion features obtained by using the feature extraction module, and the output is the final features. The calculation mechanism in each fusion module is related to the actual fusion mechanism used. If the cross-attention mechanism is used for fusion in all three fusion modules, it can be considered that the calculation formula of the cross-attention mechanism is the expression of the three fusion modules.
[0101] It is worth mentioning that the accuracy and information richness of the final features fused according to accurate image prior knowledge, character content description information, and shallow features are improved.
[0102] Step 15, use the reconstruction model to reconstruct each final feature to obtain the reconstructed image of the stele image corresponding to each final feature.
[0103] Specifically, the reconstruction model passes through the formula:
[0104] R m,SR = Pix(R m,2 + RL m,cross )
[0105] Obtain the reconstructed image R of the m-th stele image m,SR .
[0106] Wherein, Pix( ) represents upsampling, and R m,2 represents the shallow feature of the m-th stele image, and RL m,cross represents the final feature of the m-th stele image, m = 1, 2,..., M, and M represents the total number of stele images.
[0107] It can be understood that the above formula for obtaining the reconstructed image is the expression of the reconstruction model.
[0108] Exemplarily, the above formula can be run using computer software such as matlab and python to obtain the reconstructed image.
[0109] It is worth mentioning that the accuracy of the reconstructed image obtained by reconstruction based on the final feature with high accuracy and richness is improved.
[0110] Step 16, construct a reconstruction loss function and a pixel loss function based on all the reconstructed images, construct a boundary loss function based on all the final features, and optimize the fusion model and the reconstruction model using the reconstruction loss function, the pixel loss function, and the boundary loss function to obtain an optimized fusion model and an optimized reconstruction model.
[0111] The above reconstruction loss function and pixel loss function are used to describe the accuracy of all the reconstructed images, and the boundary loss function is used to describe the accuracy of all the final features.
[0112] In some embodiments of the present application, the steps of constructing a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructing a boundary loss function based on all the final features, and optimizing the fusion model and the reconstruction model using the reconstruction loss function, the pixel loss function, and the boundary loss function to obtain an optimized fusion model and an optimized reconstruction model include:
[0113] The first step is to construct a reconstruction loss function and a pixel loss function based on all the reconstructed images.
[0114] First, perform text recognition on each reconstructed image to obtain the text recognition result of each reconstructed image, and construct a reconstruction loss function based on all the text recognition results.
[0115] Specifically, the reconstruction loss function is:
[0116] l rec= 1 / m ∑(pre sr,m - pre hr,m ) 2
[0117] where L rec represents the value of the reconstruction loss function, pre sr,m represents the text recognition result of the m-th reconstructed image, and pre hr,m represents the true text label of the m-th stone tablet image.
[0118] Then, a pixel loss function is constructed based on all the reconstructed images.
[0119] Specifically, the pixel loss function is:
[0120]
[0121] where l 2 represents the value of the pixel loss function, represents the pixel value of the i-th pixel point of the m-th reconstructed image, and represents the pixel value of the i-th pixel point of the m-th stone tablet image.
[0122] Exemplarily, a convolutional recurrent neural network can be used to perform text recognition on the reconstructed images.
[0123] It should be noted that the above true text label is used to describe the true text of the stone tablet corresponding to the stone tablet image.
[0124] Second step, a boundary loss function is constructed based on all the final features.
[0125] Specifically, the boundary loss function is:
[0126]
[0127] where l box represents the value of the boundary loss function, t m represents the bounding box parameters obtained according to the final features corresponding to the m-th stone tablet image, and represents the true bounding box parameters of the m-th stone tablet image.
[0128] Third step, the fusion model and the reconstruction model are optimized using the reconstruction loss function, the pixel loss function, and the boundary loss function to obtain the optimized fusion model and the optimized reconstruction model.
[0129] Specifically, determine whether the values of the reconstruction loss function, the pixel loss function, and the boundary loss function are all less than the preset loss function value. If so, use the fusion model as the optimized fusion model and the reconstruction model as the optimized reconstruction model. Otherwise, adjust the parameters of the fusion model and the reconstruction model, and return to the step of using the fusion model to fuse the shallow features, the image prior knowledge, and the character content description information corresponding to each stele image to obtain the final features of each stele image.
[0130] Exemplarily, the preset loss function value is 1.8, and the values of the reconstruction loss function, the pixel loss function, and the boundary loss function are 2.1, 1.3, and 1.5 respectively. Among them, the value of the reconstruction loss function is greater than the preset loss function value. Then, adjust the parameters of the fusion model and the reconstruction model, and return to the step of using the fusion model to fuse the shallow features, the image prior knowledge, and the character content description information corresponding to each stele image to obtain the final features of each stele image. Re-obtain the values of the reconstruction loss function, the pixel loss function, and the boundary loss function. At this time, the values are 1.4, 1.2, and 1.6 respectively, all of which are less than 1.8. Then, use the fusion model at this time as the optimized fusion model and the reconstruction model as the optimized reconstruction model.
[0131] Step 17, use the optimized fusion model and the optimized reconstruction model to reconstruct the stele image to be reconstructed, and obtain the final reconstructed image of the stele image to be reconstructed.
[0132] The above stele image to be reconstructed is the stele image for which image reconstruction is required.
[0133] Specifically, extract the corner point map and the character content description information of the stele image to be reconstructed, then use the stele image to be reconstructed and the corner point map to obtain the shallow features and the deep features, then obtain the image prior knowledge based on the deep features, then use the optimized fusion model to fuse the shallow features, the image prior knowledge, and the character content description information to obtain the final features of the stele image to be reconstructed, and finally use the optimized reconstruction model to reconstruct the final features to obtain the final reconstructed image of the stele image to be reconstructed.
[0134] Exemplarily, computer software such as Matlab and Python can be used to run the above steps to obtain the final reconstructed image of the stele image to be reconstructed.
[0135] It is worth mentioning that obtaining the corner point map of the stone tablet image can provide the corner point information of the stone tablet image, improve the information richness of the deep features obtained based on the corner point information, and further improve the accuracy of the obtained image prior knowledge. According to the accurate image prior knowledge, character content description information, and the accuracy and information richness of the final features obtained by fusing the shallow features are improved. Based on the final features with high accuracy and richness, the accuracy of the reconstructed image obtained by reconstruction is improved.
[0136] The following is an exemplary description of the digital reconstruction device for the stone tablet image provided in this application.
[0137] As Figure 2 shown, an embodiment of this application provides a digital reconstruction device for a stone tablet image. The digital reconstruction device 200 includes:
[0138] An acquisition module 201, which acquires a plurality of stone tablet images and extracts the corner point map and character content description information of each stone tablet image; the corner point map is an image in which the corner points of all characters in the stone tablet image are marked;
[0139] A feature extraction module 202, which respectively obtains the shallow features of the stone tablet image according to the stone tablet image and the corner point map of the stone tablet image, and performs deep feature extraction on the shallow features to obtain the deep features of the stone tablet image;
[0140] A prediction module 203, which respectively predicts each deep feature to obtain the predicted characters of the stone tablet image corresponding to the deep feature, and acquires the image prior knowledge of the predicted characters; the image prior knowledge is used to describe the image information generated according to the predicted characters;
[0141] A fusion module 204, which uses a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stone tablet image to obtain the final feature of each stone tablet image;
[0142] A first reconstruction module 205, which uses a reconstruction model to reconstruct each final feature to obtain the reconstructed image of the stone tablet image corresponding to each final feature;
[0143] An optimization module 206, which constructs a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructs a boundary loss function based on all the final features, and uses the reconstruction loss function, pixel loss function, and boundary loss function to optimize the fusion model and the reconstruction model to obtain an optimized fusion model and an optimized reconstruction model; the reconstruction loss function and the pixel loss function are used to describe the accuracy of all the reconstructed images, and the boundary loss function is used to describe the accuracy of all the final features;
[0144] The second reconstruction module 207 uses the optimized fusion model and the optimized reconstruction model to reconstruct the stone tablet image to be reconstructed, and obtains the final reconstructed image of the stone tablet image to be reconstructed.
[0145] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0146] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details are not described herein again.
[0147] As Figure 3 shown, an embodiment of the present application provides a terminal device. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 3 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the steps in any of the above method embodiments.
[0148] Specifically, when the processor D100 executes the computer program D102, it obtains multiple stele images, extracts the corner map and character content description information of each stele image, and then for each stele image, based on the stele image and its corner map, obtains the shallow features of the stele image, performs deep feature extraction on the shallow features to obtain the deep features of the stele image, then for each deep feature, makes predictions on the deep feature to obtain the predicted characters of the stele image corresponding to the deep feature, and obtains the image prior knowledge of the predicted characters. Then, it uses a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each stele image to obtain the final features of each stele image, and then uses a reconstruction model to reconstruct each final feature to obtain the reconstructed image of the stele image corresponding to each final feature. Then, it constructs a reconstruction loss function and a pixel loss function based on all the reconstructed images, constructs a boundary loss function based on all the final features, and uses the reconstruction loss function, pixel loss function, and boundary loss function to optimize the fusion model and the reconstruction model to obtain an optimized fusion model and an optimized reconstruction model. Finally, it uses the optimized fusion model and the optimized reconstruction model to reconstruct the stele image to be reconstructed to obtain the final reconstructed image of the stele image to be reconstructed. Among them, obtaining the corner map of the stele image can provide the corner information of the stele image, improve the information richness of the deep features obtained based on the corner information, and further improve the accuracy of the obtained image prior knowledge. The accuracy and information richness of the final features fused according to the accurate image prior knowledge, character content description information, and shallow features are improved, and the accuracy of the reconstructed image obtained by reconstructing based on the final features with high accuracy and richness is improved.
[0149] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0150] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of the computer program. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0151] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0152] An embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device can implement the steps in the above method embodiments when executed.
[0153] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the digital reconstruction method device / terminal device of the stone tablet image, a recording medium, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0154] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0155] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0156] The above is the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for digital reconstruction of a stone tablet image, characterized in that: include: Acquire multiple stone tablet images, and extract corner point maps and character content description information of each stone tablet image; The corner point map is an image in which the corner points of all characters in the stele image are marked; For each of the stone tablet images, respectively, shallow features of the stone tablet image are obtained according to the stone tablet image and the corner point map of the stone tablet image, and deep features are extracted from the shallow features to obtain deep features of the stone tablet image; For each of the deep features, the deep features are predicted to obtain predicted characters of the stone tablet image corresponding to the deep features, and image prior knowledge of the predicted characters is obtained; the image prior knowledge is used to describe image information generated according to the predicted characters; Using a fusion model to fuse shallow features, image prior knowledge, and character content description information corresponding to each of the stele images, to obtain final features of each of the stele images; Reconstructing each of the final features using a reconstruction model to obtain a reconstructed image of the stone tablet image corresponding to each of the final features; A reconstruction loss function and a pixel loss function are constructed based on all reconstructed images, a boundary loss function is constructed based on all final features, and the fusion model and the reconstruction model are optimized using the reconstruction loss function, the pixel loss function and the boundary loss function to obtain an optimized fusion model and an optimized reconstruction model; the reconstruction loss function and the pixel loss function are used to describe the accuracy of all reconstructed images, and the boundary loss function is used to describe the accuracy of all final features; The stone tablet image to be reconstructed is reconstructed using the optimized fusion model and the optimized reconstruction model to obtain a final reconstructed image of the stone tablet image to be reconstructed.
2. The digital reconstruction method according to claim 1, characterized in that: The obtaining of shallow features of the stone tablet image according to the stone tablet image and the corner point map of the stone tablet image comprises: By formula: R m,2 ,C2=UNet(R m,1 ,C1) R m,1 ,C1=ResNet(R m,0 ,C0) R m,0 ,C0=Init(I stone,m ,Cor stone,m ) Get the shallow feature R of the mth stele image m,2 ; Among them, C2, C1, and C0 all represent the number of feature map channels, UNet () represents the UNet network operation, and R m,1 Represents secondary features, ResNet( ) represents ResNet network operation, R m,0 Init() represents the initial features, and convolution and pooling operations. stone,m represents the mth stele image, Cor stone,m represents the corner point map of the mth stele image, m=1, 2, ..., M, and M represents the total number of stele images.
3. The digital reconstruction method according to claim 1, characterized in that: The fusion model is used to fuse the shallow features, image prior knowledge, and character content description information corresponding to each of the stele images to obtain the final features of each of the stele images, including: For each of the stele images, perform the following steps: The shallow features and the corner map corresponding to the stele image are fused to obtain an initial fused feature, and the initial fused feature is fused with the image prior knowledge corresponding to the stele image to obtain a secondary fused feature; Feature extraction is performed on the character content description information corresponding to the stele image to obtain text features, and the text features are fused with the secondary fusion features to obtain final features of the stele image.
4. The digital reconstruction method according to claim 1, characterized in that: The step of reconstructing each of the final features using the reconstruction model to obtain a reconstructed image of the stone tablet image corresponding to each of the final features includes: By formula: R m,SR =Pix(R m,2 +RL m,cross ) Get the reconstructed image R of the mth stele image m,SR ; Among them, Pix( ) represents upsampling, R m,2 represents the shallow features of the m-th stele image, RL m,cross represents the final feature of the mth stele image, m=1, 2, ..., M, and M represents the total number of stele images.
5. The digital reconstruction method according to claim 4, characterized in that: The reconstruction loss function and the pixel loss function are constructed based on all the reconstructed images, including: Performing text recognition on each of the reconstructed images to obtain a text recognition result for each of the reconstructed images, and constructing a reconstruction loss function based on all the text recognition results; A pixel loss function is constructed based on all reconstructed images.
6. The digital reconstruction method according to claim 5, characterized in that: The reconstruction loss function is: the rec =1 / m∑(before sr,m -pre hr,m ) 2 Among them, L rec Represents the value of the reconstruction loss function, pre sr,m represents the text recognition result of the mth reconstructed image, pre hr,m represents the true text label of the m-th stele image; The pixel loss function is: Among them, l2 represents the value of the pixel loss function, represents the pixel value of the i-th pixel point of the m-th reconstructed image, Represents the pixel value of the i-th pixel point of the m-th stone tablet image.
7. The digital reconstruction method according to claim 1, characterized in that: The boundary loss function is: Among them, l box represents the value of the boundary loss function, t m represents the bounding box parameters obtained according to the final features corresponding to the m-th stele image, represents the true bounding box parameters of the mth stele image, and M represents the total number of stele images.
8. A digital reconstruction device for a stone tablet image, characterized in that: include: An acquisition module, which acquires a plurality of stele images and extracts a corner point map and character content description information of each stele image; The corner point map is an image in which the corner points of all characters in the stele image are marked; A feature extraction module, for each of the stele images, obtains shallow features of the stele image according to the stele image and the corner map of the stele image, and performs deep feature extraction on the shallow features to obtain deep features of the stele image; A prediction module predicts each of the deep features, obtains the predicted characters of the stone tablet image corresponding to the deep features, and acquires image prior knowledge of the predicted characters; the image prior knowledge is used to describe image information generated according to the predicted characters; A fusion module, which uses a fusion model to fuse the shallow features, image prior knowledge, and character content description information corresponding to each of the stone tablet images to obtain the final features of each of the stone tablet images; A first reconstruction module reconstructs each of the final features using a reconstruction model to obtain a reconstructed image of the stone tablet image corresponding to each of the final features; An optimization module, constructing a reconstruction loss function and a pixel loss function based on all reconstructed images, constructing a boundary loss function based on all final features, and optimizing the fusion model and the reconstruction model using the reconstruction loss function, the pixel loss function and the boundary loss function to obtain an optimized fusion model and an optimized reconstruction model; the reconstruction loss function and the pixel loss function are used to describe the accuracy of all reconstructed images, and the boundary loss function is used to describe the accuracy of all final features; The second reconstruction module reconstructs the to-be-reconstructed stone tablet image by using the optimized fusion model and the optimized reconstruction model to obtain a final reconstructed image of the to-be-reconstructed stone tablet image.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for digitally reconstructing a stone tablet image as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for digitally reconstructing a stone tablet image as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for extracting stele inscription digital rubbing based on three-dimensional data scanning
CN104268924A
Font image generation model training method and device, electronic equipment and medium
CN118297790A