Method for realizing digitalization of cultural resources based on image recognition
By combining a style transfer model based on generative adversarial networks and a DB segmentation algorithm with knowledge graph-based text recognition optimization, the problem of high text recognition difficulty in ancient book handwriting images is solved, achieving efficient and accurate ancient book text recognition and database construction.
Patent Information
- Application Number
- CN202511150753.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-18
AI Technical Summary
In existing technologies, the handwriting style of ancient books and documents makes text recognition in images difficult and hard to achieve ideal recognition results.
A style transfer model based on generative adversarial networks is adopted, and temporal feature constraints are introduced. Standard font style images are generated through content encoder and style encoder. The DB segmentation algorithm and text recognition model are combined, and knowledge graph is used to optimize the recognition results.
It effectively reduced the recognition difficulty caused by the handwriting style, improved the accuracy of ancient text recognition, and constructed a high-quality cultural resource database.
Smart Images

Figure CN120747986B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of image processing, in particular, to a cultural resource digitization implementation method based on image recognition. BACKGROUND
[0002] With the development of information technology, people have higher requirements for the protection, inheritance and utilization of cultural resources, and cultural resource digitization has become an important means to meet these needs. Digitization can achieve permanent preservation, wide dissemination and in-depth research of cultural resources, providing strong support for the inheritance and development of culture.
[0003] Ancient books are an important carrier of Chinese civilization, and it is urgent to strengthen the protection of ancient books through digital technology. Using modern technologies such as document scanning, OCR recognition, and intelligent analysis for ancient book digitization has important significance for improving the quality of ancient book work and promoting ancient book work.
[0004] However, in related technologies, for the collected image to be recognized, since ancient books are often in written form, it increases the difficulty of recognizing the text content of the image, and often it is difficult to achieve the ideal recognition effect. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide a cultural resource digitization implementation method based on image recognition to at least solve the technical problem of high recognition difficulty in related technologies for image text recognition.
[0006] To achieve the above purpose, the embodiments of the present application provide the following technical solutions.
[0007] According to an embodiment of the present application, a cultural resource digitization implementation method based on image recognition is provided;
[0008] including the following steps:
[0009] A style transfer model based on a generative adversarial network is built, and a time sequence feature constraint is introduced in the style transfer model, which is constructed based on the layout time sequence of the ancient book text and used to constrain the text time sequence consistency in the style transfer process;
[0010] The content encoder extracts the text structure of the input image, the style encoder converts the features in the pre-constructed style feature library into a style vector, and the transfer decoder generates a text image fused with the target style in combination with the text structure, the style vector and the time sequence feature constraint;
[0011] The trained style transfer model is used to generate a stylized training data set, and the data set is associated with the information label of the ancient book resource;
[0012] The text recognition model is used for text recognition on the first text segment and the second text segment respectively, the overlapping text segment between the tail of the first text segment and the head of the second text segment is obtained, the recognition result of the overlapping text segment is compared with the preset knowledge graph, the text recognition model is optimized through semantic conflict detection, and the text recognition model is trained by using a gradual style fusion mechanism and introducing a stylized training data set, so that the feature vector output by the model and the semantic features of the knowledge graph are bidirectionally matched;
[0013] The new image is subjected to dynamic style adaptation processing, the text style features are extracted and matched with the style types in the style feature library, the style transfer model is called to generate a verification image, the verification image and the new image are input into the trained text recognition model, the text recognition result of the current image to be recognized is obtained, after the text recognition of the image data set is completed, a cultural resource database containing the text recognition result is constructed and output.
[0014] Preferably, before the text recognition model is used for text recognition on the first text segment and the second text segment respectively, the following steps are further included:
[0015] Based on text edge positioning, non-text area filtering and layout positioning;
[0016] The text candidate area in the target image is extracted, a DB segmentation algorithm introducing resource characteristics is used for text segmentation on the text candidate area, and a text line image with a feature label is obtained;
[0017] The text line image is subjected to segmentation processing, and the first text segment and the second text segment with different lengths are obtained.
[0018] Preferably, the text candidate area in the target image is obtained through text edge positioning and non-text area filtering processing, wherein:
[0019] The text edge positioning and the non-text area filtering both have a learnable weight parameter, the weight parameter is dynamically adjusted based on the characteristics of the ancient book carrier, the outputs of the parts are multiplied by the weight parameter and then added element by element, and is represented as:
[0020]
[0021] In the formula, The text candidate area is represented, ReLU represents an activation function, and LayerNorm represents a normalization operation; 、 The weight parameter dynamically adjusted based on the ancient book carrier material m is represented as and The text edge positioning processing and the non-text area filtering processing are represented respectively, and X represents the target image;
[0022] In the text edge positioning branch, the target image is taken as the input of the text edge positioning branch, multiplied by the text edge positioning branch weight parameters after passing through the 1*5 convolution, 5*1 convolution and residual connection feature extraction module to obtain the edge result output;
[0023] In the non-text area filtering branch, the feature information of the text area is enhanced by the inversion of the background feature, the same feature extraction module as the text edge positioning branch is used, and the feature elements are mapped to [0, 1] through the Sigmoid function, and then the text feature is enhanced by taking the inverse way, which is represented as:
[0024]
[0025] In the formula, represents the original output of the feature extraction module mapped to [0, 1], represents the non-text area filtering branch processing of the image X, ReLU represents the activation function, and LayerNorm represents the normalization operation.
[0026] Preferably, the step of performing text segmentation on the text candidate area by using the DB segmentation algorithm introducing the resource feature to obtain the text line image with the feature label includes:
[0027] A DB model containing a deformable convolution DCNv2 module is constructed, a scroll curvature factor of ancient books is introduced in the deformable convolution DCNv2 module, and the spatial sensitivity of the convolution to the vertical text scroll area is controlled through the curvature perception strength, wherein the scroll curvature factor of ancient books is determined based on the carrier material feature related physical form parameters of the target image;
[0028] The page edge curl mask generated by the ancient book 3D scanning data is collected, the ancient book 3D scanning data is obtained through the carrier material feature related to the target image, and the page edge curl mask is generated in combination with the edge features of the text candidate area to segment the target image to obtain a text area mask;
[0029] According to the text area mask, a contour detection method based on Canny edge detection and ancient book format features is used to segment the text area into multiple text lines, and the ancient book format features include text edge positioning features and format structure information retained after non-text area filtering;
[0030] According to the boundary box of the segmented text line, a plurality of text line images are extracted from the image, and the feature label of the text line image contains the spatial position information of the corresponding text line in the target image and the associated ancient book carrier material feature parameters.
[0031] Preferably, in the DCNv2 module of the DB model, the DCNv2 module comprises a deformable convolution layer, batch normalization and a Hardswish activation function;
[0032] In the deformable convolution layer, curvature perception strength is introduced, the curvature perception strength is positively correlated with the ancient book wrinkle curvature factor, and the curvature perception strength is used for controlling the sensitivity of the deformable convolution to spatial changes;
[0033] In the Hardswish activation function, an ancient book ink aging coefficient is introduced , which is expressed as:
[0034]
[0035] Wherein, x represents a feature value input into the Hardswish activation function, represents the ancient book ink aging coefficient.
[0036] Preferably, the text recognition model is used to perform text recognition on the first text segment and the second text segment respectively, the recognition result of the coincident text segment obtained is compared with a preset knowledge graph, and the step of optimizing the text recognition model through semantic conflict detection comprises:
[0037] The recognition result of the coincident text segment in the first text segment recognized by the text recognition model is taken as a first recognition result;
[0038] The recognition result of the coincident text segment in the second text segment recognized by the text recognition model is taken as a second recognition result;
[0039] When the first recognition result and the second recognition result are different, a first difference rate is calculated, the length of the coincident text segment of the next text line image segmentation is increased, and a second difference rate is calculated;
[0040] In the calculation of the first difference rate and the second difference rate, an ancient book homophone correlation weight is introduced, wherein the ancient book homophone correlation weight is determined based on the homophone corresponding relationship recorded in the preset knowledge graph; when the second difference rate is greater than the first difference rate, the attention weight of the model to the annotation text is adjusted in combination with the ink difference positioning error source of the ancient book annotation and the main text, the text recognition model is optimized and adjusted according to the adjusted attention weight, and the new text segment image is recognized based on the optimized and adjusted text recognition model.
[0041] Preferably, the step of optimizing and adjusting the text recognition model comprises:
[0042] The Bayesian optimization is used to set and optimize an optimal hyperparameter combination of the text recognition model, and an optimization target is to minimize an error; wherein, the hyperparameter combination includes an attention weight parameter, and the attention weight parameter is input as an initial prior value of the Bayesian optimization;
[0043] In the optimization process, a Gaussian process is used as a probability agent model to fit a target function, wherein, the target function of the Gaussian process incorporates a variant processing rule of the ancient book collation; and the error-free calculation of the target function includes a difference between recognition results before and after attention weight adjustment; a verification information construction function based on the probability agent model is constructed, and the verification information construction function is represented as:
[0044]
[0045] In the formula, f(x) represents a target value of the Gaussian process, including a model output after attention weight optimization, represents a current optimal target value in the Gaussian process, represents a Gaussian distribution cumulative density function, represents a balance parameter, represents a mean value of the target function, represents a variance of the target function; represents an ancient book version coefficient.
[0046] Preferably, the text recognition model includes a recognition encoder and a recognition decoder.
[0047] The recognition encoder is used to extract features from a text segment image, including a convolution layer and a bidirectional LSTM layer, and an attention mechanism network is introduced in the bidirectional LSTM layer, and the attention mechanism network includes layer normalization and double attention, and residual connection is used between the two.
[0048] The recognition decoder includes a full connection layer, and is used to decode the extracted features into a text sequence.
[0049] Preferably, in the recognition encoder,
[0050] An ancient book ink layer extraction module is introduced in the convolution layer, and a difference in ink concentration between a main text and a note is captured through multi-spectral imaging data, and the convolution layer also adopts an inception structure of a GoogLeNet model, including an inception1 module, an inception2 module and an inception3 module.
[0051] The bidirectional LSTM layer includes a forward LSTM module, a backward LSTM module and an attention mechanism module, the forward LSTM module processes a text sequence in a top-to-bottom vertical order, and the backward LSTM module fuses a sentence reading symbol feature of an ancient book.
[0052] The forward LSTM module is used for forward LSTM unit processing on the input feature vector sequence, and the output of each time step is a hidden layer vector.
[0053] The backward LSTM module is used for backward LSTM unit processing on the input feature vector sequence, and the output of each time step is a hidden layer vector.
[0054] The attention mechanism module is used to calculate attention weights from the outputs of the forward LSTM and the backward LSTM, respectively. For each time step, there are two attention weights, which correspond to the forward and backward LSTMs, respectively. After calculating the attention weights, the weights are normalized by layer normalization. The normalized attention weights are connected in residual to the original hidden state. For each time step, the normalized attention weights are added to the corresponding forward and backward hidden states, and the hidden state after residual connection is used for weighted summation to obtain the final text feature.
[0055] In the full connection layer of the recognition decoder, the text feature is mapped to the category space through the full connection layer, and the category probability is output using the softmax function. The recognition result of the text feature is determined based on the category probability, which is represented as:
[0056]
[0057] In the formula, represents a bias term; represents a weight matrix, represents the mapped text feature, represents the mapped text feature.
[0058] Preferably, the bidirectional LSTM layer introduces vertical text timing attention, including:
[0059] A vertical text timing attention weight calculation module is constructed,
[0060] Based on the line sequence priority and the inter-character topological relationship of the vertical text of the ancient book, a timing attention mask is generated, wherein the line sequence priority is dynamically adjusted by the ratio of the column spacing to the line height in the ancient book format, and the inter-character topological relationship is represented by the deviation amount of the barycentric coordinates of adjacent characters.
[0061] An ancient book line style feature coefficient θ is introduced in the timing attention weight calculation, which is determined in combination with the line spacing specification, the character spacing ratio and the format layout features of the ancient book, and its calculation formula is represented as:
[0062]
[0063] In the formula, and represent weight coefficients, + =1; denotes the standard line spacing of the vertical text of the ancient book, denotes the actual measured line spacing; denotes the average single-character width, denotes the average single-character height; denotes the tilt angle of the page of the ancient book, and k denotes a tilt influence factor;
[0064] The hidden layer vectors output by the forward LSTM module and the hidden layer vectors output by the backward LSTM module of the bidirectional LSTM layer are spliced to obtain a fusion vector sequence, the time sequence attention mask is multiplied element by element with the fusion vector sequence to obtain a feature vector with a time sequence weight; the feature vector with the time sequence weight is input into the output layer of the bidirectional LSTM layer, and after being added to the original fusion vector sequence through a residual connection and being subjected to layer normalization processing, a final time sequence enhanced feature vector is obtained, which is used for subsequent decoding and recognition of the text sequence.
[0065] Compared with the prior art, the cultural resource digitization implementation method based on image recognition has the following beneficial effects:
[0066] The generation adversarial network style transfer model with a time sequence feature constraint is introduced in the embodiment of the present application, which can sufficiently learn the stroke trend and character structure features of the ancient book writing style, and through learning a large number of ancient book samples, the writing style text can be converted into a more easily recognizable standard font style, effectively reducing the recognition difficulty caused by the writing style form; the embodiment can understand the time sequence logic of the writing style text, accurately capture the character features, and greatly improve the recognition accuracy of the ancient book writing style text;
[0067] In the image preprocessing stage, the DB segmentation algorithm with introduced resource characteristics is used to accurately segment the image in combination with the page edge curl mask generated based on the material characteristics of the ancient book carrier and the 3D scanning data, effectively remove background noise interference, and highlight the text area; in the text recognition model, by introducing the related feature parameters of the ancient book, the model can adapt to the feature changes of the ancient book text under different conditions, and stably output high-quality recognition results;
[0068] The present application closely combines text recognition and semantic understanding. In the recognition process, the recognition result is detected and optimized for semantic conflict by comparing with a preset knowledge graph and introducing ancient book homonym association weights; when the recognition model has different recognition results for a text segment, the semantic logic more consistent with the ancient book is obtained by using the homonym association weights and the semantic relationship in the knowledge graph, so as to correct the recognition error; at the same time, for the optimization of the recognition model, the Bayesian optimization combined with the processing rules of ancient book collation and variant text is used, so that the model continuously learns and adapts to the semantic characteristics of the ancient book text, and the recognition effect is further improved.
[0069] In summary, the embodiment of the present application realizes efficient and accurate text recognition, and can digitize image information of ancient literature and the like to construct a cultural resource database convenient for management and application. BRIEF DESCRIPTION OF DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application.
[0071] Figure 1 An implementation flowchart of the cultural resource digitization implementation method based on image recognition provided by the embodiment of the present application is shown in FIG. 1.
[0072] Figure 2 A sub-flowchart of the cultural resource digitization implementation method based on image recognition of the embodiment of the present application is shown in FIG. 2.
[0073] Figure 3 Another sub-flowchart of the cultural resource digitization implementation method based on image recognition of the embodiment of the present application is shown in FIG. 3.
[0074] Figure 4 A hardware structure block diagram of a computer terminal according to the cultural resource digitization implementation method based on image recognition of the embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION
[0075] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0076] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0077] According to the embodiment of the present application, the method embodiment of the cultural resource digitization method based on image recognition is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from here.
[0078] Figure 1 is a flowchart of the cultural resource digitization method based on image recognition according to the embodiment of the present application;
[0079] As shown in Figure 1 , the digitization method of the present application comprises the following steps:
[0080] S101: Build a style transfer model based on a generative adversarial network, introduce a time sequence feature constraint in the style transfer model, and construct the time sequence feature constraint based on the layout time sequence of the ancient text, which is used to constrain the text time sequence consistency in the style transfer process;
[0081] Between step S101, the acquired image is preprocessed first;
[0082] Specifically, the images in the collected image dataset are enhanced one by one; in one implementation, when the image is enhanced, it includes contrast enhancement, image noise reduction and image correction processing;
[0083] In one implementation of contrast enhancement, the original scanned image is first divided into multiple sub-regions of equal size, and the histogram of pixel gray value is calculated for each sub-region. Based on the histogram of each sub-region, the histogram equalization algorithm is used to recombine all sub-regions to obtain the contrast-enhanced image.
[0084] In one implementation of image noise reduction, the Wiener filter is used to process the contrast-enhanced image data to obtain the noise-reduced image.
[0085] In some embodiments, in step S101, the style transfer model built uses an improved GAN architecture, including a generator and a discriminator; wherein the generator includes an encoder, a style fusion layer and a decoder;
[0086] The encoder includes 8 layers of convolutional neural networks, each layer having a convolution kernel size of 3*3 and a step size of 2, and extracts text content features of the input ancient book image through BatchNormalization and LeakyReLU activation functions, the text content features including stroke contours and character structures; the style fusion layer introduces a target style vector, the style vector including Songti, Kaishu standard font and other features, and further realizes weighted fusion of the content features and the style features through feature matrix multiplication; the decoder adopts a symmetric transpose convolution structure to reconstruct the fused features into a stylized text image with a size consistent with that of the input image, and the output layer adopts a Tanh activation function to normalize pixel values to the range of [0, 255];
[0087] In addition, the discriminator in the embodiment of the present application includes 6 layers of convolutional layers, each layer having a convolution kernel size of 4*4 and a step size of 2, and the output layer adopts a Sigmoid activation function to determine whether the input image is a real stylized text image or a fake image generated by the generator;
[0088] In addition, a timing feature constraint is introduced in the style transfer model, and a timing feature constraint module is constructed based on the top-down and right-to-left layout timing of the ancient book vertical text.
[0089] In one implementation, the constructed timing feature constraint module performs text line detection on the input ancient book image, obtains the bounding box coordinates of each text line, and calculates a line order priority parameter; the single-character barycenter coordinates are extracted through connected component analysis, the barycenter deviation of adjacent characters is calculated, and a character-to-character topological relationship matrix is constructed; the timing consistency loss is introduced in the total loss function of the GAN; then the timing feature constraint module is embedded at the end of the decoder of the generator, the original timing features extracted by the encoder are received through the skip connection, and the generator parameters are optimized through back propagation before the generated image is output, so that the text line order and the character-to-character arrangement order remain unchanged after the style transfer; the style transfer model constructed in the embodiment can not only transfer the ancient book writing text into a standard printed text style, but also ensure the text line order and the character-to-character relationship and other layout logics unchanged through the timing feature constraint, thereby providing high-quality timing regular images for subsequent text recognition.
[0090] Please continue to refer to Figure 1 In some embodiments of the present application, the implementation method further includes the following steps:
[0091] S102: extracting the text structure of the input image through the content encoder, converting the features in the pre-constructed style feature library into a style vector through the style encoder, and generating a text image with a fused target style through the migration decoder in combination with the text structure, the style vector and the timing feature constraint;
[0092] In this embodiment, the content encoder adopts a deep residual network (ResNet-50) structure to extract features related to the structure of the text in the ancient book image; after processing by the content encoder, a feature matrix is obtained; to enhance the discriminability of the features, a global average pooling (GAP) is used to generate a structure feature vector before output, which is a compact representation of the text structure;
[0093] In addition, in this embodiment, for the problem of blurred ink in ancient books, an attention gate module is added after the residual block to strengthen the extraction of features in clear stroke areas and suppress the interference of background noise and faded areas by learning an ink concentration weight matrix.
[0094] Further, the style encoder of this embodiment adopts a feature mapping structure based on Transformer to realize the conversion of features in the style feature library to style vectors. In the construction of the style feature library, standard font image libraries of multiple representative styles of ancient book printed matter are pre-collected, and the global features of each style are extracted by a VGG-19 network to construct the style feature library.
[0095] The style encoder supports style recommendation based on the input ancient book image. By calculating the cosine similarity between the structure feature vector output by the content encoder and each style feature matrix in the style feature library, the most matched style is automatically selected, and a hybrid style vector is generated by weighted fusion to improve the flexibility of style transfer.
[0096] Further, in some embodiments of the present application, the migration decoder generates a target image in combination with the text structure, style vector and timing feature constraint. The specific implementation process includes: the outputs of the content encoder, style encoder and timing feature constraint module are combined to obtain a fused feature matrix; in addition, the migration decoder includes multiple up-sampling modules, each of which doubles the size of the feature map through transposed convolution, and introduces high-resolution features of the corresponding level of the content encoder through a skip connection to restore the local details of the text; in the last up-sampling module, the timing feature map is converted into a pixel-level weight mask through a spatial attention mechanism to adjust the line sequence area and inter-word gap of the generated image, ensuring that the layout timing of the generated image is consistent with the original ancient book.
[0097] In the output of step S102, a double-branch design is adopted, the main branch generates a stylized image through a Tanh activation function, and the auxiliary branch outputs a text region mask (a binary image) for positioning the effective area in subsequent text segmentation. The mask generation process incorporates timing feature constraints.
[0098] The text image generated by the migration decoder provided in this embodiment not only retains the original text structure of the ancient book, but also fuses the visual features of the target style, while ensuring that the layout logic such as line sequence and inter-word relationship remains unchanged through timing feature constraints.
[0099] Please continue to refer to Figure 1 The implementation method of the embodiment also includes the following steps:
[0100] S103: generating a stylized training data set using the trained style transfer model, the data set being associated with information labels of the ancient book resources;
[0101] In step S103, the text structure features of each image are extracted by the content encoder of the style transfer model, the preset style vector is output in combination with the style encoder, and the stylized text image is generated through the transfer decoder, so as to ensure that each original image corresponds to the output of multiple stylized text images, expand the diversity of the data set, and add multi-dimensional information labels to the generated stylized image, such as ancient book metadata labels, content feature labels and style attribute labels;
[0102] For example, in the embodiment of the present application, the ancient book metadata labels include book name, volume, and book writing year; the content feature labels include text line sequence and font topological relationship; and the style attribute labels include style type and style transfer confidence;
[0103] Please continue to refer to Figure 1 The implementation method of the embodiment also includes the following steps:
[0104] S104: performing text recognition on the first text segment and the second text segment respectively using a text recognition model, wherein the first text segment has a coincident text segment with the second text segment between the end of the first text segment and the beginning of the second text segment, the recognition result of the coincident text segment is compared with a preset knowledge graph, the text recognition model is optimized through semantic conflict detection, and the stylized training data set is introduced into the text recognition model using a gradual style fusion mechanism to train the text recognition model, so that the feature vector output by the model and the semantic features of the knowledge graph are bidirectionally matched;
[0105] S105: performing dynamic style adaptation processing on the new image, extracting text style features and matching the style types in the style feature library, calling the style transfer model to generate a verification image, inputting the verification image and the new image into the trained text recognition model at the same time to obtain the text recognition result of the current image to be recognized, and after completing the text recognition of the image data set, constructing a cultural resource database containing the text recognition result and outputting.
[0106] In step S104, before text recognition is performed on the first text segment and the second text segment respectively by using the text recognition model, the text candidate region in the target image is extracted based on text edge positioning, non-text region filtering and layout positioning, the text candidate region is subjected to text segmentation by using a DB segmentation algorithm with introduced resource characteristics, and a text line image with a feature label is obtained; the text line image is subjected to segmentation processing, and the first text segment and the second text segment with different lengths are obtained.
[0107] In one implementation, the present embodiment extracts the text candidate region in the enhanced image through text edge positioning and non-text region filtering, and then subjects the text candidate region to text segmentation by using a DB segmentation algorithm, so that automatic text region extraction reduces manual intervention, improves processing speed and efficiency; the DB segmentation algorithm can accurately segment the text line, and in addition, the segmented text line is subjected to segmentation processing, and the recognition results of the two text segments are compared, so that the model parameters can be optimized and the overall recognition accuracy can be improved;
[0108] Specifically, in the present embodiment, the text candidate region in the enhanced image is obtained through text edge positioning and non-text region filtering processing, and the text edge positioning and the non-text region filtering both have a learnable weight parameter, the weight parameter is dynamically adjusted according to the characteristics of the ancient book carrier, the outputs of the parts are multiplied by the weight parameter and then added element by element, and is represented as:
[0109]
[0110] In the formula, represents the text candidate region, ReLU represents an activation function, and LayerNorm represents a normalization operation; 、 represents the weight parameter dynamically adjusted based on the material m of the ancient book carrier, and respectively represent the text edge positioning processing and the non-text region filtering processing, and X represents the target image;
[0111] In the text edge positioning branch, the enhanced image is taken as the input of the text edge positioning branch, is subjected to a 1x5 convolution, a 5x1 convolution and a feature extraction module with residual connection, and then is multiplied by the weight parameter of the text edge positioning branch to obtain an edge result output;
[0112] In the non-text region filtering branch, the feature information of the text region is enhanced by reversing the background feature, the same feature extraction module as that of the text edge positioning branch is adopted, and the feature elements are mapped to [0, 1] through a Sigmoid function, and then the text feature is indirectly enhanced by taking the inverse, and is represented as:
[0113]
[0114] wherein, denotes mapping the raw output of the feature extraction module to [0, 1], denotes non-text region filtering branch processing on the image X, ReLU denotes an activation function, and LayerNorm denotes a normalization operation.
[0115] Further, as Figure 2 indicated, the step of obtaining the text line image with a feature label by the DB segmentation algorithm with the introduced resource characteristics, according to the embodiment of the present application, specifically comprises the following steps:
[0116] S1021: constructing a DB model containing a deformable convolution DCNv2 module;
[0117] In the deformable convolution DCNv2 module, a scroll curvature factor of ancient books is introduced to control the spatial sensitivity of convolution to the scroll area of the vertical text by the curvature perception strength, wherein the scroll curvature factor of ancient books is determined based on the physical form parameters of the target image associated with the carrier material characteristics;
[0118] Specifically, for the input image of ancient books, the edge features are extracted by Canny edge detection, and the scroll area of the image can be segmented to obtain a binary scroll mask, wherein "1" in the mask represents the scroll area, and "0" represents the flat area; for the scroll area in the mask, a discrete point set on the edge contour thereof is extracted, the curvature of each point is calculated, the average curvature of all points is taken as the basic curvature, and then the carrier material of the target image is taken as a coefficient, and the product of the coefficient and the basic curvature is taken as the scroll curvature factor of ancient books;
[0119] Further, the step of obtaining the text line image with a feature label by the DB segmentation algorithm with the introduced resource characteristics, according to the embodiment of the present application, specifically further comprises the following steps:
[0120] S1022: segmenting the target image using the DB model to obtain a mask of the text area;
[0121] Preferably, in step S1022, a page edge curl mask generated by the ancient book 3D scanning data is collected, the ancient book 3D scanning data is obtained by the carrier material characteristics associated with the target image, and the page edge curl mask is generated in combination with the edge features of the text candidate area to segment the target image to obtain the text area mask;
[0122] S1023: according to the mask obtained by the DB segmentation, using a contour detection method based on Canny edge detection, the text area is segmented into multiple text lines;
[0123] Preferably, in step S1023, according to the text region mask, the text region is segmented into a plurality of text lines using a contour detection method based on Canny edge detection and a book format feature, the book format feature including a text edge positioning feature and a format structure information retained after filtering a non-text region;
[0124] S1024: extracting a plurality of text line images from the image according to the boundary box of the segmented text line;
[0125] In step S1024, the feature label of the text line image includes spatial position information of the corresponding text line in the target image and associated book carrier material feature parameters.
[0126] In one preferred implementation, in the DCNv2 module of the DB model of the embodiment, the DCNv2 module includes a deformable convolution layer, batch normalization, and a Hardswish activation function.
[0127] A curvature perception strength is introduced in the deformable convolution layer, the curvature perception strength is positively correlated with the book wrinkle curvature factor, and the curvature perception strength is used to control the sensitivity of the deformable convolution to spatial changes.
[0128] The curvature perception strength is used to measure the sensitivity of the model to spatial changes in the wrinkle region, and the curvature perception strength is positively correlated with the book wrinkle curvature factor in a linear relationship. When the book wrinkle is more obvious, the value of the curvature perception strength increases accordingly, indicating that the model needs to capture the text deformation caused by the wrinkle more sensitively. That is, the more obvious the wrinkle (the greater the book wrinkle curvature factor), the higher the perception requirement of the model to spatial changes (the greater the curvature perception strength).
[0129] In the embodiment, the curvature perception strength controls the sensitivity of the convolution by adjusting the sampling offset of the deformable convolution DCNv2 module (i.e., the offset of each sampling point of the convolution kernel relative to the normal position). Specifically, the deformable convolution DCNv2 module predicts the sampling offset so that the convolution kernel can adapt to the spatial deformation in the image. The curvature perception strength is used as a scaling factor for the sampling offset to adjust the original offset.
[0130] When processing the wrinkle region, the original offset is enlarged by the scaling factor, and the sampling point of the convolution kernel moves more in the direction away from the normal position, at which time the convolution sensitivity is high.
[0131] When processing the flat region, the original offset is reduced by the scaling factor, and the sampling point of the convolution kernel approaches the normal position, at which time the convolution sensitivity is low, avoiding meaningless offset.
[0132] Further, in the embodiment, a book ink aging coefficient is introduced in the Hardswish activation function is represented as:
[0133]
[0134] wherein x represents a feature value input into a Hardswish activation function, represents an ancient book ink trace aging coefficient, and the value range is 0.1 to 1.
[0135] In the embodiment, the ancient book ink trace aging coefficient is a parameter for quantifying the aging degree of the ancient book ink trace caused by long time, storage environment, etc. The ancient book ink trace aging coefficient is inversely proportional to the ink trace effective information retention rate, that is, the greater the value of the ancient book ink trace aging coefficient , the more serious the aging of the ancient book is, and the less the effective information retained on the ancient book is; on the contrary, the smaller the value of the ancient book ink trace aging coefficient , the better the preservation of the ancient book is, and the more information retained on the ancient book is.
[0136] Further, as shown in Figure 3 , in step S104, the text recognition model is used to recognize the first text segment and the second text segment respectively, and the recognition result of the coincident text segment obtained is compared with the preset knowledge graph, and the step of optimizing the text recognition model through semantic conflict detection includes the following steps:
[0137] S1031: taking the recognition result of the coincident text segment in the first text segment by the text recognition model as a first recognition result;
[0138] S1032: taking the recognition result of the coincident text segment in the second text segment by the text recognition model as a second recognition result;
[0139] S1033: when there is a difference between the first recognition result and the second recognition result, calculating a first difference rate, increasing the length of the coincident text segment of the next text line image segment, and calculating a second difference rate;
[0140] S1034: when calculating the first difference rate and the second difference rate, introducing an ancient book common false word correlation weight, wherein the ancient book common false word correlation weight is determined based on the corresponding relationship of the common false words recorded in the preset knowledge graph; when the second difference rate is greater than the first difference rate, combining the ink trace difference positioning error source of the ancient book annotation and the main text, adjusting the attention weight of the model to the annotation text, adjusting the text recognition model according to the adjusted attention weight, and recognizing the new text segment image based on the optimized and adjusted text recognition model.
[0141] In step S1034, the ancient book Tong-Fake character association weight is determined based on the corresponding relationship of Tong-Fake characters included in the preset knowledge graph.
[0142] The preset knowledge graph containing the structured Tong-Fake character relationship database is constructed based on ancient book literature corpus and tool books, etc.
[0143] In the preset knowledge graph, the corresponding relationship of Tong-Fake characters is stored in the form of triples. For example, the triple composed of the main character, the relationship attribute, and the Tong-Fake character. In the relationship attribute, the Tong-Fake frequency and the time adaptation degree are included.
[0144] The Tong-Fake frequency represents the number of occurrences in the literature corpus.
[0145] The time adaptation degree represents the coincidence degree of the usage time of the Tong-Fake relationship and the time of the target ancient book. For example, if the image ancient book to be processed is of the Eastern Han Dynasty, and the time label of the Tong-Fake relationship in the knowledge graph is also of the Eastern Han Dynasty, then the time adaptation degree is closer to 1, indicating a high coincidence degree. On the contrary, if a Tang Dynasty manuscript is identified, and the target ancient book is of the Tang Dynasty, while the reference label of the Tong-Fake character in the knowledge graph is of the Qin Dynasty, then the coincidence degree is extremely low, even 0.
[0146] Then, the Tong-Fake frequency and the time adaptation degree are normalized, so that the parameters of the Tong-Fake frequency and the time adaptation degree are mapped to the interval [0, 1].
[0147] Finally, the Tong-Fake character association weight of the ancient book is calculated by combining the time adaptation degree and the normalized Tong-Fake frequency , which is represented as: wherein, , represents the weight coefficient.
[0148] Further, in the process of optimizing and adjusting the text recognition model, Bayesian optimization is used to set and optimize the optimal hyperparameter combination of the text recognition model, and the optimization target is to minimize the error. The hyperparameter combination includes the attention weight parameter, which is input as the initial prior value of Bayesian optimization.
[0149] In the optimization process, Gaussian process is used as a probability agent model to fit the target function, wherein the target function of the Gaussian process incorporates the variant processing rules of ancient book collation; and the no-arbitrage calculation of the target function includes the difference between the recognition results before and after the attention weight adjustment; the verification information based on the probability agent model is used to construct the acquisition function, which is represented as:
[0150]
[0151] In the formula, f(x) represents the target value of the Gaussian process, including the model output after attention weight optimization, represents the current optimal target value in the Gaussian process, represents the cumulative density function of the Gaussian distribution, represents the balance parameter, represents the mean of the target function, represents the variance of the target function; represents the ancient book version coefficient.
[0152] Further, the text recognition model in step S104 of the embodiment includes a recognition encoder and a recognition decoder;
[0153] The recognition encoder is used to extract features from the text segment image, including a convolutional layer and a bidirectional LSTM layer, and an attention mechanism network is introduced in the bidirectional LSTM layer, which includes layer normalization and double attention, and residual connection is used between the two;
[0154] The recognition decoder includes a fully connected layer, which is used to decode the extracted features into a text sequence.
[0155] Specifically, in the recognition encoder:
[0156] The ancient manuscript ink layer extraction module is introduced in the convolutional layer, which captures the difference in ink concentration between the main text and the notes through multispectral imaging data, and the convolutional layer also adopts the inception structure of the GoogLeNet model, including the inception1 module, the inception2 module and the inception3 module.
[0157] Preferably, in the encoder of the text recognition model, the convolutional layer adopts the inception structure of the GoogLeNet model, including the inception1 module, the inception2 module and the inception3 module.
[0158] In the Inception1 module, it includes: two parallel 1x1 convolutional operations for increasing the number of feature channels; a 3x3 convolutional operation followed by a 1x1 convolutional operation for extracting spatial features and increasing the number of feature channels through 1x1 convolution; a 5x5 convolutional operation followed by a 1x1 convolutional operation; a 3x3 max pooling operation for performing a 3x3 max pooling operation on the input feature map and connecting a 1x1 convolutional operation, the max pooling operation is used to reduce the spatial dimension of the feature map without changing the number of feature channels; a connection layer for connecting the outputs of the above four parallel convolutional operations in the channel dimension to form a feature map with multi-scale information;
[0159] In the inception2 module, the 5*5 convolution kernel in the inception1 module is replaced by two layers of 3*3 convolution superposition, and the 3*3 convolution is divided into two groups of 3*1 and 1*3 asymmetric convolution
[0160] In the inception3 module, the serial asymmetric convolution in the inception2 module is replaced by parallel, and a 1*1 residual connection structure is introduced.
[0161] The bidirectional LSTM layer includes a forward LSTM module, a backward LSTM module and an attention mechanism module; wherein the forward LSTM module processes the text sequence in the order of vertical arrangement from top to bottom, and the backward LSTM module fuses the sentence reading symbol features of the ancient book;
[0162] The forward LSTM module is used for forward LSTM unit processing of the input feature vector sequence, and the output of each time step is a hidden layer vector;
[0163] The backward LSTM module is used for backward LSTM unit processing of the input feature vector sequence, and the output of each time step is a hidden layer vector;
[0164] The attention mechanism module is used to calculate attention weights from the outputs of the forward LSTM and the backward LSTM, respectively. For each time step, there are two attention weights, which correspond to the forward and backward LSTMs respectively. After calculating the attention weights, the weights are processed by layer normalization. The normalized attention weights are connected in residual with the original hidden state. For each time step, the normalized attention weights are added to the corresponding forward and backward hidden states, and the residual connected hidden state is used for weighted summation to obtain the final text feature.
[0165] In the full connection layer of the recognition decoder, the text feature is mapped to the category space through the full connection layer, and the class probability is output using the softmax function. The recognition result of the text feature is determined based on the class probability, which is represented as:
[0166]
[0167] In the formula, represents a bias term; represents a weight matrix, represents the text feature before mapping, represents the mapped text feature.
[0168] In an implementation manner of the present application, the bidirectional LSTM layer introduces vertical text timing attention, which includes:
[0169] The vertical text time sequence attention weight calculation module is constructed,
[0170] The time sequence attention mask is generated based on the line sequence priority and the inter-character topological relationship of the vertical text of the ancient book, wherein the line sequence priority is dynamically adjusted through the ratio of the column spacing to the line height in the ancient book format, and the inter-character topological relationship is represented by the deviation amount of the gravity centers of adjacent characters.
[0171] The ancient book line style feature coefficient θ is introduced in the time sequence attention weight calculation, and the ancient book line style feature coefficient θ is determined in combination with the line spacing specification, the character spacing ratio and the format layout features of the ancient book, and the calculation formula is represented as:
[0172]
[0173] In the formula, and represent the weight coefficients, + = 1. represents the standard line spacing of the vertical text of the ancient book, represents the actually measured line spacing; represents the average width of a single character, represents the average height of a single character; represents the inclination angle of the ancient book page, and k represents the inclination influence factor.
[0174] The hidden layer vectors output by the forward LSTM module and the hidden layer vectors output by the backward LSTM module of the bidirectional LSTM layer are spliced to obtain a fusion vector sequence, the time sequence attention mask is multiplied with the fusion vector sequence element by element to obtain a feature vector with time sequence weight, and the feature vector with time sequence weight is input into the output layer of the bidirectional LSTM layer, and the original fusion vector sequence is added through residual connection and then subjected to layer normalization processing to obtain a final time sequence enhanced feature vector, which is used for subsequent decoding and recognition of the text sequence.
[0175] In some embodiments of the present application, in step S105, the newly acquired ancient book image is preprocessed, such as size normalization, inclination correction, noise filtering, etc., and then the style features in the image are extracted through a lightweight feature extraction network, including stroke thickness distribution, corner roundness, ink density gradient, etc., to generate a style feature vector, calculate the cosine similarity of the feature vector with the style feature library, select the two styles with the highest similarity as the target style, and output the style matching confidence.
[0176] The trained style transfer model is called to generate two style verification images from the new image as input combined with the matched target style vector and the timing feature constraint. The new image and the verification image are input into the trained text recognition model, and the model performs text line segmentation, feature extraction and sequence decoding on the two images respectively to output two initial recognition results.
[0177] Then the two results are processed by a dynamic weighted fusion strategy: for single character recognition, the result with higher confidence is taken as the final output; for the segment with homophonic characters and different texts, the conflict is resolved by combining the semantic association of the knowledge graph, and the text recognition result with confidence is output.
[0178] In addition, for the construction of the cultural database, the embodiment adopts a MySQL+MongoDB hybrid architecture, MySQL stores structured data such as ancient book metadata and recognition result index, and MongoDB stores unstructured data such as original images, verification images and complete recognition text; finally, a visual query interface is provided for the database: supporting retrieval by ancient book name, age and keyword, displaying original images, recognition text and confidence heat map; the embodiment realizes the automatic recognition and structured storage of new ancient book images, and the constructed cultural resource database can support efficient retrieval, comparison and in-depth research of ancient book text, and complete the whole process of digital transformation from image to usable data.
[0179] Figure 4 A hardware structure block diagram of a computer terminal for implementing an image recognition-based cultural resource digitization method is shown.
[0180] As shown in Figure 4 , the computer terminal 30 can include one or more (shown in the figure as 302a, 302b, …, 302n) processors 302 (the processor 302 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission module 306 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 4 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device.
[0181] For example, the computer terminal 30 can include more or fewer components than those shown in Figure 4 , or have a different configuration from that shown in Figure 4 .
[0182] It should be noted that the one or more processors 302 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 30. As referred to in embodiments of the present application, the data processing circuitry functions as a processor to control, for example, the selection of the variable resistance terminal path in connection with the interface.
[0183] The memory 304 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the image recognition based cultural resource digitization implementation method in embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, i.e. implements the image recognition based cultural resource digitization implementation method described above.
[0184] The memory 304 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 304 can further include memory disposed remotely with respect to the processor 302, which can be connected to the computer terminal 30 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0185] The transmission module 306 is configured to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 30.
[0186] In one example, the transmission module 306 includes a network interface controller (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet.
[0187] In one example, the transmission module 306 can be a radio frequency (RF) module, which is configured to communicate with the Internet in a wireless manner.
[0188] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the computer terminal 30.
[0189] It should be noted that, Figure 4 The computer terminal shown is configured to execute Figure 1The image recognition-based cultural resource digitization implementation method shown is applicable to the electronic device, and the related explanations in the execution method of the above commands are also applicable to the electronic device, and will not be described here.
[0190] The embodiments of the present application also provide a non-volatile storage medium, which comprises a stored program, wherein the program controls a device where the storage medium is located to execute the above image recognition-based cultural resource digitization implementation method when the program is running.
[0191] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical or other forms.
[0192] The above is only the preferred embodiment of the present application, and it should be pointed out that, for those skilled in the art, without departing from the principle of the present application, some improvements and refinements can be made, which should be regarded as the protection scope of the present application.
Claims
1. A method for realizing cultural resource digitization based on image recognition, characterized in that, The method comprises the following steps: A style transfer model based on a generative adversarial network is built, and a time sequence feature constraint is introduced into the style transfer model, the time sequence feature constraint being constructed based on a layout time sequence of the ancient book text and used to constrain text time sequence consistency in the style transfer process; A content encoder is used to extract a text structure of an input image, a style encoder is used to convert features in a pre-constructed style feature library into a style vector, and a transfer decoder is used to generate a text image with a fused target style in combination with the text structure, the style vector and the time sequence feature constraint; A trained style transfer model is used to generate a stylized training data set, the data set being associated with information labels of ancient book resources; A text recognition model is used to perform text recognition on a first text segment and a second text segment respectively, there is an overlapping text segment between a segment end of the first text segment and a segment start of the second text segment, a recognition result of the overlapping text segment is compared with a preset knowledge graph, a text recognition model is optimized through semantic conflict detection, and the text recognition model is trained by using a gradual style fusion mechanism and introducing the stylized training data set, so that a feature vector output by the model is bidirectionally matched with semantic features of the knowledge graph; A new image is subjected to dynamic style adaptation processing, text style features are extracted and matched with style types in a style feature library, a style transfer model is called to generate a verification image, the verification image and the new image are input into a trained text recognition model at the same time, a text recognition result of a current image to be recognized is obtained, after text recognition of an image data set is completed, a cultural resource database containing the text recognition result is constructed and output. 2.The image recognition-based cultural resource digitization implementation method of claim 1, wherein, Before the text recognition model is used to perform text recognition on the first text segment and the second text segment respectively, the following steps are further included: Text candidate regions in a target image are extracted based on text edge positioning, non-text region filtering and layout positioning; A DB segmentation algorithm with introduced resource characteristics is used to perform text segmentation on the text candidate regions, to obtain text line images with feature labels; The text line images are subjected to segmentation processing, to obtain the first text segment and the second text segment with different lengths. 3.The image recognition-based cultural resource digitization implementation method of claim 2, wherein, The text candidate regions in the target image are obtained through text edge positioning and non-text region filtering processing, wherein: Both the text edge positioning and the non-text region filtering have a learnable weight parameter, the weight parameter is dynamically adjusted based on characteristics of ancient book carriers, and the outputs of the parts are multiplied by the weight parameter and then added element by element, which is represented as: ; In the formula, F total , ReLU represents an activation function, and LayerNorm represents a normalization operation. , , represents a weight parameter based on the dynamic adjustment of the material m of the ancient book carrier, and respectively represent text edge positioning processing and non-text region filtering processing, and X represents a target image. In the text edge positioning branch, the target image is taken as an input of the text edge positioning branch, and after passing through a feature extraction module with 1x5 convolution, 5x1 convolution and residual connection, the edge result output is obtained by multiplying the text edge positioning branch weight parameter; In the non-text region filtering branch, the feature information of the text region is enhanced through inversion of the background feature, the same feature extraction module as that of the text edge positioning branch is adopted, and the feature elements are mapped to [0, 1] through a Sigmoid function, and then the text feature is enhanced through inversion, which is represented as: ; wherein denotes mapping the raw output of the feature extraction module to [0, 1], denotes non-text region filtering branch processing of the image X. 4.The image recognition-based cultural resource digitization implementation method of claim 3, wherein, The step of using the DB segmentation algorithm with introduced resource characteristics to perform text segmentation on the text candidate regions to obtain the text line images with feature labels comprises: A DB model comprising a deformable convolution DCNv2 module is constructed, and a scroll curvature factor is introduced in the deformable convolution DCNv2 module, so as to control the spatial sensitivity of the convolution to the scroll area of the vertical text by the curvature perception strength; wherein the scroll curvature factor is determined based on the scroll physical form parameters associated with the carrier material characteristics of the target image; The page edge curl mask generated by the ancient book 3D scanning data is obtained by associating the carrier material characteristics of the target image, and the generation of the page edge curl mask is combined with the edge features of the text candidate area The target image is segmented to obtain a text region mask. According to the text region mask, a profile detection method based on Canny edge detection and scroll format features is used to segment the text region into multiple text lines, and the scroll format features include text edge positioning features and format structure information retained after filtering the non-text region; According to the bounding box of the segmented text line, a plurality of text line images are extracted from the image, and the feature label of the text line image includes the spatial position information of the corresponding text line in the target image and the associated scroll carrier material characteristic parameters. 5.The image recognition-based cultural resource digitization implementation method of claim 4, wherein, In the DCNv2 module of the DB model, the DCNv2 module comprises a deformable convolution layer, a batch normalization and a Hardswish activation function; The curvature perception strength is introduced in the deformable convolution layer, the curvature perception strength is positively correlated with the scroll curvature factor, and the curvature perception strength is used to control the sensitivity of the deformable convolution to the spatial variation; In the Hardswish activation function, the ancient manuscript ink aging coefficient is introduced is expressed as: ; wherein x represents a feature value input into a Hardswish activation function, represents the aging coefficient of the ancient book ink marks. 6.The image recognition-based cultural resource digitization implementation method of claim 5, wherein, The text recognition model is used to recognize the first text segment and the second text segment respectively, the recognition result of the overlapping text segment obtained is compared with the preset knowledge graph, and the text recognition model is optimized through the semantic conflict detection, including: The recognition result of the overlapping text segment in the first text segment by the text recognition model is taken as the first recognition result; The recognition result of the overlapping text segment in the second text segment by the text recognition model is taken as the second recognition result; When there is a difference between the first recognition result and the second recognition result, a first difference rate is calculated, the length of the overlapping text segment of the next text line image segmentation is increased, and a second difference rate is calculated; In the calculation of the first difference rate and the second difference rate, a scroll homophone association weight is introduced, wherein the scroll homophone association weight is determined based on the homophone corresponding relationship recorded in the preset knowledge graph; When the second difference rate is greater than the first difference rate, the attention weight of the model to the annotation text is adjusted in combination with the ink difference error source of the scroll annotation and the main text, the text recognition model is optimized and adjusted according to the adjusted attention weight, and the new text segment image is recognized by the optimized and adjusted text recognition model. 7.The image recognition-based cultural resource digitization implementation method of claim 6, wherein, The step of optimizing and adjusting the text recognition model, including: Bayesian optimization is used to set and optimize the optimal hyperparameter combination of the text recognition model, and the optimization target is to minimize the error; wherein the hyperparameter combination includes an attention weight parameter, and the attention weight parameter is input as the initial prior value of Bayesian optimization; In the optimization process, a Gaussian process is used as a probability agent model to fit the target function; wherein the target function of the Gaussian process incorporates the variant processing rules of the scroll collation; and the no-arbitrage calculation of the target function includes the difference between the recognition results before and after the attention weight adjustment; the verification information based on the probability agent model is used to construct a collection function, and the collection function is represented as: ; In the formula, f(x) represents the target value of the Gaussian process, including the model output after attention weight optimization, represents the current optimal target value in the Gaussian process, represents the cumulative density function of the Gaussian distribution, represents the balance parameter, represents the mean of the target function, represents the variance of the target function; represents the ancient book version coefficient. 8.The image recognition-based cultural resource digitization implementation method of claim 7, wherein, The text recognition model comprises a recognition encoder and a recognition decoder; The recognition encoder is configured to extract features from the text segment image, and comprises a convolutional layer and a bidirectional LSTM layer, wherein an attention mechanism network is introduced in the bidirectional LSTM layer, and the attention mechanism network comprises layer normalization and double attention, and the two are connected by a residual connection; The recognition decoder comprises a fully connected layer, and is configured to decode the extracted features into a text sequence. 9.The image recognition-based cultural resource digitization implementation method of claim 8, wherein, In the recognition encoder: The convolutional layer introduces a hierarchical extraction module for ancient book ink, which captures the difference in ink concentration between the main text and the notes through multispectral imaging data, and adopts the inception structure of the GoogLeNet model, including an inception1 module, an inception2 module and an inception3 module; The bidirectional LSTM layer comprises a forward LSTM module, a backward LSTM module and an attention mechanism module; wherein the forward LSTM module processes the text sequence in the order of top-to-bottom vertical arrangement, and the backward LSTM module fuses the sentence reading symbol features of the ancient book; The forward LSTM module is configured to perform forward LSTM unit processing on the input feature vector sequence, and the output of each time step is a hidden layer vector; the backward LSTM module is configured to perform backward LSTM unit processing on the input feature vector sequence, and the output of each time step is a hidden layer vector; The attention mechanism module is configured to calculate attention weights from the outputs of the forward LSTM and the backward LSTM respectively, and for each time step, there are two attention weights, which correspond to the forward and backward LSTMs respectively, and after calculating the attention weights, the weights are subjected to layer normalization processing; the normalized attention weights are connected with the original hidden state by a residual connection, and for each time step, the normalized attention weights are added to the corresponding forward and backward hidden states, the hidden states after the residual connection are weighted and summed to obtain the final text features; In the fully connected layer of the recognition decoder, the text features are mapped to a category space through the fully connected layer, a softmax function is used to output a category probability, and the recognition result of the text features is determined based on the category probability, which is represented as: ; In the formula, represents a bias term; represents a weight matrix, represents a text feature before mapping, represents a text feature after mapping. 10.The image recognition-based cultural resource digitization implementation method of claim 9, wherein, The bidirectional LSTM layer introduces vertical text timing attention, including: A vertical text timing attention weight calculation module is constructed, Based on the line sequence priority and the inter-character topological relationship of the vertical text of the ancient book, a timing attention mask is generated; wherein the line sequence priority is dynamically adjusted by the ratio of the column spacing to the line height in the format of the ancient book, and the inter-character topological relationship is represented by the deviation amount of the barycentric coordinates of adjacent characters; An ancient book line layout feature coefficient θ is introduced in the timing attention weight calculation, and the ancient book line layout feature coefficient θ is determined in combination with the line spacing specification, the character spacing ratio and the format layout features of the ancient book, and the calculation formula thereof is represented as: ; where λ1 and λ2 represent weight coefficients, λ1 + λ2 = 1; denotes the standard line spacing of the vertical text of the ancient book, denotes the actual measured line spacing; denotes the average width of a single character, denotes the average height of a single character; denotes the angle of inclination of the page of the ancient book, and k denotes an inclination influence factor; The hidden layer vectors output by the forward LSTM module and the hidden layer vectors output by the backward LSTM module of the bidirectional LSTM layer are spliced to obtain a fusion vector sequence, and the timing attention mask is multiplied element by element with the fusion vector sequence to obtain a feature vector with timing weights. The feature vector with time sequence weight is input into an output layer of a bidirectional LSTM layer, is added to an original fusion vector sequence through a residual connection, and is subjected to layer normalization processing to obtain a final time sequence enhanced feature vector, which is used for subsequent decoding and recognition of the text sequence.
Citation Information
Patent Citations
Ancient document digitization method
CN115797946A
Image style migration method and device, electronic equipment and storage medium
CN118570050A