Image stitching method for dunhuang manuscripts fragments based on sentence coherence

By using a sentence fluency-based joining method, combined with missing type classification, ratio filtering, and neural network models, the problem of low accuracy and efficiency in joining images of Dunhuang manuscript fragments in existing technologies has been solved, achieving a more efficient and accurate joining process.

CN115620057BActive Publication Date: 2026-01-16HENAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211276002.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2026-01-16
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

Existing methods for piecing together images of Dunhuang manuscript fragments rely on single physical information, resulting in low accuracy and inefficiency. In particular, when the fragments are flush, they are prone to generating interfering candidate results. Furthermore, differences in image scaling caused by the shooting tools increase the difficulty of piecing together.

Method used

A sentence fluency-based merging method is adopted, which improves the accuracy and efficiency of merging by classifying missing types, filtering ratio ranges, judging handwriting, and detecting sentence fluency, combined with a neural network model, and comprehensively considering text content and edge similarity.

Benefits of technology

It improved the accuracy and efficiency of image stitching of Dunhuang manuscript fragments, reduced the possibility of incorrect stitching, and reduced the amount of manual work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620057B_ABST
    Figure CN115620057B_ABST
Patent Text Reader

Abstract

The application discloses a Dunhuang manuscript fragment image splicing method based on sentence fluency, and comprises the following steps: A: classifying the fragment images according to missing types; B: acquiring the column width, column height and gap of the fragment images; C: judging the fragment images A and B according to whether the missing types correspond; D: judging the ratio of the column width and column height of the fragment images A and B and the ratio of the column width and gap; E: judging the column height ratio, gap ratio and column width ratio of the fragment image B after equal-ratio enlargement and the fragment image A; F: judging the handwriting similarity of the fragment images A and B by using a handwriting judgment neural network model; G: performing edge similarity calculation and sentence fluency detection on the fragment images A and B; and H: after the comparison of all the fragment images to be spliced is completed, obtaining all the candidate spliced images of the fragment image A of the Dunhuang manuscripts. The application comprehensively considers the sentence fluency and edge similarity of the text content of the fragment images to be spliced, and improves the efficiency and accuracy of the Dunhuang manuscript fragment image splicing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for splicing images of broken articles, and in particular to a method for splicing images of Dunhuang Manuscript fragments based on sentence fluency. BACKGROUND

[0002] Dunhuang Manuscripts are important research materials for studying the history, archaeology, religion, anthropology, sociology, linguistics, literary history, art history, science and technology history, and national history of China, Central Asia, East Asia, and South Asia in the middle and ancient periods, and have very high cultural relic value and literature research value. Dunhuang Manuscript fragment images are the main materials for Dunhuang Manuscript research. In the professional field of Dunhuang Manuscript research, the original edge of a Dunhuang Manuscript fragment image refers to the book edge, which is the edge of the Dunhuang Manuscript fragment paper in the Dunhuang Manuscript fragment image, and the broken edge is not naturally formed, but is the edge formed by the Dunhuang Manuscript paper due to damage. There are book horizontal boundary lines and book vertical grid lines on the Dunhuang Manuscript paper, which are the upper and lower boundaries of the Dunhuang Manuscript paper and the vertical alignment straight lines drawn by ancient paper mills on the Dunhuang Manuscript paper. However, on many Dunhuang Manuscript papers, the book vertical grid lines have disappeared or become difficult to identify.

[0003] In the existing Dunhuang Manuscript research process, researchers usually manually splice Dunhuang Manuscripts using domain expertise to determine whether two Dunhuang Manuscript fragments belong to the same place before being damaged. The above-mentioned manual splicing method has low accuracy and efficiency, and requires high work intensity.

[0004] The invention patent with the application number CN202110440552.4 and the name "Automatic Splicing Method for Dunhuang Manuscript Fragment Images" discloses an automatic splicing method for Dunhuang Manuscript fragment images, which can balance the broken edge and the accuracy of the grid unit width formed after splicing, and improve the efficiency and accuracy of Dunhuang Manuscript fragment image splicing. However, the above-mentioned patent only uses the physical information of Dunhuang Manuscript fragment images as the splicing reference factor, relies on the book vertical grid lines, and the reference factor is relatively single, and the splicing accuracy needs to be further improved. Moreover, when the edge of a Dunhuang Manuscript fragment image to be spliced is flat, the edge feature is extremely insignificant, which easily leads to other Dunhuang Manuscript fragment images with flat edges as candidate results, thereby bringing a large number of interfering candidate results, and finally leading to poor splicing effect.

[0005] On the other hand, due to the different reasons of the collection unit, the shooting tool and technology, the photographer, the shooting distance, the shooting angle, the shooting specification and the standard, the image captured by the shooting tool has different degrees of difference with the real physical size of the Dunhuang Manuscript, which causes the Dunhuang Manuscript image not to be the original size, and there is different scaling and different degrees of scaling, so that two or more Dunhuang Manuscript fragments originally belonging to the same Dunhuang Manuscript may not be the original size after being shot by the shooting tool, and the scaling degrees are different. The above-mentioned situation puts forward new challenges for the splicing of the Dunhuang Manuscript image. SUMMARY

[0006] The purpose of the present application is to provide a Dunhuang Manuscript fragment image splicing method based on the sentence fluency, which can comprehensively consider the sentence fluency and edge similarity of the Dunhuang Manuscript fragment image to be spliced, and improve the efficiency and accuracy of the Dunhuang Manuscript fragment image splicing.

[0007] The present application adopts the following technical scheme:

[0008] A Dunhuang Manuscript fragment image splicing method based on the sentence fluency, comprising the following steps:

[0009] A: According to the position and the broken edge shape of the missing part of the Dunhuang Manuscript fragment, the Dunhuang Manuscript fragment images to be spliced are artificially classified into the following missing types:

[0010] a type: upper missing; b type: lower missing; c type: left lower missing; d type: left upper missing; e type: right lower missing; f type: right upper missing; g type: left straight missing; h type: right straight missing; i type: left zigzag missing; j type: right zigzag missing;

[0011] B: Labeling each column of text in the Dunhuang Manuscript fragment image to be spliced to obtain the minimum bounding rectangle and the text content of each column of text; taking the width and height of the minimum bounding rectangle Q with the maximum height value as the column width w and the column height h respectively, and taking the horizontal distance between the minimum bounding rectangle Q and the minimum bounding rectangle adjacent to Q as the gap d;

[0012] C: For the specified Dunhuang Manuscript fragment image A and the Dunhuang Manuscript fragment image B to be spliced, first, it is judged whether the missing types of the Dunhuang Manuscript fragment image A and the Dunhuang Manuscript fragment image B correspond to each other, if the missing types correspond to each other, then step D is entered; if the missing types do not correspond to each other, then according to whether the missing types correspond to each other, the Dunhuang Manuscript fragment image A and the next Dunhuang Manuscript fragment image to be spliced are continuously compared;

[0013] D: judge whether the ratio P1 of the column width and height ratio of Dunhuang Manuscript Fragment Image A and the column width and height ratio of Dunhuang Manuscript Fragment Image B and the ratio P2 of the column width gap ratio of Dunhuang Manuscript Fragment Image A and the column width gap ratio of Dunhuang Manuscript Fragment Image B are in the set first ratio range at the same time; if the ratio P1 and the ratio P2 are in the set first ratio range at the same time, go to step E; if not, return to step C and continue to compare Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be combined;

[0014] E: scale Dunhuang Manuscript Fragment Image B to make the column width and gap of the scaled Dunhuang Manuscript Fragment Image B consistent with the column width and gap of Dunhuang Manuscript Fragment Image A in turn; when the column width is consistent, calculate the ratio of the column height of the scaled Dunhuang Manuscript Fragment Image B and the column height of Dunhuang Manuscript Fragment Image A and the ratio of the gap of the scaled Dunhuang Manuscript Fragment Image B and the gap of Dunhuang Manuscript Fragment Image A; when the gap is consistent, calculate the ratio of the column height of the scaled Dunhuang Manuscript Fragment Image B and the column height of Dunhuang Manuscript Fragment Image A and the ratio of the column width of the scaled Dunhuang Manuscript Fragment Image B and the column width of Dunhuang Manuscript Fragment Image A;

[0015] If the ratio of the column height and the ratio of the gap are in the set second ratio range at the same time when the column width is consistent, and the ratio of the column height and the ratio of the column width are in the set third ratio range at the same time when the gap is consistent, go to step F; otherwise, return to step C;

[0016] F: train the handwriting judgment neural network model, and then input Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B into the trained handwriting judgment neural network model for judgment; if the similarity of the handwriting in Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B is greater than or equal to the set handwriting similarity threshold, go to step G; if it is less than the set handwriting similarity threshold, return to step C;

[0017] G: first, calculate the edge similarity of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B to obtain the time sequence matching degree of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B; then judge whether the missing type of Dunhuang Manuscript Fragment Image A belongs to type i or type j; if not, continue to detect the sentence fluency of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B to obtain the sentence fluency value of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B;

[0018] When the missing type of the Dunhuang Manuscript Fragment Image A belongs to the i type or the j type, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, the Dunhuang Manuscript Fragment Image B is taken as a candidate conjugation image, then step C is returned to continue comparing the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; otherwise, step C is directly returned to continue comparing the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated.

[0019] When the missing type of the Dunhuang Manuscript Fragment Image A does not belong to the i type or the j type, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, and the maximum value in the sentence fluency value is greater than or equal to the set sentence fluency threshold value, the Dunhuang Manuscript Fragment Image B is taken as a candidate conjugation image, then step C is returned to continue comparing the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; otherwise, step C is directly returned to continue comparing the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated.

[0020] H: When the comparison of the Dunhuang Manuscript Fragment Image A with all the Dunhuang Manuscript Fragment Images to be conjugated is completed, all the candidate conjugation images of the Dunhuang Manuscript Fragment Image A are obtained.

[0021] In step A, the missing position and the broken edge shape of the Dunhuang Manuscript Fragment are determined by manual judgment.

[0022] When the broken edge is in a horizontal direction and the broken edge crosses several text columns in the Dunhuang Manuscript Fragment Image, if the missing position of the Dunhuang Manuscript Fragment is missing upper part, it is determined as a type: missing upper part; if the missing position of the Dunhuang Manuscript Fragment is missing lower part, it is determined as b type: missing lower part.

[0023] When the broken edge is in a vertical direction and the broken edge does not cross any text column in the Dunhuang Manuscript Fragment Image, that is, the broken edge is located at the gap on one side of the text column, if the missing position of the Dunhuang Manuscript Fragment is missing left side, it is determined as g type: missing left side straightly; if the missing position of the Dunhuang Manuscript Fragment is missing right side, it is determined as h type: missing right side straightly.

[0024] When the broken edge is in a vertical direction and the broken edge is jagged and repeatedly crosses one or more text columns in the Dunhuang Manuscript Fragment Image, if the missing position of the Dunhuang Manuscript Fragment is missing left side, it is determined as i type: missing left side jaggedly; if the missing position of the Dunhuang Manuscript Fragment is missing right side, it is determined as j type: missing right side jaggedly.

[0025] When the broken edge is inclined and the broken edge crosses several text columns in the Dunhuang Manuscript Fragment image, if the missing part of the Dunhuang Manuscript Fragment is missing from the lower left, it is determined to be type c: missing from the lower left; if the missing part of the Dunhuang Manuscript Fragment is missing from the upper left, it is determined to be type d: missing from the upper left; if the missing part of the Dunhuang Manuscript Fragment is missing from the lower right, it is determined to be type e: missing from the lower right; if the missing part of the Dunhuang Manuscript Fragment is missing from the upper right, it is determined to be type f: missing from the upper right.

[0026] In step G, when performing the sentence fluency detection of the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B, the following steps are performed:

[0027] G2-1: A training set is constructed using an existing Dunhuang Manuscript Fragment image annotation data set, and the BERT pre-training language model is fine-tuned and trained using the training set to obtain a fine-tuned and trained BERT language model;

[0028] G2-2: The fine-tuned and trained BERT language model is used to perform sentence fluency detection on the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B to obtain a sentence fluency value of the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B.

[0029] Step G2-1 includes the following specific steps:

[0030] G2-1-1: An existing Dunhuang Manuscript Fragment image annotation data set is obtained, and the Dunhuang Manuscript Fragment image annotation data set is composed of several sentences;

[0031] G2-1-2: The traditional Chinese characters in the Dunhuang Manuscript Fragment image annotation data set are converted to simplified Chinese characters, and then the characters in the data set after the traditional-to-simplified processing that do not appear in the BERT word table are de-duplicated and supplemented into the BERT word table;

[0032] G2-1-3: In the Dunhuang Manuscript Fragment image annotation data set after the traditional-to-simplified processing, only each sentence in the paragraph with a number of characters greater than or equal to a character number threshold is retained, and a training set is established according to each retained sentence;

[0033] G2-1-4: According to the obtained training set, positive and negative samples required for BERT language model training are constructed;

[0034] G2-1-5: The obtained positive and negative samples are used to fine-tune and train the BERT pre-training language model, and finally a fine-tuned and trained BERT language model is obtained.

[0035] Step G2-1-4 includes the following specific steps:

[0036] First, each sentence in the training set is input into the sentence reading system, and the punctuation marks in each sentence are marked by the sentence reading system;

[0037] Then, according to the whole sentence after marking the punctuation marks, the positive samples are constructed by the following method:

[0038] (1) Find the period in the whole sentence, and divide the whole sentence into one or more sentences by one or more periods;

[0039] (2) Find the comma in each sentence;

[0040] If there is no comma in the whole sentence, randomly select the first several characters in the whole sentence as the first divided sentence of the positive sample, and the remaining characters as the second divided sentence;

[0041] If there is a comma in the whole sentence, divide the whole sentence into two or more sub-sentences by one or more commas; then divide the two or more sub-sentences in the whole sentence into the first divided sentence and the second divided sentence of the positive sample in order, wherein the first divided sentence of the positive sample contains at least one sub-sentence, and the second divided sentence of the positive sample contains at least one sub-sentence;

[0042] Finally, the CSV format positive sample dataset is constructed, and the expression of the positive sample is [S1, S2, 1]; wherein S1 represents the first divided sentence of the positive sample, S2 represents the second divided sentence of the positive sample, and label 1 represents the positive sample;

[0043] Finally, the negative samples are constructed by the following method:

[0044] (1) Randomly select sub-sentences from two different sentences as the first divided sentence and the second divided sentence of the negative sample, wherein the first divided sentence of the negative sample contains at least one sub-sentence, and the second divided sentence contains at least one sub-sentence;

[0045] (2) From the whole sentence containing a period, select the last one or more sub-sentences in the sentence before the period as the first divided sentence of the negative sample, and then select the first one or more sub-sentences in the sentence after the period as the second divided sentence of the negative sample;

[0046] Finally, the CSV format negative sample dataset is constructed, and the expression of the negative sample is [S3, S4, 0]; wherein S3 represents the first divided sentence of the negative sample, S4 represents the second divided sentence of the negative sample, and label 0 represents the negative sample.

[0047] The step G2-2 includes the following specific steps:

[0048] G2-2-1: Extract the text information from images A and B of the Dunhuang manuscript fragments; then convert the traditional Chinese characters in the text information into simplified Chinese characters;

[0049] G2-2-2: Based on the position of the text information in the image A of the Dunhuang manuscript fragment, divide the text information of the image A (after the complex-to-simplified conversion) into several text columns in a right-to-left and top-to-bottom order, and construct a set S of text columns for the image A, S = {S1, S2, ..., S...} M}, S1 to S m These represent the text information from column 1 to column m from right to left in image A of the Dunhuang manuscript fragment;

[0050] Construct a text column set T for image B of Dunhuang manuscript fragments using the same method, where T = {T1, T2, ..., T...} n}, T1 to T n These represent the text information in the first to nth columns from right to left in image B of the Dunhuang manuscript fragment;

[0051] G2-2-3: Calculate the maximum sentence fluency of Dunhuang manuscript fragment image A and Dunhuang manuscript fragment image B under various mutual positional relationships and text column correspondence states;

[0052] G2-2-4: Among the sentence fluency values ​​of the obtained Dunhuang manuscript fragment images A and B under various relative positions and alignment states, select the highest value as the sentence fluency value of Dunhuang manuscript fragment images A and B.

[0053] Step G2-2-3 includes the following specific steps:

[0054] If Dunhuang manuscript fragment A and Dunhuang manuscript fragment B are of type b and a, type c and f, or type e and d respectively; then the first text column S1 of Dunhuang manuscript fragment A is aligned vertically with the first text column T1 of Dunhuang manuscript fragment B. The text in text columns S1 and T1 is then concatenated sequentially to form a string. This string is then input into a punctuation system for punctuation. Finally, it is determined whether the connection points between the text in text columns S1 and T1 are marked with punctuation marks.

[0055] If no symbol is marked at the connection point, the clause containing the connection point is selected, and the text before and after the connection point of the clause is taken as the clause pair to be predicted at position S1-T1.

[0056] If the symbol marked at the connection is a period or a comma, a clause before and after the period or the comma is selected as the pair of clauses to be predicted at the S1-T1 position;

[0057] The pair of clauses to be predicted at the S1-T1 position is input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S1, T1); then the pair of clauses to be predicted at the S2-T2 position is obtained according to the above method and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S2, T2), and so on, until the pair of clauses to be predicted at the Sm-Tm position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m -T n The pair of clauses to be predicted at the S1-T1 position is input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S1, T1); then the pair of clauses to be predicted at the S2-T2 position is obtained according to the above method and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S2, T2), and so on, until the pair of clauses to be predicted at the Sm-Tm position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m , T n ); (S m , T n ) represents that the current alignment state is that the text column S m corresponds to the text column T n ;

[0058] Then, the second column text column S2 of the Dunhuang Manuscript Fragment Image A is corresponded to the first column text column T1 of the Dunhuang Manuscript Fragment Image B according to the above method, the pair of clauses to be predicted at the S2-T1 position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S2, T1); the pair of clauses to be predicted at the S3-T2 position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S3, T2), and so on, until the pair of clauses to be predicted at the Sm-Tm position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m -T n-1 The pair of clauses to be predicted at the S1-T1 position is input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S1, T1); then the pair of clauses to be predicted at the S2-T2 position is obtained according to the above method and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S2, T2), and so on, until the pair of clauses to be predicted at the Sm-Tm position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m , T n-1 );

[0059] By analogy, the first column text column S1 of the Dunhuang Manuscript Fragment Image A is corresponded to the second column text column T2 of the Dunhuang Manuscript Fragment Image B according to the above method, the corresponding sentence fluency value NSP(S1, T2) at the S1-T2 position, the corresponding sentence fluency value NSP(S2, T3) at the S2-T3 position, and so on, until the corresponding sentence fluency value NSP(Sm, Tm) at the Sm-Tm position is obtained. m -T1 position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m -T1 position is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(Sm, Tm). m

[0060] By analogy, the first column text column S1 of the Dunhuang Manuscript Fragment Image A is corresponded to the second column text column T2 of the Dunhuang Manuscript Fragment Image B according to the above method, the corresponding sentence fluency value NSP(S1, T2) at the S1-T2 position, the corresponding sentence fluency value NSP(S2, T3) at the S2-T3 position, and so on, until the corresponding sentence fluency value NSP(Sm, Tm) at the Sm-Tm position is obtained. m-1 -T​n The corresponding statement fluency value (NSP) for the position m-1 T n );

[0061] And so on; until the first text column S1 of image A of Dunhuang manuscript fragments is combined with the nth text column T of image B of Dunhuang manuscript fragments. n After the top and bottom are aligned, S1-T n The corresponding statement fluency value NSP(S1, T) at the given position n );

[0062] Finally, the sentence fluency values ​​NSP(S1, T1), NSP(S2, T2), ..., NSP(S) are obtained for all aligned states of Dunhuang manuscript fragment image A and Dunhuang manuscript fragment image B. m T n ); NSP(S2, T1), NSP(S3, T2),…, NSP(S m T n-1 ); ..., NSP(S) m , T1); NSP (S1, T2), NSP (S2, T3), ..., NSP (S m-1 T n ); ..., NSP(S1, T n );

[0063] If Dunhuang manuscript fragment image A and Dunhuang manuscript fragment image B are of type a and b, type f and c, or type d and e respectively; then, following the method described above, sequentially obtain the sentence fluency values ​​NSP(T1, S1), NSP(T2, S2), ..., NSP(T... n S m ); NSP(T2, S1), NSP(T3, S2),…, NSP(T n S m-1 ); ..., NSP(T n , S1); NSP (T1, S2), NSP (T2, S3),..., NSP (T n-1 S m ); ..., NSP(T1, S m );

[0064] If image A of the Dunhuang manuscript fragment is of type g and image B is of type h, then the text column S of image A will be selected from the m-th column. m S obtained from the first text column T1 of image B of the Dunhuang manuscript fragment mThe clause to be predicted at position -T1 is input into the optimized and trained BERT language model to obtain the sentence fluency score NSP(S). m ,T1), at this time (S m T1) indicates that the current alignment state is text column S. m The text column T1 is distributed right to left;

[0065] If image A of the Dunhuang manuscript fragment is of type h and image B of type g, then the first text column S1 of image A and the nth text column T of image B will be used. n S1-T obtained from n The clause to be predicted at the given position is input into the optimized and trained BERT language model to obtain the sentence fluency score NSP(S1, T). n At this time (S1, T) n This indicates that the current alignment is between text column S1 and text column T. n Distributed left and right.

[0066] In step D, the ratio ratio The first ratio range is set to [0.8, 1.25]. w1, h1 and d1 are the column width, column height and spacing of the text column of image A of Dunhuang manuscript fragments, respectively. w2, h2 and d2 are the column width, column height and spacing of the text column of image B of Dunhuang manuscript fragments, respectively.

[0067] In step E, the second ratio ranges from [0.8 to 1.25], and the third ratio ranges from [0.8 to 1.25].

[0068] In step G2-1-4, the ratio of positive samples to negative samples is 1:3.

[0069] This invention comprehensively considers the sentence fluency and edge similarity of the text content of the Dunhuang manuscript fragments to be joined, thereby improving the efficiency and accuracy of joining Dunhuang manuscript fragment images. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0071] The present invention will now be described in detail with reference to the accompanying drawings and embodiments:

[0072] like Figure 1 As shown, the method for image joining of Dunhuang manuscript fragments based on sentence fluency according to the present invention includes the following steps:

[0073] A: According to the position and edge shape of the missing part of the Dunhuang Manuscript fragments, the image of the Dunhuang Manuscript fragments to be spliced is manually classified into the following missing types:

[0074] a type: upper missing; b type: lower missing; c type: left lower missing; d type: left upper missing; e type: right lower missing; f type: right upper missing; g type: left straight missing; h type: right straight missing; i type: left jagged missing; j type: right jagged missing;

[0075] In this embodiment, the position and edge shape of the missing part of the Dunhuang Manuscript fragments are manually judged:

[0076] When the edge is horizontally oriented and the edge horizontally spans several text columns in the Dunhuang Manuscript fragment image, if the missing part of the Dunhuang Manuscript fragment is missing up, it is determined to be a type: upper missing; if the missing part of the Dunhuang Manuscript fragment is missing down, it is determined to be b type: lower missing;

[0077] When the edge is vertically oriented and the edge does not span any text column in the Dunhuang Manuscript fragment image, that is, the edge is located at the gap on one side of the text column, if the missing part of the Dunhuang Manuscript fragment is missing left, it is determined to be g type: left straight missing; if the missing part of the Dunhuang Manuscript fragment is missing right, it is determined to be h type: right straight missing;

[0078] When the edge is vertically oriented and the edge is jagged and repeatedly spans one or more text columns in the Dunhuang Manuscript fragment image, if the missing part of the Dunhuang Manuscript fragment is missing left, it is determined to be i type: left jagged missing; if the missing part of the Dunhuang Manuscript fragment is missing right, it is determined to be j type: right jagged missing;

[0079] When the edge is obliquely oriented and the edge horizontally spans several text columns in the Dunhuang Manuscript fragment image, if the missing part of the Dunhuang Manuscript fragment is missing left down, it is determined to be c type: left lower missing; if the missing part of the Dunhuang Manuscript fragment is missing left up, it is determined to be d type: left upper missing; if the missing part of the Dunhuang Manuscript fragment is missing right down, it is determined to be e type: right lower missing; if the missing part of the Dunhuang Manuscript fragment is missing right up, it is determined to be f type: right upper missing;

[0080] B: Label each text column in the image of the Dunhuang Manuscript fragments to be spliced to obtain the minimum bounding rectangle and the text content of each text column; the width and height of the minimum bounding rectangle Q with the largest height value are taken as the column width w and the column height h, respectively, and the horizontal distance between the minimum bounding rectangle Q and the minimum bounding rectangle adjacent to Q is taken as the gap d;

[0081] In the present application, when obtaining the minimum circumscribed rectangle frame of the text column, all the characters in each text column are framed by the rectangle frame, and the upper edge of the rectangle frame is close to the top edge of the uppermost character in the text column, the lower edge of the rectangle frame is close to the bottom edge of the lowermost character in the text column, the left edge of the rectangle frame is close to the leftmost edge of all the characters in the text column, and the right edge of the rectangle frame is close to the rightmost edge of all the characters in the text column.

[0082] C: For the specified Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B to be spliced, first determine whether the missing types of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B correspond to each other, if the missing types correspond to each other, go to step D; if the missing types do not correspond to each other, continue to compare the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be spliced according to whether the missing types correspond to each other;

[0083] In the present embodiment, the a type corresponds to the b type, the c type corresponds to the f type, the d type corresponds to the e type, the g type corresponds to the h type, and the i type corresponds to the j type.

[0084] In the present embodiment, the missing type is used to first perform primary image filtering on the two Dunhuang Manuscript Fragment Images to be spliced, so that the two Dunhuang Manuscript Fragment Images that cannot be spliced can be quickly eliminated, the workload in the subsequent splicing process is reduced, and the splicing efficiency is improved.

[0085] D: Determine whether the ratio P1 of the column width to column height ratio of the Dunhuang Manuscript Fragment Image A to the column width to column height ratio of the Dunhuang Manuscript Fragment Image B and the ratio P2 of the column width gap ratio of the Dunhuang Manuscript Fragment Image A to the column width gap ratio of the Dunhuang Manuscript Fragment Image B are both within the set first ratio range; if the ratio P1 and the ratio P2 are both within the set first ratio range, go to step E; if not, return to step C and continue to compare the Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be spliced;

[0086] In the present embodiment, the ratio P1 of the column width to column height ratio of the Dunhuang Manuscript Fragment Image A to the column width to column height ratio of the Dunhuang Manuscript Fragment Image B is The ratio P2 of the column width gap ratio of the Dunhuang Manuscript Fragment Image A to the column width gap ratio of the Dunhuang Manuscript Fragment Image B is The set first ratio range is [0.8, 1.25], w1, h1 and d1 are respectively the column width, the column height and the gap of the text column of the Dunhuang Manuscript Fragment Image A, and w2, h2 and d2 are respectively the column width, the column height and the gap of the text column of the Dunhuang Manuscript Fragment Image B.

[0087] Considering the varying degrees of difference between the images of Dunhuang manuscripts captured by the imaging tools and the actual physical dimensions of the Dunhuang manuscripts, this embodiment employs a second image filtering process. This involves extracting the column width-to-height ratio and column width-to-gap ratio of two images of Dunhuang manuscripts to be joined, using the condition that the proportions of the same Dunhuang manuscript are basically consistent. This process can eliminate two images of Dunhuang manuscript fragments that clearly do not belong to the same fragment, reducing the workload in the subsequent joining process and improving joining efficiency.

[0088] E: Scale the image B of the Dunhuang manuscript fragments, and make the column width and spacing of the scaled Dunhuang manuscript fragment image B consistent with the column width and spacing of fragment image A.

[0089] When the column widths are consistent, calculate the ratio of the column height of the scaled Dunhuang manuscript fragment image B to the column height of the Dunhuang manuscript fragment image A, as well as the ratio of the gap between the scaled Dunhuang manuscript fragment image B and the gap between the Dunhuang manuscript fragment image A.

[0090] When the gaps are consistent, calculate the ratio of the column height of the scaled Dunhuang manuscript fragment image B to the column height of the Dunhuang manuscript fragment image A, and the ratio of the column width of the scaled Dunhuang manuscript fragment image B to the column width of the Dunhuang manuscript fragment image A.

[0091] If, when the column widths are consistent, the ratio of column height to column width is simultaneously within the set second ratio range, and when the column widths are consistent, the ratio of column height to column width is simultaneously within the set third ratio range, then proceed to step F; otherwise, return to step C.

[0092] In this invention, using the column width w1 of the Dunhuang manuscript fragment image A as a reference, the column width w2 of the Dunhuang manuscript fragment image B is scaled by a factor of r, so that the column width rw2 of the scaled Dunhuang manuscript fragment image B is consistent with the column width w1 of the Dunhuang manuscript fragment image A. Then, it is determined whether the ratio P3 of the column height rh2 of the scaled Dunhuang manuscript fragment image B to the column height h1 of the Dunhuang manuscript fragment image A, and the ratio P4 of the gap rd2 of the scaled Dunhuang manuscript fragment image B to the gap d1 of the Dunhuang manuscript fragment image A, are simultaneously within a set second ratio range. If they are simultaneously within the set second ratio range, the determination continues; if they are not simultaneously within the set second ratio range, the process returns to step C.

[0093] Then continue to take the gap d1 of Dunhuang Manuscript Fragment Image A as the benchmark, scale the gap d2 of Dunhuang Manuscript Fragment Image B by t times, so that the gap td2 of Dunhuang Manuscript Fragment Image B is consistent with the gap d1 of Dunhuang Manuscript Fragment Image A, and then judge whether the ratio P5 of the column height th2 of the scaled Dunhuang Manuscript Fragment Image B and the column height h1 of Dunhuang Manuscript Fragment Image A and the ratio P6 of the column width tw2 of the scaled Dunhuang Manuscript Fragment Image B and the column width w1 of Dunhuang Manuscript Fragment Image A are in the set third ratio range at the same time; if they are in the set third ratio range at the same time, go to step F; if they are not in the set third ratio range at the same time, return to step C;

[0094] In this embodiment, first scale the column width of Dunhuang Manuscript Fragment Image B by r times, so that w1 = rw2, and then judge whether and are in the set second ratio range at the same time; the second ratio range is [0.8, 1.25];

[0095] Then scale the gap of Dunhuang Manuscript Fragment Image B by t times, so that d1 = td2, and then judge whether and are in the set third ratio range at the same time; the third ratio range is [0.8, 1.25];

[0096] Considering the different degrees of differences between the Dunhuang Manuscript images captured by the shooting tool and the true physical size of the Dunhuang Manuscript, in this embodiment, by taking the column width and gap of Dunhuang Manuscript Fragment Image A as the benchmark, Dunhuang Manuscript Fragment Image B is restored to the same size as Dunhuang Manuscript Fragment Image A, and then the third image filtering is performed by taking whether the sizes of the two Dunhuang Manuscript Fragment images match as the screening condition, which can eliminate two Dunhuang Manuscript Fragment images that obviously do not belong to the same Dunhuang Manuscript Fragment, reduce the workload in the subsequent splicing process, and improve the splicing efficiency.

[0097] F: Train the handwriting judgment neural network model, and then input Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B into the trained handwriting judgment neural network model for judgment. If the similarity of the handwriting in Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B is greater than or equal to the set handwriting similarity threshold, go to step G; if it is less than the set handwriting similarity threshold, return to step C;

[0098] The handwriting judgment neural network model adopts a classic twin neural network model. The input of the model is a pair of images. The intermediate steps include extraction and fusion of features of the two images. The output is the similarity of the handwriting between the two images. When training the twin neural network model, positive and negative sample pairs need to be constructed. The construction of the positive sample pair includes two parts. One is each pair of images of the combined Dunhuang manuscript fragments. The other is to randomly divide each Dunhuang manuscript fragment image into two parts to form a positive sample pair. The negative sample pair is the combination between different Dunhuang manuscript fragment images of two types of missing types that do not correspond to each other. The number ratio of the positive sample pair to the negative sample pair is 1:1.

[0099] In this embodiment, the existing twin neural network model is used to perform the fourth image filtering by taking the similarity of the handwriting in the two Dunhuang manuscript fragment images as the screening condition. The Dunhuang manuscript fragment images with obviously different handwriting can be eliminated. The workload in the subsequent combining process is reduced, and the combining efficiency is improved.

[0100] G: First, the edge similarity of the Dunhuang manuscript fragment image A and the Dunhuang manuscript fragment image B is calculated to obtain the time sequence matching degree of the Dunhuang manuscript fragment image A and the Dunhuang manuscript fragment image B. Then, it is judged whether the missing type of the Dunhuang manuscript fragment image A belongs to type i or type j. If not, the Dunhuang manuscript fragment image A and the Dunhuang manuscript fragment image B are continuously subjected to the sentence fluency detection to obtain the sentence fluency value of the Dunhuang manuscript fragment image A and the Dunhuang manuscript fragment image B.

[0101] When the missing type of the Dunhuang manuscript fragment image A belongs to type i or type j, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, the Dunhuang manuscript fragment image B is taken as the candidate combining image. Then, step C is returned to continue comparing the Dunhuang manuscript fragment image A with the next Dunhuang manuscript fragment image to be combined. Otherwise, step C is directly returned to continue comparing the Dunhuang manuscript fragment image A with the next Dunhuang manuscript fragment image to be combined.

[0102] When the missing type of the Dunhuang manuscript fragment image A does not belong to type i or type j, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, and the sentence fluency value is greater than or equal to the set sentence fluency threshold value, the Dunhuang manuscript fragment image B is taken as the candidate combining image. Then, step C is returned to continue comparing the Dunhuang manuscript fragment image A with the next Dunhuang manuscript fragment image to be combined. Otherwise, step C is directly returned to continue comparing the Dunhuang manuscript fragment image A with the next Dunhuang manuscript fragment image to be combined.

[0103] When the edge similarity of the Dunhuang manuscript fragment image A and the Dunhuang manuscript fragment image B is calculated, the following steps are performed:

[0104] G1-1: Manually determine the reference lines at the upper, lower, left and right edges of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B, and the middle reference line adjacent to the left reference line, to obtain the reference reference image of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B;

[0105] G1-2: Obtain the position coordinates of the upper reference line U, the lower reference line D, the left reference line L, the right reference line R and the middle reference line M in the Dunhuang Manuscript Fragment Reference Reference Image obtained by computer positioning;

[0106] G1-3: Calculate the width of the grid unit in the Dunhuang Manuscript Fragment Reference Reference Image by computer;

[0107] G1-4: Obtain the real physical width of the grid unit according to the real physical size of Dunhuang Manuscript Fragment Image A, and then calculate the width value of the grid unit of the Dunhuang Manuscript Fragment Reference Reference Image corresponding to Dunhuang Manuscript Fragment Image A to restore the scaling ratio γ of the real physical size; wherein γ=β 2 , β is the multiple relationship between the width value of the grid unit of the Dunhuang Manuscript Fragment Image and the real physical width value of the grid unit of the Dunhuang Manuscript Fragment;

[0108] G1-5: Perform edge detection on Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B to extract the edge lines of Dunhuang Manuscript Fragment Image, to obtain the edge line image corresponding to Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B;

[0109] G1-6: Obtain the edge line skeleton in the edge line image corresponding to Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B by computer, to obtain the edge line skeleton image corresponding to each Dunhuang Manuscript Fragment Image, and the edge line skeleton refers to the central pixel point in the edge line;

[0110] G1-7: Manually determine the left and right broken edge parts in the edge line skeleton image of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B, to obtain the edge line skeleton annotation image corresponding to each Dunhuang Manuscript Fragment Image;

[0111] G1-8: The left and right broken edge parts of the edge line skeleton in the edge line skeleton annotation image obtained in step G1-7 are respectively processed by time series to obtain the corresponding two-dimensional numerical time series data;

[0112] G1-9: Use the scaling ratio γ and the multiple relationship β obtained in step G1-4 to obtain the position coordinates of the left reference line L: (l x, l y ), M: (m x , m y ), R: (r x , r y ), U: (u x , u y ) and D: (d x , d y ), into the position coordinate points L': (l' x , l' y ), M': (m' x , m' y ), R': (r' x , r' y ), U': (u' x , u' y ) and D': (d' x , d' y ) of the reference line after being restored to the real physical size; then the width G w of the grid cell obtained in step G1-3 is converted into the width G' w of the grid cell after being restored to the real physical size; and the two-dimensional time-sequenced data T l and T r corresponding to the broken edge portions on the left and right sides of the edge line skeleton obtained in step G1-8 are respectively converted into the two-dimensional time-sequenced data T' l and T' r corresponding to the broken edge portions on the left and right sides of the edge line skeleton after being restored to the real physical size; T' l = {(V' l1 , W' l1 ), (V' l2 , W' l2 ), (V' l3 , W' l3 ), …, (V' li , W' li )}, T' r = {(V' r1 , W' r1 ), (V' r2 , W' r2 ), (V' r3 , W' r3 ), …, (V' ri , W' ri )}, i is a positive integer, (V' li , W' li ) and (V' ri , W' ri )respectively represent the pixel positions of the i-th pixel data of the left and right broken edge parts after being restored to the true physical size;

[0113] G1-10: The two-dimensional time-sequenced data T' obtained in step G1-9 l V' in T' li and T' r V' in T' ri , and the position coordinate points L': (l' x , l' y ) and R': (r' x , r' y ) of the reference line are normalized, respectively, to obtain the normalized time-sequenced edge curve data T l " and T" r , and the normalized relative position coordinates L": (l" x , l" y ) and R": (r" x , r" y ) of the reference line, T" l = {(V" l1 , W" l1 ), (V" l2 , W' l2 ), (V" l3 , W" l3 ), …, (V" li , W" li )}, T" r = {(V" r1 , W" r1 ), (V" r2 , W" r2 ), (V" r3 , W" r3 ), …, (V" ri , W" ri )};

[0114] G1-11: For the to-be-concatenated Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B, the broken edge part time-sequenced edge curve data T" l and T" r of the edge line skeleton of the two Dunhuang Manuscript Fragment Images are obtained by normalization according to step G1-10, and then the time-sequenced matching degree s of the two Dunhuang Manuscript Fragment Images is calculated.

[0115] The specific details of the above steps belong to the prior art, which have been disclosed in detail in the granted patent with the application number CN202110440552.4 and the name of “Automatic Concatenation Method of Dunhuang Manuscript Fragment Image”, and will not be repeated here.

[0116] In the sentence fluency detection of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B, the following steps are performed:

[0117] G2-1: Construct a training set using an existing Dunhuang Manuscript Fragment Image annotation data set, and use the training set to fine-tune the BERT pre-training language model to obtain a fine-tuned BERT language model;

[0118] The step G2-1 includes the following specific steps:

[0119] G2-1-1: Obtain an existing Dunhuang Manuscript Fragment Image annotation data set, which is composed of a plurality of sentences;

[0120] In the Dunhuang Manuscript Fragment Image annotation data set, the text information in each sentence corresponds to the text content in a Dunhuang Manuscript Fragment Image.

[0121] G2-1-2: Convert the traditional Chinese characters in the Dunhuang Manuscript Fragment Image annotation data set to simplified Chinese characters, and then supplement the characters in the data set that do not appear in the BERT word table after de-duplication into the BERT word table;

[0122] In this embodiment, the OPENCC program can be used for traditional Chinese to simplified Chinese conversion;

[0123] G2-1-3: In the Dunhuang Manuscript Fragment Image annotation data set after traditional Chinese to simplified Chinese conversion, only keep each sentence in the paragraph whose number of characters is greater than or equal to a character number threshold, and establish a training set according to each retained sentence;

[0124] In this embodiment, the character number threshold is 4;

[0125] G2-1-4: According to the obtained training set, construct positive and negative samples required for BERT language model training;

[0126] First, input each sentence in the training set into a sentence reading system, and use the sentence reading system to mark the punctuation in each sentence;

[0127] The sentence reading system uses an existing ancient book automatic punctuation system developed by the Ancient Union Intelligent Data Research Room and the "Ancient Union-Beijing Normal University Joint Laboratory" based on different training methods. The system uses the unique 1.5 billion data of the "Chinese Classic Ancient Book Library" as the training set, and the model effect on the validation set is more than 92% in punctuation F1 value and more than 96% in sentence breaking F1 value.

[0128] Then, according to the entire sentence after marking the punctuation, the positive samples are constructed according to the following method:

[0129] (1) Find the period in the whole sentence, and divide the whole sentence into one or more sentences by one or more periods;

[0130] (2) Find the comma in each sentence;

[0131] If there is no comma in the whole sentence, randomly select the first several characters in the whole sentence as the first divided sentence of the positive sample, and the remaining characters as the second divided sentence;

[0132] If there is a comma in the whole sentence, divide the whole sentence into two or more sub-sentences by one or more commas; then divide the two or more sub-sentences in the whole sentence into the first divided sentence and the second divided sentence of the positive sample in order, wherein the first divided sentence of the positive sample contains at least one sub-sentence, and the second divided sentence of the positive sample contains at least one sub-sentence;

[0133] Finally, the CSV format positive sample dataset is constructed, and the expression of the positive sample is [S1, S2, 1]; wherein S1 represents the first divided sentence of the positive sample, S2 represents the second divided sentence of the positive sample, and the label 1 represents the positive sample;

[0134] For example, a certain sentence is composed as follows: SS1, SS2, SS3, SS4. SS5, SS6, SS7. SS8, SS9. Wherein SS1 to SS4 constitute the first sentence, SS5 to SS7 constitute the second sentence, SS8 and SS9 constitute the third sentence, and the first sentence to the third sentence constitute the sentence.

[0135] Then the sentence can exhaust the following positive samples:

[0136] [SS1, SS2+SS3+SS4], [SS1+SS2, SS3+SS4], [SS1+SS2+SS3, SS4], [SS1, SS2], [SS2, SS3], [SS3, SS4], [SS2+SS3, SS4], [SS2, SS3+SS4], [SS5, SS6+SS7], [SS5+SS6, SS7], [SS5, SS6], [SS6, SS7], [SS8, SS9];

[0137] Finally, the negative sample is constructed according to the following method:

[0138] (1) Randomly select sub-sentences from two different sentences as the first divided sentence and the second divided sentence of the negative sample, wherein the first divided sentence of the negative sample contains at least one sub-sentence, and the second divided sentence contains at least one sub-sentence;

[0139] (2) from the whole sentence containing a period, select the last one or more clauses in the sentence before the period as the first divided sentence of the negative sample, and then select the first one or more clauses in the sentence after the period as the second divided sentence of the negative sample;

[0140] Finally, the CSV format negative sample data set is constructed, and the expression of the negative sample is [S3, S4, 0]; wherein, S3 represents the first divided sentence of the negative sample, S4 represents the second divided sentence of the negative sample, and the label 0 represents the negative sample;

[0141] In this embodiment, the number ratio of the positive sample and the negative sample is 1:3.

[0142] G2-1-5: using the obtained positive sample and negative sample, the BERT pre-training language model is fine-tuned and trained, and finally the fine-tuned and trained BERT language model is obtained;

[0143] G2-2: using the fine-tuned and trained BERT language model, the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B are subjected to sentence fluency detection to obtain the sentence fluency value of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B;

[0144] The step G2-2 includes the following specific steps:

[0145] G2-2-1: extracting the character information on the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B; and then converting the traditional Chinese characters in the character information into simplified Chinese characters;

[0146] G2-2-2: according to the position of the character information in the Dunhuang Manuscript Fragment Image A, the character information of the Dunhuang Manuscript Fragment Image A after the traditional-to-simplified processing is sequentially divided into a plurality of text columns in the order of right-to-left and top-to-bottom, and a text column set S of the Dunhuang Manuscript Fragment Image A is constructed by using the text columns, S={S1, S2, …, Sm}, wherein, S1 to Sm respectively represent the 1st to mth column character information from right to left in the Dunhuang Manuscript Fragment Image A; m m

[0147] The Dunhuang Manuscript Fragment Image B is constructed into a text column set T of the Dunhuang Manuscript Fragment Image B by using the same method, T={T1, T2, …, Tn}, wherein, T1 to Tn respectively represent the 1st to n th column character information from right to left in the Dunhuang Manuscript Fragment Image B; n n

[0148] G2-2-3: calculating the maximum value of the sentence fluency of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B in various mutual position relationships and text column corresponding states;

[0149] ​​​​If Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B are of types b and a, c and f, or e and d, respectively; then the first column of text S1 of Dunhuang Manuscript Fragment Image A is corresponded with the first column of text T1 of Dunhuang Manuscript Fragment Image B, and then the characters in text column S1 and text column T1 are connected in order to form a string, and then the string is input into a sentence reading system for punctuation symbol annotation, and then it is determined whether the connection of the characters in text column S1 and text column T1 is annotated with a symbol:

[0150] If no symbol is annotated at the connection, a clause at the connection is selected, and the characters before and after the connection are taken as a to-be-predicted clause pair at S1-T1 position;

[0151] If the symbol annotated at the connection is a period or a comma, one clause before and after the period or the comma is selected as a to-be-predicted clause pair at S1-T1 position;

[0152] The to-be-predicted clause pair at S1-T1 position is input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S1, T1); then the to-be-predicted clause pair at S2-T2 position is obtained according to the above method and input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S2, T2), and so on, until the to-be-predicted clause pair at S m -T n position is obtained and input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S m , T n ); (S m , T n ) represents that the current alignment state is that text column S m is corresponded with text column T n ;

[0153] Then, the second column of text S2 of Dunhuang Manuscript Fragment Image A is corresponded with the first column of text T1 of Dunhuang Manuscript Fragment Image B according to the above method, the to-be-predicted clause pair at S2-T1 position is obtained and input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S2, T1); the to-be-predicted clause pair at S3-T2 position is obtained and input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S3, T2), and so on, until the to-be-predicted clause pair at S m -T n-1 position is obtained and input into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S m , T n-1 );

[0154] and so on; until the first column of the Dunhuang Manuscript Fragment Image A text column S m corresponding to the first column of the Dunhuang Manuscript Fragment Image B text column T1, and then S m -T1 position under the predicted clause pair and input into the BERT language model after tuning training, the sentence fluency value NSP(S m , T1) is obtained.

[0155] Similarly, according to the above method, after the first column of the Dunhuang Manuscript Fragment Image A text column S1 and the second column of the Dunhuang Manuscript Fragment Image B text column T2 are corresponded, the sentence fluency value NSP(S1, T2) corresponding to the S1-T2 position, the sentence fluency value NSP(S2, T3) corresponding to the S2-T3 position, …, the sentence fluency value NSP(S m-1 -T n position corresponding to the sentence fluency value NSP(S m-1 , T n ) is obtained.

[0156] Similarly, according to the above method, after the first column of the Dunhuang Manuscript Fragment Image A text column S1 and the second column of the Dunhuang Manuscript Fragment Image B text column T2 are corresponded, the sentence fluency value NSP(S1, T2) corresponding to the S1-T2 position, the sentence fluency value NSP(S2, T3) corresponding to the S2-T3 position, …, the sentence fluency value NSP(S n -T n position corresponding to the sentence fluency value NSP(S1, T n ) is obtained.

[0157] Finally, the sentence fluency values NSP(S1, T1), NSP(S2, T2), …, NSP(S m , T n ) of all alignment states of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B are obtained; NSP(S2, T1), NSP(S3, T2), …, NSP(S m , T n-1 ); …, NSP(S m , T1); NSP(S1, T2), NSP(S2, T3), …, NSP(S m-1 , T n ); …, NSP(S1, T n ).

[0158] If the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B are a type and b type, f type and c type, or d type and e type respectively; then according to the above method, the sentence fluency values NSP(T1, S1), NSP(T2, S2), …, NSP(T n , S m ) of all alignment states of the Dunhuang Manuscript Fragment Image B and the Dunhuang Manuscript Fragment Image A are obtained; NSP(T2, S1), NSP(T3, S2), …, NSP(Tn , S m-1 );..., NSP(T n , S1); NSP(T1, S2), NSP(T2, S3),..., NSP(T n-1 , S m );..., NSP(T1, S m );

[0159] If Dunhuang Manuscript Fragment Image A is of type g and Dunhuang Manuscript Fragment Image B is of type h, then the to-be-predicted clause at the S m -T1 position obtained from the first column of text S m 1 of Dunhuang Manuscript Fragment Image A and the first column of text T1 of Dunhuang Manuscript Fragment Image B is input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S m , T1), where (S m , T1) indicates that the current alignment state is that the text column S m is distributed to the right of the text column T1.

[0160] If Dunhuang Manuscript Fragment Image A is of type h and Dunhuang Manuscript Fragment Image B is of type g, then the to-be-predicted clause at the S1-T n position obtained from the first column of text S1 of Dunhuang Manuscript Fragment Image A and the nth column of text T n of Dunhuang Manuscript Fragment Image B is input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S1, T n ), where (S1, T n ) indicates that the current alignment state is that the text column S1 is distributed to the left of the text column T n .

[0161] G2-2-4: In the corresponding sentence fluency values obtained under various mutual position relationships and alignment states of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B, the highest value in the sentence fluency values is selected as the sentence fluency value of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B.

[0162] In this embodiment, the time sequence matching degree threshold value and the sentence fluency threshold value can be set according to experience. By fully considering the sentence fluency of the to-be-concatenated Dunhuang Manuscript Fragment Images at the concatenation position, interference items in which the missing parts and broken edge shapes are both matched but the semantics of the Dunhuang Manuscript Fragment Image characters are obviously irrelevant can be eliminated, greatly improving the efficiency and accuracy of Dunhuang Manuscript Fragment Image concatenation.

[0163] H: After Dunhuang Manuscript Fragment Image A is compared with all to-be-concatenated Dunhuang Manuscript Fragment Images, all candidate concatenated images of Dunhuang Manuscript Fragment Image A are obtained.

Claims

1. A method for image stitching of Dunhuang Manuscript fragments based on sentence well-formedness, characterized in that, The method comprises the following steps: A: According to the position and edge shape of the missing part of the Dunhuang Manuscript Fragment, the image of the Dunhuang Manuscript Fragment to be spliced is manually classified into the following missing types: a type: upper missing; b type: lower missing; c type: left lower missing; d type: left upper missing; e type: right lower missing; f type: right upper missing; g type: left straight missing; h type: right straight missing; i type: left jagged missing; j type: right jagged missing; B: Labeling each column of text in the image of the Dunhuang Manuscript Fragment to be spliced to obtain the minimum bounding rectangle and the text content of each column of text; taking the width and height of the minimum bounding rectangle Q with the maximum height value as the column width w and the column height h respectively, and taking the horizontal distance between the minimum bounding rectangle Q and the minimum bounding rectangle adjacent to Q as the gap d; C: For the specified Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B to be spliced, first determine whether the missing types of the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B correspond to each other, if the missing types correspond to each other, then proceed to step D; if the missing types do not correspond to each other, continue to compare the Dunhuang Manuscript Fragment image A with the next Dunhuang Manuscript Fragment image to be spliced according to whether the missing types correspond to each other; D: Determine whether the ratio P1 of the column width to column height ratio of the Dunhuang Manuscript Fragment image A and the column width to column height ratio of the Dunhuang Manuscript Fragment image B, and the ratio P2 of the column width gap ratio of the Dunhuang Manuscript Fragment image A and the column width gap ratio of the Dunhuang Manuscript Fragment image B, are both within the set first ratio range; if the ratio P1 and the ratio P2 are both within the set first ratio range, proceed to step E; if not, return to step C and continue to compare the Dunhuang Manuscript Fragment image A with the next Dunhuang Manuscript Fragment image to be spliced; E: Scale the Dunhuang Manuscript Fragment image B to make the column width and the gap of the scaled Dunhuang Manuscript Fragment image B consistent with the column width and the gap of the Dunhuang Manuscript Fragment image A in sequence; When the column width is consistent, calculate the ratio of the column height of the scaled Dunhuang Manuscript Fragment image B to the column height of the Dunhuang Manuscript Fragment image A, and the ratio of the gap of the scaled Dunhuang Manuscript Fragment image B to the gap of the Dunhuang Manuscript Fragment image A; When the gap is consistent, calculate the ratio of the column height of the scaled Dunhuang Manuscript Fragment image B to the column height of the Dunhuang Manuscript Fragment image A, and the ratio of the column width of the scaled Dunhuang Manuscript Fragment image B to the column width of the Dunhuang Manuscript Fragment image A; If the ratio of the column height and the ratio of the gap are both within the set second ratio range when the column width is consistent, and the ratio of the column height and the ratio of the column width are both within the set third ratio range when the gap is consistent, then proceed to step F; otherwise, return to step C; F: Train the handwriting judgment neural network model, and then input the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B into the trained handwriting judgment neural network model for judgment; if the similarity of the handwriting in the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B is greater than or equal to the set handwriting similarity threshold, proceed to step G; if less than the set handwriting similarity threshold, return to step C; G: first, the edge similarity of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B is calculated to obtain the time sequence matching degree of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B; then, it is judged whether the missing type of Dunhuang Manuscript Fragment Image A belongs to type i or type j; if not, the sentence fluency of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B is detected to obtain the sentence fluency value of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B; When the missing type of Dunhuang Manuscript Fragment Image A belongs to type i or type j, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, Dunhuang Manuscript Fragment Image B is taken as the candidate conjugation image, and then step C is returned to continue comparing Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; otherwise, step C is directly returned to continue comparing Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; When the missing type of Dunhuang Manuscript Fragment Image A does not belong to type i or type j, if the obtained time sequence matching degree is greater than or equal to the set time sequence matching degree threshold value, and the maximum value in the sentence fluency value is greater than or equal to the set sentence fluency threshold value, Dunhuang Manuscript Fragment Image B is taken as the candidate conjugation image, and then step C is returned to continue comparing Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; otherwise, step C is directly returned to continue comparing Dunhuang Manuscript Fragment Image A with the next Dunhuang Manuscript Fragment Image to be conjugated; H: when Dunhuang Manuscript Fragment Image A is compared with all the Dunhuang Manuscript Fragment Images to be conjugated, all the candidate conjugation images of Dunhuang Manuscript Fragment Image A are obtained. 2.The Dunhuang Manuscript Fragment Image Splicing Method Based on Sentence Well-formedness According to Claim 1, characterized in that: In step A, the position of the missing part of the Dunhuang Manuscript Fragment and the shape of the broken edge are determined by manual operation: When the broken edge is horizontal and the broken edge crosses several text columns in the Dunhuang Manuscript Fragment Image, if the missing part of the Dunhuang Manuscript Fragment is missing above, it is determined to be type a: missing above; if the missing part of the Dunhuang Manuscript Fragment is missing below, it is determined to be type b: missing below; When the broken edge is vertical and the broken edge does not cross any text column in the Dunhuang Manuscript Fragment Image, that is, the broken edge is located at the gap on one side of the text column, if the missing part of the Dunhuang Manuscript Fragment is missing left, it is determined to be type g: missing left straight; if the missing part of the Dunhuang Manuscript Fragment is missing right, it is determined to be type h: missing right straight; When the broken edge is vertical and the broken edge is jagged and repeatedly crosses one or more text columns in the Dunhuang Manuscript Fragment Image, if the missing part of the Dunhuang Manuscript Fragment is missing left, it is determined to be type i: missing left jagged; if the missing part of the Dunhuang Manuscript Fragment is missing right, it is determined to be type j: missing right jagged; When the broken edge is inclined and the broken edge crosses several text columns in the Dunhuang Manuscript Fragment image, if the missing part of the Dunhuang Manuscript Fragment is missing from the lower left, it is determined to be type c: missing from the lower left; if the missing part of the Dunhuang Manuscript Fragment is missing from the upper left, it is determined to be type d: missing from the upper left; if the missing part of the Dunhuang Manuscript Fragment is missing from the lower right, it is determined to be type e: missing from the lower right; if the missing part of the Dunhuang Manuscript Fragment is missing from the upper right, it is determined to be type f: missing from the upper right. 3.The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 1, characterized in that, In the step G, when the sentence fluency detection of the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B is performed, the following steps are performed: G2-1: a training set is constructed by using an existing Dunhuang Manuscript Fragment image annotation data set, and the BERT pre-training language model is fine-tuned and trained by using the training set to obtain the fine-tuned and trained BERT language model; G2-2: the fine-tuned and trained BERT language model is used to perform sentence fluency detection on the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B to obtain the sentence fluency values of the Dunhuang Manuscript Fragment image A and the Dunhuang Manuscript Fragment image B.

4. The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 3, characterized in that, The step G2-1 includes the following specific steps: G2-1-1: an existing Dunhuang Manuscript Fragment image annotation data set is obtained, and the Dunhuang Manuscript Fragment image annotation data set is composed of several sentences; G2-1-2: the traditional Chinese characters in the Dunhuang Manuscript Fragment image annotation data set are converted into simplified Chinese characters, and then the characters in the data set after the conversion from traditional Chinese characters to simplified Chinese characters that do not exist in the BERT word table are de-duplicated and supplemented into the BERT word table; G2-1-3: in the Dunhuang Manuscript Fragment image annotation data set after the conversion from traditional Chinese characters to simplified Chinese characters, only each sentence in the paragraph with the number of characters greater than or equal to the character number threshold is retained, and a training set is established according to each retained sentence; G2-1-4: according to the obtained training set, positive and negative samples required for training the BERT language model are constructed; G2-1-5: the obtained positive and negative samples are used to fine-tune and train the BERT pre-training language model, and finally the fine-tuned and trained BERT language model is obtained.

5. The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 4, characterized in that, The step G2-1-4 includes the following specific steps: First, each sentence in the training set is input into a sentence reading system, and the punctuation marks in the characters in each sentence are marked by using the sentence reading system; Then, according to the whole sentence after the punctuation marks are marked, the positive samples are constructed according to the following method: (1) find the period in the whole sentence, and divide the whole sentence into one or more sentences by one or more periods; (2) find the comma in each sentence; If there is no comma in the whole sentence, randomly select the first several characters in the whole sentence as the first divided sentence of the positive sample, and the remaining characters as the second divided sentence; If there is a comma in the whole sentence, divide the whole sentence into two or more sub-sentences by one or more commas; then divide the two or more sub-sentences in the whole sentence into the first divided sentence and the second divided sentence of the positive sample in order, wherein the first divided sentence of the positive sample contains at least one sub-sentence, and the second divided sentence of the positive sample contains at least one sub-sentence; Finally, the positive sample dataset in CSV format is constructed, and the expression of the positive sample is [S1, S2, 1]; wherein S1 represents the first divided sentence of the positive sample, S2 represents the second divided sentence of the positive sample, and label 1 represents the positive sample; Finally, the negative sample is constructed according to the following method: (1) randomly select clauses from two different sentences as the first divided sentence and the second divided sentence of the negative sample, wherein the first divided sentence of the negative sample contains at least one clause, and the second divided sentence contains at least one clause; (2) from the whole sentence containing a period, select the last one or more clauses in the sentence before the period as the first divided sentence of the negative sample, and then select the first one or more clauses in the sentence after the period as the second divided sentence of the negative sample; Finally, the negative sample dataset in CSV format is constructed, and the expression of the negative sample is [S3, S4, 0]; wherein S3 represents the first divided sentence of the negative sample, S4 represents the second divided sentence of the negative sample, and label 0 represents the negative sample.

6. The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 3, characterized in that, The step G2-2 includes the following specific steps: G2-2-1: extracting the character information on the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B; and then converting the traditional Chinese characters in the character information into simplified Chinese characters; G2-2-2: According to the position of the text information appearing in Dunhuang Manuscript Fragment Image A, the text information of Dunhuang Manuscript Fragment Image A after the traditional Chinese to simplified Chinese conversion is divided into a plurality of text columns in sequence from right to left and from top to bottom, and a text column set S of Dunhuang Manuscript Fragment Image A is constructed by using the text columns, S = {S1, S2, …, S m}, S1 to S m respectively represent the 1st column to the mth column of text information in Dunhuang Manuscript Fragment Image A from right to left. The image B of the Dunhuang Manuscript Fragment is constructed into a text column set T of the image B of the Dunhuang Manuscript Fragment by the same method, T={T1, T2, …, Tn}, T1 to Tn respectively represent the 1st column to the nth column of character information from right to left in the image B of the Dunhuang Manuscript Fragment. n}T1 to Tn respectively represent the 1st column to the nth column of character information from right to left in the image B of the Dunhuang Manuscript Fragment. n}T1 to Tn respectively represent the 1st column to the nth column of character information from right to left in the image B of the Dunhuang Manuscript Fragment. G2-2-3: calculating the maximum value of the sentence fluency of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B in various mutual position relationships and text column corresponding states; G2-2-4: selecting the highest value in the sentence fluency values of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B in various mutual position relationships and alignment states as the sentence fluency value of the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B.

7. The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 6, characterized in that, The step G2-2-3 includes the following specific steps: If the Dunhuang Manuscript Fragment Image A and the Dunhuang Manuscript Fragment Image B are of types b and a, types c and f, or types e and d, respectively; then the first column of text S1 of the Dunhuang Manuscript Fragment Image A is corresponded with the first column of text T1 of the Dunhuang Manuscript Fragment Image B, then the characters in the text column S1 and the text column T1 are connected in order to form a string, and then the string is input into a sentence reading system for punctuation symbol annotation, and then it is judged whether the connection position of the characters in the text column S1 and the text column T1 is annotated with a symbol: If no symbol is annotated at the connection position, the clause at the connection position is selected, and the characters before and after the connection position are selected as the to-be-predicted clause pair in the S1-T1 position; If the symbol annotated at the connection position is a period or a comma, one clause before and after the period or the comma is selected as the to-be-predicted clause pair in the S1-T1 position. The BERT language model after tuning and training is inputted with the to-be-predicted clause pair under the S1-T1 position to obtain a sentence fluency value NSP(S1, T1); then the to-be-predicted clause pair under the S2-T2 position is obtained according to the above method and inputted into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S2, T2), and so on, until the to-be-predicted clause pair under the S m -T n 2 position is obtained and inputted into the BERT language model after tuning and training to obtain a sentence fluency value NSP(S m 2, T n 2). (S m , T n ) represents that the current alignment state is that the text column S m corresponds to the text column T n . Then, the second column text column S2 of Dunhuang Manuscript Fragment Image A is corresponded with the first column text column T1 of Dunhuang Manuscript Fragment Image B according to the above method, the to-be-predicted clause pair at the position of S2-T1 is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S2, T1); the to-be-predicted clause pair at the position of S3-T2 is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S3, T2), …, the to-be-predicted clause pair at the position of S m -T n-1 is obtained and input into the BERT language model after tuning and training, to obtain the sentence fluency value NSP(S m , n-1 T ) By analogy; until the Dunhuang Manuscript Fragment Image A, the first column of the text column S m corresponding to the Dunhuang Manuscript Fragment Image B, the first column of the text column T1, and then S m -T1 position under the predicted clause pair and input into the BERT language model after tuning training, get the sentence fluency value NSP(S m , T1); Similarly, following the above method, after obtaining the text column S1 of image A of Dunhuang manuscript fragments and the text column T2 of image B of Dunhuang manuscript fragments, the sentence fluency values ​​NSP(S1, T2) corresponding to positions S1-T2, NSP(S2, T3) corresponding to positions S2-T3, ..., S m-1 -T n The corresponding statement fluency value (NSP) for the position m-1 T n ); and so on; until the first column text column S1 of Dunhuang Manuscript Fragment Image A is matched with the n-th column text column T of Dunhuang Manuscript Fragment Image B n After the up-down correspondence, S1-T n The statement fluency value NSP(S1, T n ) of the position-down correspondence Finally, the sentence fluency values NSP(S1, T1), NSP(S2, T2), …, NSP(S m , T n ) in all alignment states of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B are obtained. m , T n-1 ) in all alignment states of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B are obtained. m , T1) in all alignment states of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B are obtained. m-1 , T n ) in all alignment states of Dunhuang Manuscript Fragment Image A and Dunhuang Manuscript Fragment Image B are obtained. n , If Dunhuang manuscript fragment image A and Dunhuang manuscript fragment image B are of type a and b, type f and c, or type d and e respectively; then, following the method described above, sequentially obtain the sentence fluency values ​​NSP(T1, S1), NSP(T2, S2), ..., NSP(T... n S m ); NSP(T2, S1), NSP(T3, S2),…, NSP(T n S m-1 ); ..., NSP(T n , S1); NSP (T1, S2), NSP (T2, S3),..., NSP (T n-1 S m ); ..., NSP(T1, S m ); If image A of the Dunhuang manuscript fragment is of type g and image B is of type h, then the text column S of image A will be selected from the m-th column. m S obtained from the first text column T1 of image B of the Dunhuang manuscript fragment m The clause to be predicted at position -T1 is input into the optimized and trained BERT language model to obtain the sentence fluency score NSP(S). m ,T1), at this time (S m T1) indicates that the current alignment state is text column S. m The text column T1 is distributed right to left; If Dunhuang Manuscript Fragment Image A is of type h and Dunhuang Manuscript Fragment Image B is of type g, then the first column of text S1 from Dunhuang Manuscript Fragment Image A and the nth column of text Tn from Dunhuang Manuscript Fragment Image B are obtained n S1-Tn n under the position to be predicted, a BERT language model after inputting and tuning is trained, to obtain a sentence fluency value NSP(S1, Tn n ), at this time (S1, Tn n ) represents that the current alignment state is that the text column S1 and the text column Tn n are distributed on the left and right. 8.The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 1, characterized in that: The ratio in step D The ratio The first ratio range is set as [0.8, 1.25], w1, h1 and d1 are respectively the column width, column height and gap of the text column of the Dunhuang Manuscript Fragment Image A, and w2, h2 and d2 are respectively the column width, column height and gap of the text column of the Dunhuang Manuscript Fragment Image B.

9. The Dunhuang Manuscript Fragment Image Mosaic Method Based on Sentence Well-formedness According to Claim 1, characterized in that: In the step E, the second ratio range is [0.8, 1.25], and the third ratio range is [0.8, 1.25].

10. The Dunhuang Manuscript Fragment Image Splicing Method Based on Sentence Fluency According to Claim 4, characterized in that In the step G2-1-4, the number ratio of the positive sample and the negative sample is 1:3.

Citation Information

Patent Citations

  • Automatic conjugation method of Dunhuang book fragment image

    CN112991185A