Ancient book restoration management system containing mathematical symbols and restoration method

By combining multispectral data acquisition and linear unmixing methods with artificial intelligence to identify text and symbols in ancient book images, and using historical symbol templates for digital restoration, the problem of low efficiency in traditional ancient book restoration and difficulty in recognizing mathematical ancient books has been solved, achieving efficient ancient book image restoration and management.

CN120496078BActive Publication Date: 2025-12-16INNER MONGOLIA NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510590818.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-12-16
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional methods of ancient book restoration are inefficient and costly. Mathematical ancient books contain complex symbols that are difficult to identify effectively, which affects the quality of restoration.

Method used

The system employs a multispectral data acquisition module to obtain ultraviolet, visible, and near-infrared data of ancient book images. Combined with an end-member library module and artificial intelligence, it utilizes a linear unmixing method for text and symbol recognition and combines historical symbol templates for digital restoration.

Benefits of technology

It has improved the accuracy and precision of restoration and recognition of ancient mathematical texts, significantly enhanced the sensitivity of detection of faded text or symbols, and achieved effective digital management of ancient text images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496078B_ABST
    Figure CN120496078B_ABST
Patent Text Reader

Abstract

The application discloses an ancient book restoration management system containing mathematical symbols and a restoration method, and relates to the technical field of ancient book restoration management systems, and aims to improve the digital recognition accuracy of mathematical ancient books and improve the restoration quality of mathematical ancient books.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of image recognition, and particularly relates to a system for repairing and managing ancient books containing mathematical symbols and a restoration method. BACKGROUND

[0002] Ancient books are important carriers of human civilization, but due to the reasons of long time, limited preservation conditions, etc., a large number of ancient books are facing the problems of damage, blurred characters, etc., and need to be repaired. Traditional ancient book repair mainly relies on manual work, which is low in efficiency and high in cost. Benefiting from the rapid development of artificial intelligence, artificial intelligence can be used to realize effective digital repair of part of ancient books. However, compared with general text ancient books, mathematical ancient books contain a large number of complex mathematical symbols, and it is difficult to identify them, which easily leads to errors and affects the repair quality. SUMMARY

[0003] In order to solve the above problems existing in the prior art, the application provides a system for repairing and managing ancient books containing mathematical symbols and a restoration method.

[0004] The technical problem to be solved by the application is solved by the following technical scheme:

[0005] A system for repairing and managing ancient books containing mathematical symbols comprises:

[0006] A multi-spectral data acquisition module is configured to acquire images in ultraviolet, visible light and near-infrared bands of ancient book images to obtain multi-spectral data.

[0007] An end member library module is configured to store end member spectral data, wherein the end member spectral data comprises ink end member spectral data, paper end member spectral data, interference end member spectral data and symbol end member spectral data.

[0008] An identification module is configured to identify characters and symbols of the ancient book images by using artificial intelligence and linear unmixing method according to the multi-spectral data and the end member spectral data to obtain character symbol identification results.

[0009] A template management module is configured to store historical symbol templates.

[0010] An image repair module is configured to digitally repair the ancient book images according to the character symbol identification results and the historical symbol templates.

[0011] An ancient book management module is configured to manage digital ancient books according to the digitally repaired ancient book images.

[0012] Optionally, the symbol end member spectral data comprises straight line end member spectral data, circle end member spectral data, angle end member spectral data and point end member spectral data.

[0013] Optionally, the ink end-member spectral data comprises end-member spectral data of a plurality of different materials of ink.

[0014] Optionally, the ink end-member spectral data comprises: non-penetrating ink end-member spectral data, slightly penetrating ink end-member spectral data, moderately penetrating ink end-member spectral data, and heavily penetrating ink end-member spectral data.

[0015] Optionally, the identification module, according to the multi-spectral data and the end-member spectral data, uses artificial intelligence and linear unmixing method to perform character and symbol identification on the ancient book image, to obtain a character and symbol identification result, comprising:

[0016] According to the multi-spectral data, the ink end-member spectral data, the paper end-member spectral data, and the interference end-member spectral data, a full-constrained least squares model is used to solve the proportion of pixels corresponding to ink, paper, and interference in the multi-spectral data, to obtain a first solving result, and an ink identification result is obtained according to the first solving result;

[0017] According to the ink identification result, image segmentation is performed on the multi-spectral data to obtain segmented multi-spectral data, and artificial intelligence is used to identify the characters in the segmented multi-spectral data to obtain a character identification result;

[0018] According to the segmented multi-spectral data and the symbol end-member spectral data, a sparse-constrained least squares model is used to solve the proportion of pixels corresponding to symbols or non-symbols in the segmented multi-spectral data, to obtain a second solving result, and a symbol identification result is obtained according to the second solving result.

[0019] Optionally, the full-constrained least squares model is:

[0020]

[0021] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R1 is a first end-member matrix constructed according to the ink end-member spectral data, the paper end-member spectral data, and the interference end-member spectral data, A = {a k}, a k is an abundance map of the multi-spectral data corresponding to the kth end-member, is the (i,j)th element in the abundance map, m is the number of end-members contained in the first end-member matrix, and β is a parameter used to control the smoothing intensity.

[0022] Optionally, the sparse-constrained least squares model is:

[0023]

[0024] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R2 is a second endmember matrix constructed according to the symbol endmember spectral data, A = {a k} and a k is an abundance map of the kth endmember corresponding to the multi-spectral data, is an (i, j)th element in the abundance map, n is the number of symbol endmembers, and λ is a parameter for controlling sparsity.

[0025] The application also provides a method for restoring and repairing ancient books containing mathematical symbols, comprising:

[0026] obtaining multi-spectral data, wherein the multi-spectral data is obtained by image acquisition in ultraviolet, visible light and near-infrared bands on an image of an ancient book;

[0027] obtaining endmember spectral data and historical symbol templates, wherein the endmember spectral data comprises ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data;

[0028] performing character and symbol recognition on the image of the ancient book by using artificial intelligence and a linear unmixing method according to the multi-spectral data and the endmember spectral data, to obtain a character and symbol recognition result;

[0029] digitally restoring and repairing the image of the ancient book according to the character and symbol recognition result and the historical symbol templates.

[0030] Optionally, the symbol endmember spectral data comprises straight line endmember spectral data, circular endmember spectral data, angular endmember spectral data and point endmember spectral data.

[0031] Optionally, the ink endmember spectral data comprises endmember spectral data of ink of multiple different materials.

[0032] Optionally, the recognition module performs character and symbol recognition on the image of the ancient book by using artificial intelligence and a linear unmixing method according to the multi-spectral data and the endmember spectral data, to obtain a character and symbol recognition result, comprising:

[0033] solving the proportion of pixels corresponding to ink, paper and interference in the multi-spectral data by using a total least squares model according to the multi-spectral data, the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, to obtain a first solving result, and obtaining an ink recognition result according to the first solving result;

[0034] performing image segmentation on the multi-spectral data according to the ink recognition result, to obtain segmented multi-spectral data, and recognizing characters in the segmented multi-spectral data by using artificial intelligence, to obtain a character recognition result;

[0035] According to the segmented multi-spectral data and the symbol endmember spectral data, a sparse constraint least square model is used to solve the proportion of pixels corresponding to the symbol or non-symbol in the segmented multi-spectral data, to obtain a second solving result, and a symbol recognition result is obtained according to the second solving result.

[0036] Optionally, the full constraint least square model is:

[0037]

[0038] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R1 is a first endmember matrix constructed according to the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, A={a k} and a k is an abundance map of the multi-spectral data corresponding to the kth endmember, is an (i,j)th element in the abundance map, m is the number of endmembers contained in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

[0039] Optionally, the sparse constraint least square model is:

[0040]

[0041] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R2 is a second endmember matrix constructed according to the symbol endmember spectral data, B={b k} and b k is an abundance map of the multi-spectral data corresponding to the kth symbol endmember, is an (i,j)th element in the abundance map, n is the number of symbol endmembers, and λ is a parameter used to control the sparsity.

[0042] The ancient book repair management system containing mathematical symbols provided by the application has the following beneficial effects:

[0043] (1) The multi-spectral data is collected by ultraviolet / visible / near-infrared multi-band cooperation, the spectral feature differences of different materials (ink, paper and interference) are obtained, and the detection sensitivity of faded text or symbols is significantly improved.

[0044] (2) The endmember library module is set, the linear unmixing method is used to recognize the ancient book image based on the endmember spectral data, the mathematical symbols are effectively recognized, and the effective recognition of the text and symbols of the ancient book image is realized in combination with artificial intelligence.

[0045] (3) The historical symbol template is further combined with the character symbol recognition result, the ancient book image is digitally repaired, and the repair accuracy is improved.

[0046] The application will be further described in detail below with reference to the drawings. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the structure of an ancient book restoration and management system containing mathematical symbols provided in an embodiment of the present invention;

[0048] Figure 2 This is a flowchart illustrating a method for restoring and reconstructing ancient books containing mathematical symbols, provided in an embodiment of the present invention. Detailed Implementation

[0049] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0050] To improve the accuracy of digital recognition of ancient mathematical texts and enhance the quality of their restoration, this invention provides a system for the restoration and management of ancient mathematical texts that includes mathematical symbols. (See attached document.) Figure 1 As shown, the system includes: a multispectral data acquisition module, an end-member library module, a recognition module, a template management module, an image restoration module, and an ancient book management module.

[0051] The multispectral data acquisition module is used to acquire images of ancient books in the ultraviolet, visible, and near-infrared bands to obtain multispectral data.

[0052] Specifically, images of ancient books were captured using cameras in the ultraviolet, visible, and near-infrared bands, resulting in multispectral data. When capturing ultraviolet images, visible light interference must be disabled; when capturing visible light images, a D65 white light source is preferred; and when capturing near-infrared images, an infrared filter can be added to the camera used for visible light photography. Furthermore, the images captured by the cameras can be standardized to the same resolution, ensuring that the images in different bands have the same resolution in the final multispectral data.

[0053] The endmember library module is used to store endmember spectral data, which includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data, and symbol endmember spectral data.

[0054] In the present application, the end member refers to the spectral characteristics of pure substances or pure components constituting the mixed pixels in the image, which can be compared to the "basic primary colors" on a color palette. It can be understood that for a mathematical ancient book image, its spectral composition may include one or more characteristics of ink, paper, and interference (mold spots, stains, creases, etc.), and using conventional OCR (optical character recognition) tools may incorrectly identify the interference as ink, leading to recognition errors. Therefore, the present application introduces interference end members to effectively distinguish interference from ink in the ancient book image. In addition, the special structure of the symbols in the mathematical ancient book image also makes it difficult to be recognized by conventional OCR tools. For example, small symbols such as "·" and "∴" may be recognized as noise by conventional OCR recognition tools, and some similar symbols "⊥" may be incorrectly recognized as text by conventional OCR recognition tools. Therefore, the present application introduces symbol end members to effectively distinguish symbols from text.

[0055] Regarding the acquisition of end member spectral data, the average spectrum of clear ink pixels in the ancient book can be selected as the ink end member spectral data, the average spectrum of a location without words in the ancient book can be selected as the paper end member spectral data, the average spectrum of a location with mold spots in the ancient book can be selected as the mold spot end member spectral data, the average spectrum of a location with stains in the ancient book can be selected as the stain end member spectral data, and the average spectrum of a location with different types of symbols in the ancient book can be selected as the end member spectral data of the corresponding type of symbol.

[0056] The recognition module is configured to perform text and symbol recognition on the ancient book image by using artificial intelligence and a linear unmixing method based on the multi-spectral data and the end member spectral data, to obtain a text and symbol recognition result.

[0057] Specifically, the recognition module performs text and symbol recognition on the ancient book image by using artificial intelligence and a linear unmixing method based on the multi-spectral data and the end member spectral data, to obtain a text and symbol recognition result, including:

[0058] Step one, according to the multi-spectral data, ink end member spectral data, paper end member spectral data, and interference end member spectral data, a full constraint least squares model is used to solve the proportion of pixels in the multi-spectral data corresponding to ink, paper, and interference, to obtain a first solving result, and an ink recognition result is obtained based on the first solving result;

[0059] In this step S10, the full constraint least squares model can be:

[0060]

[0061] In formula (1), I is a multi-spectral tensor constructed according to multi-spectral data, the dimension of which is CxLxH, C is the number of wave bands, C=3 in the present application, LxH represents the size of a single wave band spectral image. R1 is a first end member matrix constructed according to ink end member spectral data, paper end member spectral data and interference end member spectral data, the dimension of which is mxC, m is the number of end members contained in the first end member matrix, for example, if the end members only include ink, paper and mold spots, then m=3. A={a k}, the dimension of which is mxLxH, a k is the abundance map (size LxH) of the kth end member corresponding to the multi-spectral data, is the (i,j)th element in the abundance map, and β is a parameter used to control the smoothing intensity, preferably 0.05, but of course not limited thereto.

[0062] In the full constraint least square model, in view of the possible stroke breakage in the ancient book image, a spatial continuity constraint is introduced on the basis of the conventional full constraint least square model, that is, a spatial smoothing term (the second term in formula (1)) is introduced in the objective function of the conventional full constraint least square model, so as to force the adjacent pixel abundances to be similar, thereby effectively identifying the stroke breakage.

[0063] After A is solved based on formula (1) and the first solving result, the specific value of a in A can be used to judge that the (i,j)th element specifically corresponds to ink, paper or interference, thereby obtaining the ink recognition result.

[0064] Step two, image segmentation is performed on the multi-spectral data according to the ink recognition result, to obtain segmented multi-spectral data, and the segmented multi-spectral data is recognized by artificial intelligence to obtain a character recognition result.

[0065] Here, the image segmentation performed on the multi-spectral data according to the ink recognition result is to separate the region determined as ink from its background region, which can remove the influence of interference and yellowing paper on image quality, that is, mold spots, stains and yellowing paper can all be converted into clean background through this operation. Then, the characters in the segmented multi-spectral data are further recognized by artificial intelligence, and the artificial intelligence here can adopt a conventional OCR tool, but of course is not limited thereto.

[0066] In addition, since the conventional OCR tool may incorrectly recognize some symbols as characters, in practice, the specific data recognized by the OCR can be analyzed, and the recognized characters with low reliability are not used, and only the recognized characters with high reliability are retained. In addition, it can be understood that the symbols that may be recognized as noise by the conventional OCR recognition tool will not be recognized as characters by the OCR recognition tool.

[0067] Step 3: Based on the segmented multispectral data and symbol endmember spectral data, use the sparse constrained least squares model to solve for the proportion of pixels in the segmented multispectral data corresponding to symbols or non-symbols, obtain the second solution result, and obtain the symbol recognition result based on the second solution result.

[0068] Specifically, after step S20, the text in the ink has been identified. Therefore, the remaining ink images that were not identified as text can be further segmented from the segmented multispectral data to form further segmented multispectral data. Then, the solution is obtained using a sparse constrained least squares model based on the further segmented multispectral data and the symbol endmember spectral data.

[0069] The sparse-constrained least squares model is as follows:

[0070]

[0071] Where I is the multispectral tensor constructed from the further segmented multispectral data, and its dimension is also C×L×H. R2 is the second endmember matrix constructed from the symbolic endmember spectral data, and its dimension is (n+1)×C, where n is the number of symbolic endmembers, and B={b k}, b k For the abundance plot of multispectral data corresponding to the k-th symbol endmember, For the (i,j)th element in this abundance graph, when k=0, b k The (i,j)th element in the formula does not correspond to any symbol, and λ is a parameter used to control sparsity, preferably 0.1, but it is not limited to this.

[0072] After solving B (i.e., the second solution result) based on equation (2), according to B... The specific value can be used to determine whether the (i,j)th element corresponds to a symbol endmember and what kind of symbol endmember it corresponds to, thus obtaining a preliminary symbol recognition result. That is, it is preliminarily determined which positions of the ink marks remaining after removing the text in the ancient book image correspond to symbols, and what kind of stroke structure these symbols correspond to (determined by the corresponding symbol endmember).

[0073] The template management module is used to store historical symbol templates.

[0074] Here, the template management module is actually a historical symbol template library, which stores a large number of historical symbol templates, which are obtained in the following manner: first, a high-precision scanned image of a mathematical ancient book is obtained, the image is enhanced by using an image enhancement method, and then a symbol image is extracted from the enhanced image by a mathematical ancient book research expert, and then based on the extracted symbol image, the symbol image is vectorized by Illustrator (Adobe Illustrator) or Inkscape, and the vectorized symbol is used to represent the geometric characteristics of the symbol. Inkscape is a free and open source vector graphics editor.

[0075] The image repair module is used for digital repair of the ancient book image according to the character symbol recognition result and the historical symbol template.

[0076] Specifically, for the symbol (assuming Sig) preliminarily recognized in the character symbol recognition result, the Illustrator (Adobe Illustrator) or Inkscape is also used to vectorize Sig to represent it; and according to the symbol end member corresponding to Sig, the historical symbol template containing the symbol end member is screened from the historical symbol template stored in the template management module, and the vectorized representation of Sig is matched with the screened historical symbol template (realized based on vector similarity calculation), so as to finally determine that Sig is what kind of symbol. Therefore, the symbols and characters in the ancient book image have been recognized, so that a digital restoration image can be generated for the ancient book image according to the final recognition result at this time, and a repair-like effect can be achieved.

[0077] The ancient book management module is used for digital ancient book management according to the ancient book image that has been digitally repaired.

[0078] Specifically, the digital ancient book management according to the ancient book image that has been digitally repaired can include: binding the ancient book image that has been digitally repaired into a book, adding page numbers, abstracts, covers and other information; storing the mathematical ancient book in categories, adding related search and browsing functions, and the specific management mode can be determined according to the use requirements of the mathematical ancient book, and a corresponding management scheme can be formulated, which will not be described herein.

[0079] The ancient book repair management system with mathematical symbols provided by the application adopts ultraviolet / visible / near-infrared multi-band cooperative collection of multi-spectral data, and obtains the spectral characteristic differences of different materials (ink, paper and interference), so that the detection sensitivity of faded text or symbols is significantly improved. The application sets an end member library module, recognizes the ancient book image based on end member spectral data by using a linear unmixing method, effectively identifies mathematical symbols, and realizes effective identification of text and symbols of the ancient book image in combination with artificial intelligence. In addition, the application further combines historical symbol templates and text symbol recognition results to digitally repair the ancient book image, thereby improving the repair accuracy.

[0080] Based on the same inventive concept, the application further provides an ancient book repair and restoration method with mathematical symbols, which is described in detail in the description of the ancient book repair management system with mathematical symbols. Figure 2 The method comprises the following steps:

[0081] S10, multi-spectral data is obtained, which is obtained by collecting images in ultraviolet, visible and near-infrared bands on the ancient book image;

[0082] S20, end member spectral data and historical symbol templates are obtained; the end member spectral data comprises ink end member spectral data, paper end member spectral data, interference end member spectral data and symbol end member spectral data;

[0083] S30, according to the multi-spectral data and the end member spectral data, text and symbol recognition of the ancient book image is performed by using artificial intelligence and a linear unmixing method, and text symbol recognition results are obtained;

[0084] S40, the ancient book image is digitally repaired and restored according to the text symbol recognition results and the historical symbol templates.

[0085] Optionally, the symbol end member spectral data comprises straight line end member spectral data, circle end member spectral data, angle end member spectral data and point end member spectral data.

[0086] Optionally, the ink end member spectral data comprises end member spectral data of ink of multiple different materials. Optionally, the ink end member spectral data comprises non-penetrating ink end member spectral data, slightly penetrating ink end member spectral data, moderately penetrating ink end member spectral data and heavily penetrating ink end member spectral data.

[0087] Optionally, the recognition module, according to the multi-spectral data and the end member spectral data, performs text and symbol recognition of the ancient book image by using artificial intelligence and a linear unmixing method, and obtains text symbol recognition results, which comprises:

[0088] According to the multi-spectral data, the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, a full constraint least square model is used to solve the proportion of pixels corresponding to the ink, the paper and the interference in the multi-spectral data, to obtain a first solving result, and an ink recognition result is obtained according to the first solving result.

[0089] According to the ink recognition result, image segmentation is performed on the multi-spectral data to obtain segmented multi-spectral data, and an artificial intelligence is used to recognize the text in the segmented multi-spectral data to obtain a text recognition result.

[0090] According to the segmented multi-spectral data and the symbol endmember spectral data, a sparse constraint least square model is used to solve the proportion of pixels corresponding to the symbol or the non-symbol in the segmented multi-spectral data to obtain a second solving result, and a symbol recognition result is obtained according to the second solving result.

[0091] Optionally, the full constraint least square model is:

[0092]

[0093] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R1 is a first endmember matrix constructed according to the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, A={a k} and a k is an abundance map of the multi-spectral data corresponding to the kth endmember, is the (i,j)th element in the abundance map, m is the number of endmembers contained in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

[0094] Optionally, the sparse constraint least square model is:

[0095]

[0096] wherein I is a multi-spectral tensor constructed according to the multi-spectral data, R2 is a second endmember matrix constructed according to the symbol endmember spectral data, B={b k} and b k is an abundance map of the multi-spectral data corresponding to the kth symbol endmember, is the (i,j)th element in the abundance map, n is the number of symbol endmembers, and λ is a parameter used to control the sparsity.

[0097] It should be noted that, for the method embodiment, since it is basically similar to the system embodiment, the description is relatively simple, and the related parts refer to the part of the system embodiment, and the method embodiment can achieve the same or similar beneficial effects as the system embodiment.

[0098] It should be noted that the terms "first", "second", and so on do not necessarily signify a particular order or sequence, but are used to distinguish similar objects. It should be understood that the data as used herein can be interchanged, where appropriate, so that the embodiments of the application described herein can be implemented in other than the order depicted or described herein. The implementations described in the following example embodiments are not meant to represent all implementations consistent with the present application. Rather, they are simply examples of apparatuses and methods consistent with some aspects of the present application.

[0099] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific feature or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. In the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Also, the specific feature or characteristic described can be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification.

[0100] Although the present application is described herein in conjunction with various embodiments, those skilled in the art, with the benefit of the description presented herein, will appreciate that other variations of the described embodiments are possible. In the description of the present application, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "multiple" means two or more, unless otherwise expressly specified. In addition, some measures described in different embodiments can be combined to produce good results.

[0101] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, apparatuses (devices), or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects, all of which are referred to herein as "modules" or "systems". Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The computer program can be stored / distributed in a suitable medium, provided with other hardware or as part of hardware, and can also take other distribution forms, such as through the Internet or other wired or wireless telecommunications systems.

[0102] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flows and / or block or blocks in the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams.

[0103] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart or flows and / or block or blocks of the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams.

[0104] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flows and / or block or blocks in the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams. Figure 1 one or more functions specified in the flowchart or flows and / or block or blocks of the block diagrams.

[0105] The above description is further to the present application in conjunction with specific preferred embodiments, and cannot be deemed to limit the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of simple deductions or replacements can also be made, which should be deemed to fall within the protection scope of the present application.

Claims

1. A system for the restoration and management of ancient books containing mathematical symbols, characterized in that, include: The multispectral data acquisition module is used to acquire images of ancient books in the ultraviolet, visible, and near-infrared bands to obtain multispectral data. The endmember library module is used to store endmember spectral data; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data, and symbol endmember spectral data. The recognition module is used to perform text and symbol recognition on the ancient book image based on the multispectral data and the endmember spectral data, using artificial intelligence and linear unmixing methods to obtain text and symbol recognition results. This includes: using a fully constrained least squares model to solve for the proportion of pixels in the multispectral data corresponding to ink, paper, and interference based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, obtaining a first solution result, and obtaining an ink recognition result based on the first solution result; performing image segmentation on the multispectral data based on the ink recognition result, obtaining segmented multispectral data, and using artificial intelligence to recognize the text in the segmented multispectral data, obtaining a text recognition result; and using a sparse constrained least squares model to solve for the proportion of pixels in the segmented multispectral data corresponding to symbols or non-symbols based on the segmented multispectral data and the symbol endmember spectral data, obtaining a second solution result, and obtaining a symbol recognition result based on the second solution result. The template management module is used to store historical symbol templates; The image restoration module is used to digitally restore the ancient book image based on the text symbol recognition results and the historical symbol template; The ancient book management module is used for digital management of ancient books based on the digitally restored images of ancient books.

2. The ancient book restoration and management system containing mathematical symbols according to claim 1, characterized in that, The symbol endmember spectral data includes: straight endmember spectral data, round endmember spectral data, corner endmember spectral data, and point endmember spectral data.

3. The ancient book restoration and management system containing mathematical symbols according to claim 1, characterized in that, The ink end-member spectral data includes end-member spectral data of inks made of various materials.

4. The ancient book restoration and management system containing mathematical symbols according to claim 1, characterized in that, The ink end-member spectral data includes: end-member spectral data of ink with no penetration, end-member spectral data of ink with slight penetration, end-member spectral data of ink with moderate penetration, and end-member spectral data of ink with heavy penetration.

5. The ancient book restoration management system containing mathematical symbols according to claim 1, characterized in that, The fully constrained least squares model is as follows: ; in, The multispectral tensor constructed based on the multispectral data, The first endmember matrix is ​​constructed based on the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data. , The multispectral data corresponds to the first Abundance plot of endogenous species, The first in this abundance diagram One element, The number of endmembers in the first endmember matrix. These are parameters used to control the smoothing intensity.

6. The ancient book restoration and management system containing mathematical symbols according to claim 1, characterized in that, The sparse-constrained least squares model is as follows: ; in, The multispectral tensor constructed based on the multispectral data, This is the second endmember matrix constructed based on the symbolic endmember spectral data. , The multispectral data corresponds to the first Abundance plot of symbolic endmembers, The first in this abundance diagram One element, For the number of sign-endons, These are parameters used to control sparsity.

7. A method for restoring and reconstructing ancient books containing mathematical symbols, characterized in that, include: Acquire multispectral data; the multispectral data is obtained by acquiring images of ancient books in the ultraviolet, visible, and near-infrared bands. Acquire endmember spectral data and historical symbol templates; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data, and symbol endmember spectral data; Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve for the proportion of pixels in the multispectral data corresponding to ink, paper, and interference, obtaining a first solution result. An ink stain recognition result is then obtained based on this first solution result. Based on the ink stain recognition result, the multispectral data is segmented to obtain segmented multispectral data. Artificial intelligence is then used to recognize the text in the segmented multispectral data, obtaining a text recognition result. Based on the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve for the proportion of pixels in the segmented multispectral data corresponding to symbols or non-symbols, obtaining a second solution result. A symbol recognition result is then obtained based on this second solution result. The ancient book image is digitally restored based on the text symbol recognition results and the historical symbol template.

8. The method for restoring and reconstructing ancient books containing mathematical symbols according to claim 7, characterized in that, The symbol endmember spectral data includes: straight endmember spectral data, round endmember spectral data, corner endmember spectral data, and point endmember spectral data.

9. The method for restoring and reconstructing ancient books containing mathematical symbols according to claim 7, characterized in that, The ink end-member spectral data includes end-member spectral data of inks made of various materials.

Citation Information

Patent Citations

  • Ancient tomb inscription character recognition method based on hyperspectral remote sensing technology

    CN111539409A

  • Mathematical ancient book illustration identification method and device, electronic equipment and storage medium

    CN119649375A