Mathematical symbol-containing ancient book restoration management system and restoration method

Through the method of multispectral data acquisition and artificial intelligence recognition combined with historical symbol templates, the problems of low repair efficiency and difficulty in recognition of ancient mathematical books are solved, and high-precision digital repair is achieved.

CN120496078AActive Publication Date: 2025-08-15INNER MONGOLIA NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510590818.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional ancient books have low efficiency and high cost to repair. Because ancient mathematical books contain a large number of complex mathematical symbols, they are difficult to identify and prone to errors, which affects the quality of repair.

Method used

The multi-spectral data acquisition module is used to obtain image data in the ultraviolet, visible light and near-infrared bands, combine the terminal element library module and artificial intelligence to identify text and symbols, and use historical symbol templates for digital repair.

Benefits of technology

It improves the digital recognition accuracy and repair accuracy of ancient mathematical books, and significantly improves the detection sensitivity of faded text or symbols.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496078A_ABST
    Figure CN120496078A_ABST
Patent Text Reader

Abstract

The invention discloses a mathematical symbol-containing ancient book restoration management system and restoration method, and the system comprises a multispectral data collection module which is used for carrying out the image collection of an ultraviolet wave band, a visible light wave band and a near-infrared wave band of an ancient book image, and obtaining multispectral data; the end member library module is used for storing end member spectral data; the recognition module is used for carrying out character and symbol recognition on the ancient book image by utilizing artificial intelligence and a linear unmixing method according to the multispectral data and the end member spectral data to obtain a character and symbol recognition result; the template management module is used for storing a historical symbol template; the image restoration module is used for performing digital restoration on the ancient book image according to a character symbol recognition result and a historical symbol template; and the ancient book management module is used for carrying out digital ancient book management according to the digitally restored ancient book image. According to the method, the digital identification precision of the mathematical ancient books can be improved, and the repairing quality of the mathematical ancient books is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and in particular relates to a restoration management system and a restoration method for ancient books containing mathematical symbols. Background Art

[0002] Ancient books are important carriers of human civilization. However, due to their age and limited preservation conditions, many ancient books are damaged, with blurred characters, and in urgent need of restoration. Traditional restoration of ancient books relies mainly on manual labor, which is inefficient and costly. Thanks to the rapid development of artificial intelligence, some ancient books can be effectively digitally restored using AI. However, compared to general textual ancient books, mathematical ancient books contain a large number of complex mathematical symbols, making their identification particularly difficult and prone to errors, which affects the quality of restoration. Summary of the Invention

[0003] In order to solve the above problems existing in the prior art, the present invention provides an ancient book restoration management system and restoration method containing mathematical symbols.

[0004] The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0005] An ancient book restoration and management system containing mathematical symbols, comprising:

[0006] The multispectral data acquisition module is used to collect images of ancient books in the ultraviolet band, visible light band and near-infrared band to obtain multispectral data;

[0007] An endmember library module is used to store endmember spectral data; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data;

[0008] a recognition module, configured to perform text and symbol recognition on the ancient book image using artificial intelligence and a linear unmixing method based on the multispectral data and the endmember spectral data, to obtain a text and symbol recognition result;

[0009] Template management module, used to store historical symbol templates;

[0010] An image restoration module, configured to digitally restore the ancient book image based on the text symbol recognition result and the historical symbol template;

[0011] The ancient book management module is used to manage digital ancient books based on the digitally restored ancient book images.

[0012] Optionally, the symbolic endmember spectral data includes: straight endmember spectral data, circular endmember spectral data, angle endmember spectral data and point endmember spectral data.

[0013] Optionally, the ink end-member spectral data includes end-member spectral data of inks of multiple different materials.

[0014] Optionally, the ink end-member spectrum data includes: no-penetration ink end-member spectrum data, light-penetration ink end-member spectrum data, moderate-penetration ink end-member spectrum data, and heavy-penetration ink end-member spectrum data.

[0015] Optionally, the recognition module performs text and symbol recognition on the ancient book image based on the multispectral data and the endmember spectral data using artificial intelligence and a linear unmixing method to obtain a text and symbol recognition result, including:

[0016] Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve the ratio of pixels in the multispectral data corresponding to ink, paper, and interference to obtain a first solution result, and an ink recognition result is obtained based on the first solution result;

[0017] Performing image segmentation on the multispectral data according to the ink recognition result to obtain segmented multispectral data, and using artificial intelligence to recognize text in the segmented multispectral data to obtain a text recognition result;

[0018] According to the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve the ratio of pixels in the segmented multispectral data corresponding to symbols or non-symbols to obtain a second solution result, and a symbol recognition result is obtained according to the second solution result.

[0019] Optionally, the fully constrained least squares model is:

[0020]

[0021] Wherein, I is a multispectral tensor constructed based on the multispectral data, R1 is a first endmember matrix constructed based on the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, A={a k}, a k is the abundance map of the multispectral data corresponding to the k-th end member, is the (i, j)th element in the abundance map, m is the number of endmembers in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

[0022] Optionally, the sparse constrained least squares model is:

[0023]

[0024] Wherein, I is a multispectral tensor constructed based on the multispectral data, R2 is a second endmember matrix constructed based on the symbolic endmember spectral data, and A={a k}, a k is the abundance map of the multispectral data corresponding to the k-th end member, is the (i, j)th element in the abundance map, n is the number of symbolic endmembers, and λ is a parameter used to control sparsity.

[0025] The present invention also provides a method for repairing and restoring ancient books containing mathematical symbols, comprising:

[0026] Acquiring multispectral data; the multispectral data is obtained by collecting images of ancient books in ultraviolet bands, visible light bands, and near-infrared bands;

[0027] Acquiring endmember spectral data and historical symbol templates; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data;

[0028] According to the multispectral data and the endmember spectral data, using artificial intelligence and linear unmixing method to perform text and symbol recognition on the ancient book image to obtain a text and symbol recognition result;

[0029] The ancient book image is digitally restored and repaired based on the text symbol recognition result and the historical symbol template.

[0030] Optionally, the symbolic endmember spectral data includes: straight endmember spectral data, circular endmember spectral data, angle endmember spectral data and point endmember spectral data.

[0031] Optionally, the ink end-member spectral data includes end-member spectral data of inks of multiple different materials.

[0032] Optionally, the recognition module uses artificial intelligence and linear unmixing methods to perform text and symbol recognition on the ancient book image based on the multispectral data and the endmember spectral data, and obtains text and symbol recognition results, including:

[0033] Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve the ratio of pixels in the multispectral data corresponding to the ink, paper, and interference to obtain a first solution result, and an ink recognition result is obtained based on the first solution result;

[0034] Perform image segmentation on the multispectral data according to the ink recognition result to obtain segmented multispectral data, and use artificial intelligence to recognize text in the segmented multispectral data to obtain text recognition results;

[0035] According to the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve the ratio of pixels in the segmented multispectral data corresponding to symbols or non-symbols to obtain a second solution result, and a symbol recognition result is obtained according to the second solution result.

[0036] Alternatively, the fully constrained least squares model is:

[0037]

[0038] Where I is a multispectral tensor constructed based on multispectral data, R1 is a first endmember matrix constructed based on ink endmember spectral data, paper endmember spectral data and interference endmember spectral data, A={a k}, a k is the abundance map of the multispectral data corresponding to the k-th end member, is the (i, j)th element in the abundance map, m is the number of endmembers in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

[0039] Alternatively, the sparse constrained least squares model is:

[0040]

[0041] Where I is the multispectral tensor constructed based on the multispectral data, R2 is the second endmember matrix constructed based on the symbolic endmember spectral data, and Β={b k}, b k is the abundance map of the multispectral data corresponding to the kth symbol end member, is the (i, j)th element in the abundance map, n is the number of symbolic endmembers, and λ is a parameter used to control sparsity.

[0042] The ancient book restoration and management system containing mathematical symbols provided by the present invention has the following beneficial effects:

[0043] (1) The multi-spectral data is collected collaboratively using ultraviolet / visible light / near-infrared multi-bands to obtain the spectral characteristic differences of different materials (ink, paper, interference), and the detection sensitivity of faded text or symbols is significantly improved.

[0044] (2) Set up an end-member library module, use the linear unmixing method based on the end-member spectral data to identify ancient book images, effectively identify mathematical symbols, and combine with artificial intelligence to achieve effective recognition of text and symbols in ancient book images.

[0045] (3) The historical symbol templates were further combined with the text symbol recognition results to digitally restore the ancient book images, thereby improving the restoration accuracy.

[0046] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a schematic diagram of the structure of an ancient book restoration management system containing mathematical symbols provided by an embodiment of the present invention;

[0048] Figure 2 The present invention provides a flowchart of a method for repairing and restoring ancient books containing mathematical symbols. DETAILED DESCRIPTION

[0049] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0050] In order to improve the accuracy of digital recognition of ancient mathematical books and improve the quality of restoration of ancient mathematical books, the embodiment of the present invention provides an ancient book restoration management system containing mathematical symbols. Figure 1 As shown, the system includes: a multispectral data acquisition module, an end member library module, a recognition module, a template management module, an image restoration module and an ancient book management module.

[0051] The multispectral data acquisition module is used to collect images of ancient books in the ultraviolet band, visible light band and near-infrared band to obtain multispectral data.

[0052] Specifically, images of ancient books are captured using cameras in the ultraviolet, visible, and near-infrared bands, generating images in three different bands to form multispectral data. When capturing images in the ultraviolet band, visible light interference must be disabled; when capturing images in the visible light band, D65 white light is preferred; and when capturing images in the near-infrared band, an infrared-transmitting filter can be added to the visible light camera. Furthermore, the images captured by the cameras can be aligned to the same resolution, ensuring that the images in different bands in the resulting multispectral data have the same resolution.

[0053] The endmember library module is used to store endmember spectral data; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data.

[0054] In the present invention, end members refer to the spectral characteristics of pure substances or pure components that constitute mixed pixels in an image, which can be compared to the "basic primary colors" on a palette. It is understandable that for images of ancient mathematical books, their spectral composition may contain one or more characteristics of ink, paper, and interference (mildew, stains, creases, etc.). Conventional OCR (optical character recognition) tools may mistakenly identify interference as ink, thereby leading to recognition errors. For this reason, the present invention introduces interference end members to effectively distinguish interference from ink in ancient book images. In addition, the particularity of the stroke structure of symbols in images of ancient mathematical books also makes them difficult to recognize by conventional OCR tools. For example, small symbols such as "·" and "∴" may be recognized as noise by conventional OCR recognition tools, and some class symbols "⊥" may be mistakenly identified as text by conventional OCR recognition tools. For this reason, the present invention introduces symbol end members to effectively distinguish symbols from text.

[0055] Regarding the acquisition of end-member spectral data, we can select clear ink pixels in ancient books in advance to obtain their average spectrum as ink end-member spectral data, select the positions of blank spaces in ancient books to obtain their average spectrum as paper end-member spectral data, select the positions of mildew spots in ancient books to obtain their average spectrum as mildew end-member spectral data, select the positions of stains in ancient books to obtain their average spectrum as stain end-member spectral data, and select the average spectrum of the positions of different types of symbols in ancient books as end-member spectral data of the corresponding types of symbols in ink.

[0056] The recognition module is used to perform text and symbol recognition on ancient book images based on multispectral data and endmember spectral data using artificial intelligence and linear unmixing methods to obtain text and symbol recognition results.

[0057] Specifically, the recognition module uses artificial intelligence and linear unmixing methods to perform text and symbol recognition on ancient book images based on multispectral data and endmember spectral data, and obtains text and symbol recognition results, including:

[0058] Step 1: Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve the ratio of pixels in the multispectral data corresponding to the ink, paper, and interference, to obtain a first solution result, and an ink recognition result is obtained based on the first solution result;

[0059] In step S10, the fully constrained least squares model may be:

[0060]

[0061] In formula (1), I is a multispectral tensor constructed based on multispectral data, and its dimension is C×L×H, where C is the number of bands. In the present invention, C=3, and L×H represents the size of the spectral image of a single band. R1 is a first endmember matrix constructed based on the ink endmember spectral data, paper endmember spectral data, and interference endmember spectral data. Its dimension is m×C, where m is the number of endmembers contained in the first endmember matrix. For example, if the endmembers include only ink, paper, and mildew, then m=3. A={a k}, whose dimensions are m×L×H, a k is the abundance map of the multispectral data corresponding to the kth end member (size is L×H), is the (i, j)th element in the abundance map, β is a parameter for controlling the smoothing intensity, preferably 0.05, but of course not limited to this.

[0062] In this fully constrained least squares model, a spatial continuity constraint is introduced on the basis of the conventional fully constrained least squares model to address the possible stroke breaks in ancient book images. That is, a spatial smoothing term (the second term in Equation (1)) is introduced into the objective function of the conventional fully constrained least squares model to force the abundance of adjacent pixels to be similar, thereby effectively identifying stroke breaks.

[0063] After solving A based on formula (1), which is the first solution, according to the The specific value of can be used to determine whether the (i, j)th element corresponds to ink, paper or interference, thereby obtaining the ink recognition result.

[0064] Step 2: Perform image segmentation on the multispectral data according to the ink recognition result to obtain segmented multispectral data, and use artificial intelligence to recognize text in the segmented multispectral data to obtain text recognition results.

[0065] Here, the multispectral data is segmented based on the ink recognition results. This means separating the ink area from the background area. This operation can remove interference and the impact of yellowing paper on image quality. In other words, mold, stains, and yellowed paper can be transformed into a clean background through this step. Then, artificial intelligence is used to further recognize text in the segmented multispectral data. The artificial intelligence here can use conventional OCR tools, but of course it is not limited to this.

[0066] Furthermore, since conventional OCR tools may mistakenly identify some symbols as text, in practice, we can analyze the specific OCR data and discard the recognized text with low confidence, retaining only the recognized text with high confidence. Furthermore, it is understood that symbols that might be identified as noise by conventional OCR tools will not be recognized as text by OCR tools.

[0067] Step 3: Based on the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve the ratio of pixels in the segmented multispectral data corresponding to symbols or non-symbols, to obtain a second solution result, and a symbol recognition result is obtained based on the second solution result.

[0068] Specifically, after the operation of step S20, the text in the ink has been recognized, so the remaining ink image that has not been recognized as text can be further segmented from the segmented multispectral data to form further segmented multispectral data, and then the sparse constrained least squares model is used to solve the problem based on the further segmented multispectral data and the symbol end member spectral data.

[0069] The sparse constrained least squares model is:

[0070]

[0071] Where I is the multispectral tensor constructed based on the further segmented multispectral data, and its dimension is also C×L×H. R2 is the second endmember matrix constructed based on the symbol endmember spectral data, and its dimension is (n+1)×C, where n is the number of symbol endmembers, and B={b k}, b k is the abundance map of the multispectral data corresponding to the kth symbol end member, is the (i, j)th element in the abundance map. When k = 0, Indicates b k The (i, j)th element in does not correspond to any symbol, and λ is a parameter for controlling sparsity, which is preferably 0.1, but is not limited thereto.

[0072] After solving Β based on formula (2), which is the second solution, according to the The specific value of can be used to determine whether the (i, j)th element corresponds to a symbol end member and what symbol end member it corresponds to, thereby obtaining a preliminary symbol recognition result, that is, preliminarily determining which positions of the ink remaining after removing the text in the ancient book image correspond to symbols, and what kind of stroke structures these symbols correspond to (determined by the corresponding symbol end members).

[0073] Template management module, used to store historical symbol templates.

[0074] The template management module is actually a historical symbol template library that stores a large number of historical symbol templates. These historical symbol templates are obtained in the following way: first, high-precision scanned images of ancient mathematical texts are obtained. After the images are enhanced using image enhancement methods, symbol images are extracted from them by experts in ancient mathematical text research. Based on the extracted symbol images, the symbol images are then converted into vectors using Illustrator (Adobe Illustrator) or Inkscape, and the converted symbol vectors are used to represent the geometric characteristics of the symbols. Among them, Inkscape is a free and open-source vector graphics editor.

[0075] The image restoration module is used to digitally restore ancient book images based on text symbol recognition results and historical symbol templates.

[0076] Specifically, for the symbol initially identified in the text symbol recognition results (assuming it is represented by Sig), Illustrator (Adobe Illustrator) or Inkscape is also used to vectorize S. Then, based on the symbol end member corresponding to Sig, a historical symbol template containing the symbol end member is selected from the historical symbol templates stored in the template management module, and the vectorized representation of Sig is matched with the selected historical symbol template (based on vector similarity calculation), thereby finally determining what specific symbol Sig is. In this way, the symbols and text in the ancient book image have been recognized, and a digital restored image can be generated for the ancient book image based on the final recognition result at this time, achieving a similar restoration effect.

[0077] The ancient book management module is used to manage digital ancient books based on the digitally restored ancient book images.

[0078] Specifically, the management of digital ancient books based on the digitally restored ancient book images can include: combining the digitally restored ancient book images into volumes, adding page numbers, abstracts, covers and other information; classifying and storing ancient mathematics books, adding relevant retrieval and browsing functions, etc. The specific management method can be formulated according to the use requirements of ancient mathematics books, and the present invention will not elaborate on it.

[0079] The ancient book restoration and management system containing mathematical symbols provided by the present invention uses ultraviolet / visible light / near-infrared multi-band collaborative collection of multispectral data to obtain the spectral characteristic differences of different materials (ink, paper, interference), and significantly improves the detection sensitivity of faded text or symbols. The present invention provides an end-member library module, uses a linear unmixing method based on end-member spectral data to identify ancient book images, effectively identifies mathematical symbols, and combines artificial intelligence to achieve effective recognition of text and symbols in ancient book images. In addition, the present invention further combines historical symbol templates with text and symbol recognition results to digitally restore ancient book images, improving the restoration accuracy rate.

[0080] Based on the same inventive concept, the present invention also provides a method for repairing and restoring ancient books containing mathematical symbols, see Figure 2 , the method comprises the following steps:

[0081] S10, obtaining multispectral data; the multispectral data is obtained by collecting images of ancient books in ultraviolet bands, visible light bands, and near-infrared bands;

[0082] S20, acquiring endmember spectral data and historical symbol templates; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data, and symbol endmember spectral data;

[0083] S30, performing text and symbol recognition on the ancient book image using artificial intelligence and linear unmixing method based on the multispectral data and the endmember spectral data to obtain a text and symbol recognition result;

[0084] S40. Digitally repair and restore the ancient book image based on the text symbol recognition results and the historical symbol template.

[0085] Optionally, the symbolic endmember spectral data includes: straight endmember spectral data, circular endmember spectral data, angle endmember spectral data and point endmember spectral data.

[0086] Optionally, the ink endmember spectral data includes endmember spectral data of inks of multiple different materials. Optionally, the ink endmember spectral data includes: endmember spectral data of ink without penetration, endmember spectral data of ink with slight penetration, endmember spectral data of ink with moderate penetration, and endmember spectral data of ink with severe penetration.

[0087] Optionally, the recognition module uses artificial intelligence and linear unmixing methods to perform text and symbol recognition on the ancient book image based on the multispectral data and the endmember spectral data, and obtains text and symbol recognition results, including:

[0088] Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve the ratio of pixels in the multispectral data corresponding to the ink, paper, and interference to obtain a first solution result, and an ink recognition result is obtained based on the first solution result;

[0089] Perform image segmentation on the multispectral data according to the ink recognition result to obtain segmented multispectral data, and use artificial intelligence to recognize text in the segmented multispectral data to obtain text recognition results;

[0090] According to the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve the ratio of pixels in the segmented multispectral data corresponding to symbols or non-symbols to obtain a second solution result, and a symbol recognition result is obtained according to the second solution result.

[0091] Alternatively, the fully constrained least squares model is:

[0092]

[0093] Where I is a multispectral tensor constructed based on multispectral data, R1 is a first endmember matrix constructed based on ink endmember spectral data, paper endmember spectral data and interference endmember spectral data, A={a k}, a k is the abundance map of the multispectral data corresponding to the k-th end member, is the (i, j)th element in the abundance map, m is the number of endmembers in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

[0094] Alternatively, the sparse constrained least squares model is:

[0095]

[0096] Where I is the multispectral tensor constructed based on the multispectral data, R2 is the second endmember matrix constructed based on the symbolic endmember spectral data, and Β={b k}, b k is the abundance map of the multispectral data corresponding to the kth symbol end member, is the (i, j)th element in the abundance map, n is the number of symbolic endmembers, and λ is a parameter used to control sparsity.

[0097] It should be noted that, for the method embodiment, since it is basically similar to the system embodiment, the description is relatively simple. For relevant details, please refer to the partial description of the system embodiment. The method embodiment can achieve the same or similar beneficial effects as the system embodiment.

[0098] It should be noted that the terms "first," "second," and the like are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention.

[0099] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0100] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the word "comprising" does not exclude other components or steps, "one" or "a" does not exclude multiple situations, and "multiple" means two or more, unless otherwise clearly and specifically defined. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0101] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, devices (equipment), or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or a combination of software and hardware embodiments, all of which are collectively referred to herein as "modules" or "systems." Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The computer program may be stored / distributed in a suitable medium, provided together with other hardware or as part of the hardware, or may be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0102] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, apparatus (devices) and computer program products of the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0103] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0105] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. An ancient book restoration management system containing mathematical symbols, characterized in that: include: The multispectral data acquisition module is used to collect images of ancient books in the ultraviolet band, visible light band and near-infrared band to obtain multispectral data; An endmember library module is used to store endmember spectral data; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data; a recognition module, configured to perform text and symbol recognition on the ancient book image using artificial intelligence and a linear unmixing method based on the multispectral data and the endmember spectral data, to obtain a text and symbol recognition result; Template management module, used to store historical symbol templates; An image restoration module, configured to digitally restore the ancient book image based on the text symbol recognition result and the historical symbol template; The ancient book management module is used to manage digital ancient books based on the digitally restored ancient book images.

2. The ancient book restoration and management system containing mathematical symbols according to claim 1 is characterized in that: The symbolic end-member spectral data includes: straight end-member spectral data, circular end-member spectral data, angle end-member spectral data and point end-member spectral data.

3. The ancient book restoration and management system containing mathematical symbols according to claim 1 is characterized in that: The ink end-member spectrum data includes end-member spectrum data of inks of various materials.

4. The ancient book restoration and management system containing mathematical symbols according to claim 1 is characterized in that: The ink end-member spectrum data includes: no-penetration ink end-member spectrum data, light-penetration ink end-member spectrum data, moderate-penetration ink end-member spectrum data, and heavy-penetration ink end-member spectrum data.

5. The ancient book restoration and management system containing mathematical symbols according to claim 1 is characterized in that: The recognition module performs text and symbol recognition on the ancient book image based on the multispectral data and the endmember spectral data using artificial intelligence and a linear unmixing method to obtain a text and symbol recognition result, including: Based on the multispectral data, the ink endmember spectral data, the paper endmember spectral data, and the interference endmember spectral data, a fully constrained least squares model is used to solve the ratio of pixels in the multispectral data corresponding to ink, paper, and interference to obtain a first solution result, and an ink recognition result is obtained based on the first solution result; Performing image segmentation on the multispectral data according to the ink recognition result to obtain segmented multispectral data, and using artificial intelligence to recognize text in the segmented multispectral data to obtain a text recognition result; According to the segmented multispectral data and the symbol endmember spectral data, a sparse constrained least squares model is used to solve the ratio of pixels in the segmented multispectral data corresponding to symbols or non-symbols to obtain a second solution result, and a symbol recognition result is obtained according to the second solution result.

6. The ancient book restoration and management system containing mathematical symbols according to claim 5 is characterized in that: The fully constrained least squares model is: Wherein, I is a multispectral tensor constructed based on the multispectral data, R1 is a first endmember matrix constructed based on the ink endmember spectral data, the paper endmember spectral data and the interference endmember spectral data, A={a k }, a k is the abundance map of the multispectral data corresponding to the k-th end member, is the (i, j)th element in the abundance map, m is the number of endmembers contained in the first endmember matrix, and β is a parameter used to control the smoothing intensity.

7. The ancient book restoration and management system containing mathematical symbols according to claim 5 is characterized in that: The sparse constrained least squares model is: Wherein, I is a multispectral tensor constructed based on the multispectral data, R2 is a second endmember matrix constructed based on the symbolic endmember spectral data, and B={b k }, b k is the abundance map of the multispectral data corresponding to the kth symbol end member, is the (i, j)th element in the abundance map, n is the number of symbolic endmembers, and λ is a parameter used to control sparsity.

8. A method for restoring ancient books containing mathematical symbols, characterized in that: include: Acquiring multispectral data; the multispectral data is obtained by collecting images of ancient books in ultraviolet bands, visible light bands, and near-infrared bands; Acquiring endmember spectral data and historical symbol templates; the endmember spectral data includes ink endmember spectral data, paper endmember spectral data, interference endmember spectral data and symbol endmember spectral data; According to the multispectral data and the endmember spectral data, using artificial intelligence and linear unmixing method to perform text and symbol recognition on the ancient book image to obtain a text and symbol recognition result; The ancient book image is digitally restored based on the text symbol recognition result and the historical symbol template.

9. The method for repairing and restoring ancient books containing mathematical symbols according to claim 8, characterized in that: The symbolic end-member spectral data includes: straight end-member spectral data, circular end-member spectral data, angle end-member spectral data and point end-member spectral data.

10. The method for repairing and restoring ancient books containing mathematical symbols according to claim 8, characterized in that: The ink end-member spectrum data includes end-member spectrum data of inks of various materials.

Citation Information

Patent Citations

  • Ancient tomb inscription character recognition method based on hyperspectral remote sensing technology

    CN111539409A

  • Mathematical ancient book illustration identification method and device, electronic equipment and storage medium

    CN119649375A

  • Unmixing image data

    US20240273678A1