Archive digitization management method and system and storage medium
By performing brightness compensation and feature extraction on the original scanned image data of the handwritten archive, combined with local projection analysis and term matching processing, the problem of handwritten report recognition difficulties in traditional technology is solved, and efficient and accurate digital processing and information extraction are achieved.
Patent Information
- Application Number
- CN202510090055.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional archive digital management methods encounter recognition problems when processing handwritten reports, especially the researchers' notes and handwriting are sloppy, and the accuracy of recognition using OCR technology is difficult to meet, resulting in the need of manual proofreading, affecting the digital efficiency.
By obtaining the original scanned image data of the technological handwritten archive, regional brightness compensation processing is performed to extract handwriting feature data, including handwriting pressure change feature data and ink diffusion feature data. Then, local projection analysis and breakpoint connectivity analysis are performed based on these feature data to generate standardized character sequence data. The preset research term verification vocabulary is used to perform term matching and intelligent error correction processing, and finally structured field recognition and index construction of text data.
It significantly improves the image quality and character recognition accuracy of handwritten documents, reduces the need for manual proofreading, improves the efficiency and quality of digital processing, and ensures the accuracy and reliability of research information.
Smart Images

Figure CN120014650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information digitization technology, and in particular to a management method, system and storage medium for digitalized archives. Background Art
[0002] Archive digitization refers to the process of converting archival information on traditional carriers such as paper, audio, and video into digital information through scanning, filming, and other technical means, and storing, managing, and utilizing it in a computer system. This is like moving a traditional library into a computer, but it is smarter and more convenient than a traditional library. Before starting digitization, comprehensive research and planning are required. This includes assessing the number, type, and preservation status of existing archives, and determining the scope of archives that should be digitized first. At the same time, detailed work processes and quality standards should be formulated, just like having a complete design drawing before building a building. Before the actual scanning, the archives need to be sorted and repaired. This includes removing staples, flattening wrinkles, and repairing damage. This stage is particularly important because the physical condition of the archives directly affects the quality of digitization. Digital conversion is the core link, which mainly includes scanning, image optimization, and format conversion, and finally data management and storage.
[0003] However, traditional methods of managing digital archives often have the following problems: Research institutions often encounter recognition difficulties when dealing with handwritten reports. Researchers' notes are often illegible and use a lot of professional abbreviations and symbols. Even with the most advanced OCR technology, the recognition accuracy is difficult to be satisfactory. Studies have shown that the recognition accuracy of handwritten reports is usually only 75%-85%, which means that 15-25 pages out of every 100 pages of documents need to be manually proofread, which greatly affects the efficiency of digitization. Summary of the invention
[0004] Based on this, it is necessary for the present invention to provide a digital archive management method, system and storage medium to solve at least one of the above technical problems.
[0005] To achieve the above purpose, a digital archive management method includes the following steps:
[0006] Step S1: obtaining original scanned image data of scientific and technological handwritten files; performing regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data;
[0007] Step S2: performing local projection analysis on the original scanned image data according to the handwriting feature data, and performing breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data;
[0008] Step S3: using a preset research term verification word library to perform term matching processing on the standardized character sequence data, and performing intelligent error correction processing, thereby obtaining standardized research text data;
[0009] Step S4: Perform structured field recognition processing on the standardized research text data to obtain field type data; locate the text in the scientific and technological handwritten archives based on the field type data and the original scanned image data, and construct an index to generate R&D record index data.
[0010] The present invention can effectively improve the image quality and make the handwriting features clearer and more prominent by performing regional brightness compensation processing on the original scanned image data. The extraction of handwriting pressure change feature data and ink diffusion feature data provides a rich and accurate basis for subsequent character analysis, helps to more accurately identify and understand the morphology and writing habits of handwritten characters, lays a solid foundation for the entire digitization process, and greatly improves the accuracy and reliability of character recognition. Based on the local projection analysis and breakpoint connectivity analysis of handwriting feature data, characters can be accurately standardized. The acquisition of character inclination angle data and stroke connection feature data allows characters that may have different shapes due to non-standard writing to be unified and standardized to form standardized character sequence data. This not only facilitates subsequent term matching and intelligent error correction, but also greatly improves the consistency and readability of text data, provides a more standardized and orderly data basis for subsequent text processing and information extraction, and effectively reduces recognition errors and information deviations caused by differences in character morphology. Using the preset research term verification vocabulary for term matching processing, combined with the intelligent error correction function, professional term errors in handwritten documents can be accurately identified and corrected. The generation of standardized research text data ensures the professionalism and accuracy of the text content, which is extremely important for the transmission, sharing and subsequent decision support of research information. It can effectively avoid misunderstandings of research information caused by handwriting errors or ambiguity, improve the overall quality and credibility of research documents, and provide high-quality data guarantee for digital management and data analysis in the research industry. The standardized research text data is processed by structured field recognition, and the text positioning and index construction are combined with the original scanned image data to achieve efficient generation of R&D record index data. The acquisition of field type data enables the research text information to be accurately classified and stored according to the preset structure, which is convenient for users to quickly query and retrieve the required information. At the same time, the text positioning and index construction functions allow users to quickly find the location of specific fields and content in the document, greatly improving the operability and practicality of scientific and technological R&D records, providing great convenience for researchers to quickly obtain subject information, conduct report review and analysis in their work, and effectively improving the efficiency and quality of research work.
[0011] Preferably, the present invention further provides a digital archive management system for executing the digital archive management method described above, wherein the digital archive management system comprises:
[0012] The image preprocessing module is used to obtain the original scanned image data of the scientific and technological handwritten archives; perform regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data;
[0013] A character standardization module is used to perform local projection analysis on the original scanned image data according to the handwriting feature data, and to perform breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data;
[0014] The term processing module is used to perform term matching processing on the standardized character sequence data using the preset research term verification word library, and perform intelligent error correction processing to obtain standardized research text data;
[0015] The data structuring and indexing module is used to perform structured field recognition processing on standardized research text data to obtain field type data; locate the text in scientific and technological handwritten archives based on the field type data and the original scanned image data, and construct an index to generate R&D record index data.
[0016] The present invention also provides a storage medium storing a computer program, wherein the computer program is used to execute the above-mentioned method for managing digital archives.
[0017] The present invention can effectively improve the image quality, and significantly improve the problems of uneven brightness and blurred handwriting that may be caused by scanning conditions or writing habits, providing a high-quality image foundation for subsequent character recognition and analysis, and greatly improving the accuracy and reliability of the entire system. The character standardization module then plays a role. It performs local projection analysis and breakpoint connectivity analysis on the original scanned image data based on the handwriting feature data, and then obtains standardized character sequence data, which includes character inclination angle data and stroke connection feature data. The benefit of this module is that it can standardize handwritten characters of various shapes, making the shape and arrangement of characters more regular and unified, which is convenient for subsequent term matching and intelligent error correction, and also improves the consistency and readability of text data, laying a solid foundation for the accurate transmission and effective use of research information. The key value of the term processing module is to ensure the professionalism and accuracy of the text content. Through accurate term matching and effective intelligent error correction, it can correct professional term errors or writing errors that may appear in handwritten documents, avoid research decision errors caused by information errors, and greatly improve the quality and credibility of research documents. The outstanding advantage of the data structuring and indexing module is that it realizes the efficient organization and rapid retrieval of research text data. By structuring and accurately indexing text data, researchers can quickly find the required information, improve the efficiency of research work, and facilitate the statistical analysis and sharing of research data, providing strong support for research information construction and big data applications. In summary, the cooperation of these four modules not only improves the quality and efficiency of the digital processing of scientific and technological handwritten archives, but also ensures the accuracy, consistency and availability of research information, which is of great significance for promoting the digital transformation of the research industry and improving the level of research services. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments thereof made with reference to the following drawings:
[0019] Figure 1 A schematic diagram of the steps of the digital archive management method of the present invention;
[0020] Figure 2 for Figure 1 Detailed step flow diagram of step S1;
[0021] Figure 3 for Figure 1 Detailed step flow chart of step S2 in FIG. DETAILED DESCRIPTION
[0022] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0023] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0024] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0025] To achieve this, please refer to Figures 1 to 3 The present invention provides a method for managing digital archives, the method comprising the following steps:
[0026] Step S1: obtaining original scanned image data of scientific and technological handwritten files; performing regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data;
[0027] The embodiment of the present invention obtains the original scanned image data of the scientific and technological handwritten archive. The original scanned image data usually comes from a high-resolution scanner, and the resolution can be set to more than 300 DPI to ensure that the details of the handwriting are accurately captured. Then, the image is subjected to regional brightness compensation processing to solve the problem of blurred ink or insufficient image contrast that may appear in the handwritten archive. Brightness compensation processing is usually implemented by local image enhancement technology, using adaptive histogram equalization (AHE) or local contrast enhancement (LCE) algorithm to adjust each area in the image. The processed image can clearly show the details of the handwriting, including the pressure change of the strokes and the diffusion characteristics of the ink. In the process of extracting the handwriting feature data, the pressure change characteristics of the strokes, that is, the thickness change of the handwriting, are first calculated using the gray value change of the image, and then the ink diffusion characteristics are obtained to describe the diffusion range and direction of the ink. In actual application scenarios, for example, when processing a handwritten experimental report, through this brightness compensation processing, the unclear handwriting area caused by the scanning quality problem can be effectively restored, making the subsequent processing more accurate.
[0028] Step S2: performing local projection analysis on the original scanned image data according to the handwriting feature data, and performing breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data;
[0029] According to the handwriting feature data obtained in step S1, the embodiment of the present invention performs local projection analysis and breakpoint connectivity analysis on the original scanned image. Local projection analysis is to project on the horizontal and vertical coordinates of the image respectively, and count the brightness changes of each row or column of pixels, so as to identify the distribution of handwriting, especially the stroke outline of the character. In specific operation, by setting an appropriate window size (for example, 5px x 5px), each small area in the image is scanned, and the brightness value is summed and the change detection is performed to locate the boundary of the character. Next, through the breakpoint connectivity analysis, it is identified whether there is a break in the handwriting, such as whether there is a clear gap between the strokes. If there is a breakpoint, an interpolation algorithm (such as bilinear interpolation or Bezier curve interpolation) is used to supplement the connection and restore the broken strokes. In this way, the system can extract the tilt angle data of the character and analyze the rotation angle of the character in order to construct the standardized character sequence data; at the same time, the stroke connection feature data can also be obtained to determine the connection relationship between the strokes, such as whether the upper and lower strokes of the letter "o" are coherent. In actual application scenarios, this method is often used to identify slanted text and incompletely connected glyphs in handwritten documents, especially in the handwritten parts of scientific research reports, to ensure that each character is processed correctly.
[0030] Step S3: using a preset research term verification word library to perform term matching processing on the standardized character sequence data, and performing intelligent error correction processing, thereby obtaining standardized research text data;
[0031] The embodiment of the present invention uses a preset research term verification word library to perform term matching processing on the standardized character sequence data, and performs intelligent error correction processing. First, the character level sequence annotation model (such as CRF model or BERT model) in the natural language processing technology is used to match the characters in the standardized character sequence data to find terms related to the research field. For example, in medical or chemical literature, the word library may contain terms such as "cell division", "chemical reaction", etc. Through word library comparison, the system can identify the matching terms in the standardized character sequence. If a suspected error (such as a spelling error or character confusion) is found, an intelligent error correction algorithm (such as an algorithm based on spelling distance Levenshtein distance, Word2Vec model, etc.) is used for error correction processing. The intelligent error correction process automatically corrects errors such as "molecular substructure" being misidentified as "molecular group structure" by learning a large number of correct terms and their variants in the field, and generates standardized research text data. In applications, such as when processing handwritten medical reports, through this intelligent error correction processing, it can be ensured that the medical terms in the handwritten text are automatically corrected to ensure the accuracy and consistency of the document content.
[0032] Step S4: Perform structured field recognition processing on the standardized research text data to obtain field type data; locate the text in the scientific and technological handwritten archives based on the field type data and the original scanned image data, and construct an index to generate R&D record index data.
[0033] The embodiment of the present invention performs structured field recognition processing on the standardized research text data to obtain field type data. First, by parsing the standardized research text data, each structured field in the document, such as title, author, abstract, experimental method, research results, etc., is identified. This process usually relies on a machine learning classification model (such as SVM, decision tree or deep learning model) to train the text and identify the types of each field in the document. In the application, for a scientific research paper, the model can accurately identify the "method" field, the "result" field, etc. Then, the text is positioned according to the field type data and the original scanned image data, and the position of each field in the document is calibrated. At this time, it is necessary to spatially locate each character or word in the scanned image through OCR technology and map it to the pixel coordinates in the image. For example, the title "research method" may be located at the top of the page, while the text "experimental steps" may be located in the middle of the page. Based on these position information, the system can further perform index construction. By constructing position index data, subsequent keyword retrieval and document analysis can be effectively supported to ensure that users can quickly locate key content in the document. In practical applications, researchers can use this method to quickly locate specific content about "experimental methods" or "result analysis" in scientific research literature and accurately display the relevant locations in the image.
[0034] Preferably, step S1 comprises the following steps:
[0035] Step S11: obtaining physical scanning parameter data of the scientific and technological handwritten archive; adjusting the resolution and contrast of the scanning device according to the physical scanning parameter data, thereby obtaining scanning configuration data;
[0036] The embodiment of the present invention obtains physical scanning parameter data of scientific and technological handwritten archives. Specifically, the physical scanning parameters of the document are collected by a scanning device (such as a high-resolution scanner), including basic information such as scanning resolution, scanning contrast, and scanning brightness. These parameters can be directly read through the configuration interface of the scanner, or obtained by analyzing the scanned document image. In actual operation, the resolution parameter of the scanner is usually set to at least 300dpi to ensure that the details of the image are clearly visible. According to these physical scanning parameter data, the scanning device is further adjusted to optimize the image quality, especially under different scanning conditions, the resolution and contrast may need to be adjusted according to the specific characteristics of the document (such as text density, paper material). For example, the resolution (such as from 300dpi to 600dpi) and the contrast (such as increasing the contrast to 50) can be adjusted through the setting interface of the scanner or the driver to adapt to the scanning requirements of different documents. Finally, the scanning configuration data is obtained according to these adjustments.
[0037] Step S12: performing multi-region scanning processing on the scientific handwritten file according to the scanning configuration data, thereby obtaining original scanned image data;
[0038] The embodiment of the present invention performs multi-region scanning processing according to the scanning configuration data obtained in step S11. Specifically, the scientific and technological handwritten archives are first divided into multiple areas, including a text area, a title area, an image area, and a background area. In multi-region scanning, different scanning strategies are applied to each area by adjusting the settings of the scanning device. For example, for areas with dense text, a higher resolution scan (such as 600dpi) is used, while for images or blank areas, a lower resolution (such as 300dpi) can be used to save storage space and processing time. In actual operation, the scanning device can assign scanning tasks to different areas by means of automatic area recognition or manual division of scanning areas. After the processing is completed, the original scanned image data contains the scanning information of these different areas, ensuring the comprehensiveness and detail clarity of the image.
[0039] Step S13: performing image quality evaluation based on clarity and noise level on the original scanned image data, thereby obtaining image quality feature data;
[0040] The embodiment of the present invention evaluates the clarity of the scanned image through image processing algorithms (such as FFT transformation, edge detection, etc.). The clarity evaluation is mainly carried out by calculating the ratio of the high-frequency component (edge information) to the low-frequency component of the image, and using, for example, a gradient operator (Sobel operator) to extract the edge information of the image, thereby obtaining the clarity score of the image. In terms of noise level evaluation, a noise removal algorithm (such as median filtering or Gaussian filtering) is used to process the noise in the image, and the noise ratio is calculated, and the impact of the noise is evaluated based on the distribution characteristics of the noise (for example, based on the peak signal-to-noise ratio PSNR). Through these image quality evaluations, image quality feature data is obtained, including the clarity score and noise level score of the image. These data will provide an important basis for subsequent image processing.
[0041] Step S14: performing adaptive region division processing on the original scanned image data according to the clarity information in the image quality feature data, thereby obtaining image block data, wherein the image block data includes core text region data and edge background region data;
[0042] The embodiment of the present invention performs adaptive region division processing on the original scanned image data based on the clarity information in the image quality feature data obtained in step S13. Specifically, a clarity threshold is used to judge the quality of different areas in the image. For example, a clarity threshold is set. When the clarity score of the image is higher than the threshold, the area is considered to be a core text area; when the clarity is lower, the area is divided into a background area. In order to achieve adaptive region division, the image is firstly processed into blocks, and the image is divided into multiple small areas, and the size of each area is dynamically adjusted according to the clarity information. For areas with higher clarity, the block size is smaller to facilitate accurate analysis of text details; for areas with lower clarity, the block size is larger to avoid noise influence. Finally, the image is divided into a core text area and an edge background area, and image block data is obtained.
[0043] Step S15: calculating the local brightness mean of each area according to the image block data, thereby obtaining background brightness distribution data; constructing a brightness compensation curve using the background brightness distribution data, thereby obtaining brightness compensation parameter data;
[0044] The embodiment of the present invention uses a grayscale histogram analysis method to calculate the brightness mean value in each block area. The specific steps are: first, grayscale the block areas of the image, convert the image into a grayscale image, and then calculate the brightness mean value of each block area. For the background area, its brightness mean value is calculated and statistically analyzed to obtain the background brightness distribution data of the entire image. In practical applications, by comparing the brightness distribution of each block area of the image, the areas with uneven brightness that may exist are identified and marked. These background brightness distribution data will be used for subsequent brightness compensation calculations to ensure that the visual effect of the image is more uniform.
[0045] Step S16: Adaptively adjust the brightness of the core text area data in the image block data according to the brightness compensation parameter data, thereby obtaining the handwriting pressure change characteristic data; perform edge detection on the core text area data after brightness adjustment, thereby obtaining the ink diffusion characteristic data;
[0046] The construction of the brightness compensation curve in the embodiment of the present invention is based on the difference between the local brightness mean of each block area and the background brightness. By analyzing the brightness distribution of the background area in the image, a brightness compensation parameter can be obtained, and a linear or nonlinear fitting method (such as polynomial regression) is usually used to construct the compensation curve. The purpose of this curve is to adjust the brightness of the darker or brighter area in the image so that it reaches an ideal equilibrium state. According to the constructed compensation curve, the brightness of the core text area is adaptively adjusted. The adjusted core text area will show a more uniform brightness distribution, which helps to improve the readability of the handwriting. In this process, the functions in the image processing tool (such as OpenCV) can be used to adjust the brightness of the image.
[0047] Step S17: performing feature fusion processing on the original scanned image data according to the handwriting pressure change feature data and the ink diffusion feature data, so as to obtain the handwriting feature data.
[0048] The embodiment of the present invention performs feature fusion processing on the original scanned image data according to the handwriting pressure change feature data and ink diffusion feature data obtained in step S16. Specifically, the handwriting pressure change feature data is fused by the pressure change in the image (obtained through indicators such as grayscale difference and edge blur) and the ink diffusion feature data (ink edge information obtained by edge detection methods such as Canny operator extraction). These two types of feature data combine the details of the handwriting and the ink diffusion, which helps to accurately judge the quality and extended features of the handwriting. Through feature fusion algorithms, such as weighted average or principal component analysis (PCA), these two types of feature data are merged into a comprehensive handwriting feature data set. The fused handwriting feature data can fully reflect the handwriting details of the handwritten document and provide reliable input data for subsequent text recognition and document analysis.
[0049] The present invention can enable the scanning device to scan the document in an optimal state by acquiring the physical scanning parameter data of the scientific and technological handwritten archives and adjusting the resolution and contrast of the scanning device accordingly, thereby obtaining high-quality original scanned image data. This provides a good foundation for subsequent image processing and feature extraction, ensures that the details and layers of the image can be fully retained, and reduces image distortion and information loss caused by improper settings of the scanning device. Multi-region scanning and processing of the scientific and technological handwritten archives according to the scanning configuration data can obtain the image information of the document more comprehensively and meticulously. This regional scanning method can accurately scan according to the characteristics of different regions, further improve the integrity and accuracy of the image data, and provide a richer and more accurate data source for subsequent image quality evaluation and feature extraction. Image quality evaluation based on clarity and noise level of the original scanned image data can accurately understand the current quality status of the image. The acquisition of image quality feature data provides a clear guiding direction for subsequent image processing, so that subsequent processing can optimize the image more targeted, improve the overall quality of the image, and create more favorable conditions for feature extraction. Adaptive region division processing is performed based on the clarity information in the image quality feature data, and the image can be accurately divided into the core text area and the edge background area. This division enables subsequent processing to be performed according to the characteristics of different areas, thereby improving the efficiency and effect of the processing. For example, the core text area is processed intensively to extract more accurate handwriting features, while the edge background area is processed appropriately to reduce the interference of background noise on the handwriting feature extraction. The local brightness mean of each area is calculated based on the image block data and a brightness compensation curve is constructed, which can accurately compensate for the uneven brightness of the image. The generation of brightness compensation parameter data makes the brightness distribution of the image more uniform, and the handwriting in the core text area is more clearly visible, providing better image conditions for the extraction of handwriting features, and effectively improving the accuracy and reliability of handwriting feature extraction. Adaptive brightness adjustment processing and edge detection processing are performed on the core text area data in the image block data, which can accurately extract handwriting pressure change feature data and ink diffusion feature data. The acquisition of these two feature data provides an important basis for the comprehensive and accurate description of handwriting features, so that the subsequent feature fusion processing can more realistically reflect the writing characteristics and morphological characteristics of the handwriting, and provide richer information for subsequent character recognition and analysis. By fusing the handwriting pressure change feature data and the ink diffusion feature data, more comprehensive and accurate handwriting feature data can be obtained. This feature fusion method integrates handwriting information in multiple dimensions, making the handwriting features more complete and rich, and can better reflect the essential characteristics of handwritten characters. It provides high-quality feature input for subsequent character standardization processing, term matching, and intelligent error correction, effectively improving the performance and effect of the entire archive digital management method.
[0050] Preferably, step S15 comprises the following steps:
[0051] Step S151: Scan point by point in the core text area in the core text area data using a sliding window algorithm, extract the gray value of each pixel, and calculate the average gray value of the pixels in the window, thereby obtaining the local brightness mean data of the core area;
[0052] The embodiment of the present invention uses a sliding window algorithm to scan the core text area data point by point, extract the grayscale value of each pixel and calculate the average grayscale value of the pixels in the window. The specific operation is: first, set a sliding window of a fixed size, such as 5x5 pixels or 7x7 pixels, and select a suitable window size according to the scanning resolution of the document. The window starts from the upper left corner of the image and slides gradually to the right and downward, moving one pixel at a time. In each window, the grayscale values of all pixels in the window are extracted, and the average grayscale value of these pixels is calculated. For each sliding window position, a local brightness mean is calculated. In this way, a data set containing the local brightness mean values of each position in the core text area can be obtained, and these data can reflect the local brightness changes of the text area. In practical applications, this method can effectively detect areas with uneven grayscale in the image and provide basic data for subsequent brightness compensation.
[0053] Step S152: segmenting the edge background region data by a region growing algorithm, extracting the gray value distribution characteristics of each sub-region, and calculating the average gray value of the sub-region, thereby obtaining the local brightness mean data of the background region;
[0054] The embodiment of the present invention segments the edge background area data through a region growing algorithm, extracts the grayscale value distribution characteristics of each sub-region, and calculates the average grayscale value of the sub-region. Specifically, first starting from a certain pixel point in the background area, set the point as a seed point, and use the region growing algorithm to gradually expand the neighborhood area according to the grayscale similarity until the set stop condition is met (such as the grayscale value difference of the pixels in the region is greater than a certain threshold). During the region growing process, the grayscale mean of the expanded area is calculated for each expansion step. After the entire background area segmentation is completed, each sub-region will have an independent grayscale mean. Through this method, the complex background area can be divided into multiple relatively uniform sub-regions, and then the local brightness mean data of each sub-region is calculated, which helps to more accurately understand the brightness changes and distribution characteristics of the background area.
[0055] Step S153: Integrate the brightness information in the local brightness mean data of the core area and the local brightness mean data of the background area to establish an overall brightness distribution model, thereby obtaining background brightness distribution data;
[0056] The embodiment of the present invention integrates the brightness information in the local brightness mean data of the core area and the local brightness mean data of the background area to establish an overall brightness distribution model. First, the local brightness mean data of the core area obtained in step S151 and the local brightness mean data of the background area obtained in step S152 are collected. Then, these two parts of data are merged by a data fusion method (such as weighted average or interpolation method) to obtain a complete set of brightness distribution information. Specifically, the brightness mean data of the core area reflects the local brightness changes in the central area of the image, while the brightness mean data of the background area provides brightness change information in the edge area. By integrating the two, a complete overall brightness distribution model can be constructed, which can reflect the brightness change trend of the entire image, especially the brightness contrast of different areas during document scanning.
[0057] Step S154: performing mathematical modeling based on the background brightness distribution data, and determining curve parameters by the least square method, thereby generating a brightness compensation curve that can characterize the overall brightness change trend;
[0058] The embodiment of the present invention performs mathematical modeling based on the background brightness distribution data, and determines the parameters of the brightness compensation curve by the least square method. First, according to the background brightness distribution data obtained in step S153, a set of key data points are selected, and these data points represent the brightness information of different areas in the image. Then, assuming that the brightness change follows a certain curve trend (such as linear, quadratic curve, etc.), the least square method is used to fit these data points. The goal of the least square method is to determine the optimal curve parameters by minimizing the sum of squares of the errors between the observed data points and the fitted curve. For example, a quadratic function form can be used: y=ax 2 +bx+c, where y represents the brightness value, x is the position coordinate of the image, and a, b, and c are the parameters to be determined. The optimal parameter value is calculated by the least squares method, thereby obtaining a brightness compensation curve that characterizes the trend of brightness changes in the image.
[0059] Step S155: discretizing the brightness compensation curve into specific numerical parameters, and calculating corresponding brightness compensation coefficients for different positions of the image, thereby obtaining brightness compensation parameter data.
[0060] The embodiment of the present invention discretizes the brightness compensation curve obtained in step S154 into a set of numerical parameters, which correspond to the brightness compensation values at different positions of the image. For example, the compensation curve is sampled within a certain range, and the brightness compensation values of the sampling points are stored as discretized brightness compensation coefficients. These coefficients can be stored in the form of a table or an array, recording the brightness compensation value of each position. Then, in practical applications, for each pixel in the image, the corresponding brightness compensation coefficient is searched and applied according to its position in the image. These compensation coefficients will be adjusted according to the compensation curve to ensure the balance of image brightness and the readability of the document. In specific application scenarios, such as when processing research documents, the calculation and application of the brightness compensation coefficient can effectively reduce the brightness deviation in the document caused by uneven light, poor scanning quality, etc., thereby improving the recognizability of the text and the overall scanning quality.
[0061] The present invention utilizes a sliding window algorithm to scan point by point in the core text area in the core text area data, extracts the gray value of each pixel, and calculates the average gray value of the pixel in the window, thereby obtaining the local brightness mean data of the core area. This method can accurately analyze the brightness change of the core text area and provide detailed data support for subsequent brightness compensation. In this way, the part with uneven brightness in the image can be effectively identified, laying the foundation for further image processing. The edge background area data is segmented by a regional growing algorithm, the gray value distribution characteristics of each sub-area are extracted, and the average gray value of the sub-area is calculated, thereby obtaining the local brightness mean data of the background area. The regional growing algorithm can effectively segment the background area into multiple sub-areas and calculate their brightness mean values respectively, which helps to more accurately describe the brightness distribution of the background area. This segmentation and calculation method can improve the accuracy of background area brightness estimation and reduce the interference of background noise on text area brightness compensation. The brightness information in the local brightness mean data of the core area and the local brightness mean data of the background area is integrated to establish an overall brightness distribution model, thereby obtaining background brightness distribution data. By integrating the brightness information of the core area and the background area, a comprehensive brightness distribution model can be constructed, which can more accurately reflect the brightness change trend of the entire image. The establishment of this model provides a scientific basis for the generation of the brightness compensation curve and ensures the accuracy and effectiveness of the brightness compensation. Mathematical modeling is performed based on the background brightness distribution data, and the curve parameters are determined by the least squares method to generate a brightness compensation curve that can characterize the overall brightness change trend. The least squares method is a commonly used mathematical modeling method that can effectively fit data and determine the optimal curve parameters. The brightness compensation curve generated by this method can accurately describe the change trend of image brightness and provide an accurate mathematical tool for brightness compensation. The brightness compensation curve is discretized into specific numerical parameters, and the corresponding brightness compensation coefficients are calculated for different positions of the image to obtain brightness compensation parameter data. Discretization processing enables the brightness compensation curve to be converted into actual and operable numerical parameters, which is convenient for application in the image processing process. By calculating the brightness compensation coefficients at different positions, accurate brightness compensation can be achieved for each area in the image, improving the overall quality and readability of the image. In summary, these steps significantly improve the quality of scientific and technological handwritten archive images through precise brightness analysis and compensation, making subsequent character recognition and text processing more accurate and efficient. These methods not only improve the clarity and contrast of images, but also reduce the impact of noise and uneven brightness on image processing, providing strong technical support for the digital management of research images.
[0062] Preferably, step S2 comprises the following steps:
[0063] Step S21: performing pixel projection processing on the handwriting pressure change characteristic data in the horizontal direction and the vertical direction, thereby obtaining character projection characteristic data;
[0064] The embodiment of the present invention extracts the pressure value information of each pixel point according to the handwriting pressure change characteristic data obtained in step S16. Then, the image is projected in the horizontal direction (horizontal) and the vertical direction (vertical) respectively. The horizontal projection obtains the total pressure value of each row by summing the pixel pressure values of each row, which indicates the overall stroke density of the text in the row. The vertical projection obtains the total pressure value of each column by summing the pixel pressure values of each column, which indicates the stroke distribution of the text in the column. Through the projection in these two directions, the distribution characteristics of the characters in the horizontal and vertical directions can be effectively obtained to form the character projection feature data. These projection feature data can reflect the overall structure of the characters and the density changes of the strokes, and are the basic data for the subsequent character spacing calculation and standardization processing. In specific application scenarios, for example, when processing scientific and technological handwritten archives, this projection method can accurately capture the details of the handwriting, especially for those characters with sparse or overlapping strokes, which can provide a clear reference for subsequent processing.
[0065] Step S22: performing statistical analysis on the blank areas between adjacent characters according to the character projection feature data, thereby obtaining character spacing feature data, wherein the character spacing feature data includes line spacing data and character spacing data;
[0066] The embodiment of the present invention analyzes the projection interval between characters according to the character projection feature data obtained in step S21. The blank area between characters is represented as a part with a low or zero pressure value in the projection data. Therefore, by detecting the zero-value interval or the low-value interval in the projection data, the blank area between adjacent characters can be effectively segmented. Then, the length of these blank areas is counted, which is the character spacing; if in the same line, the length of the blank area between the lines is counted, which is the line spacing. In the calculation process, the values of the character spacing and line spacing are usually expressed in pixel units or character width units. Finally, by analyzing the character spacing data, the spacing features between each character can be obtained. These spacing features are helpful for subsequent character layout analysis and standardization processing. In actual application scenarios, such as in the process of scientific and technological handwriting file recognition, the character spacing and line spacing feature data are very important for distinguishing the layout of each character and line, especially when the handwriting is unclear or the characters overlap, accurate spacing analysis can significantly improve the accuracy of recognition.
[0067] Step S23: performing sliding window analysis on the handwriting feature data using the character spacing feature data, thereby obtaining local area feature data;
[0068] According to the character spacing and line spacing data obtained in step S22, the embodiment of the present invention sets the size and step length of the sliding window. The size of the window can be adjusted according to the average width of the character, usually 3 to 5 character widths. For example, a window size of 5 character widths is selected, and a step length of 1 character width is set. Then, in the handwriting feature data, sliding is performed according to the window size and step length, and the handwriting feature data in the window is gradually extracted. Each time the handwriting feature data in the window is slid, it will be merged and processed to generate new local area feature data. These local area feature data can effectively represent the distribution of characters in a certain area and the handwriting changes. Through this sliding window analysis method, the handwriting features of different areas in the document can be captured in detail, especially when there are complex deformations or overlaps between characters, the feature data of the local area can help analyze the precise position and morphological changes of the characters. In the application scenario, this method is particularly useful when processing low-quality scanned handwritten documents, and it can enhance the accuracy of overall recognition through local analysis.
[0069] Step S24: performing breakpoint connectivity analysis on the local area feature data, thereby obtaining standardized character sequence data including character inclination angle data and stroke connection feature data.
[0070] The embodiment of the present invention performs a breakpoint connectivity analysis on the local area feature data obtained in step S23. The analysis process identifies whether there are breaks or connections between characters by performing a connectivity analysis on the strokes in each local area. If there are obvious breakpoints in the strokes of a character, it means that the strokes of the character may not be completely connected due to scanning quality problems or irregular writing. Then, based on the connection of the strokes in the local area, the inclination angle and stroke connection features of the character are extracted. The inclination angle data is calculated by analyzing the main stroke direction of the character, and the least squares method can be used to fit the direction line of the stroke to obtain the inclination angle of each character. The stroke connection feature data determines the connection mode between characters based on the stroke connectivity of the characters, whether there are strokes connected or separated. Through this analysis, the structural characteristics and arrangement relationships of the characters can be standardized to generate standardized character sequence data, which is convenient for subsequent character recognition and text parsing. In practical applications, especially when processing handwritten characters, the tilt angle and stroke connection features are crucial for accurately recognizing character morphology. Especially in scientific or legal documents, the tilt and connectivity analysis of characters can help restore unclear characters or spliced characters, thereby improving recognition accuracy.
[0071] The present invention performs pixel projection processing in the horizontal and vertical directions on the handwriting pressure change feature data, and can effectively extract the projection features of the characters in two directions. This processing method allows the morphology and distribution features of the characters to be quantified, providing intuitive and accurate data support for subsequent character analysis and processing. Through pixel projection, the width, height and approximate position of the characters in the document can be clearly observed, which lays a foundation for further character spacing analysis and standardization processing, and helps to improve the accuracy and efficiency of character recognition. According to the character projection feature data, the blank area between adjacent characters is statistically analyzed to obtain character spacing feature data, including line spacing and word spacing data. The acquisition of character spacing features is crucial for understanding the typesetting structure and writing style of the document. Accurate line spacing and word spacing data can help identify the boundaries of different paragraphs, sentences and words in the document, making the structured analysis of the document more accurate. In addition, reasonable character spacing analysis can also assist in identifying the possible connection or break of the pen during the writing process, providing an important basis for the subsequent breakpoint connectivity analysis, and further improving the accuracy and integrity of text processing. Using the character spacing feature data to perform sliding window analysis on the handwriting feature data, the feature information of the local area can be accurately obtained. The sliding window analysis method can scan the document character by character or character group by character, so as to analyze the handwriting characteristics and arrangement of each local area in detail. This localized analysis method helps to have a deeper understanding of the relationship between characters and the writing coherence, and provides a more detailed and detailed feature description for subsequent standardization processing. In this way, abnormal characters or writing errors in the document can be effectively identified, providing support for further intelligent error correction and text optimization. Breakpoint connectivity analysis is performed on the local area feature data, and finally standardized character sequence data including character inclination angle data and stroke connection feature data is obtained. Breakpoint connectivity analysis can repair the situation of stroke breakage or improper connection that may occur during the writing process, making the character morphology more standardized and unified. The acquisition of character inclination angle data helps to correct the character inclination problem caused by writing habits or scanning angles in the document, and further improve the readability and aesthetics of the text. The extraction of stroke connection feature data ensures the continuity and integrity between characters, making the final generated standardized character sequence data more accurate and standardized, providing high-quality data input for subsequent research term matching, intelligent error correction, and scientific and technological research and development record index construction steps, greatly improving the performance and effectiveness of the entire scientific and technological handwritten archive digital processing process, and ensuring the accurate transmission and effective use of research information.
[0072] Preferably, step S24 includes the following steps:
[0073] Step S241: calculating the inclination angle of the entire text according to the local area feature data, thereby obtaining character inclination angle data;
[0074] The embodiment of the present invention selects a part of representative areas, such as the central area of the characters or the area containing most of the characters, based on the local area feature data obtained in step S23. Then, the edge detection algorithm (such as the Canny edge detection algorithm) is used to extract the edge of the characters in the area to obtain the edge information of the characters. Then, the edge information is fitted using the Hough Transform or the least squares fitting algorithm to calculate the main stroke directions of the characters in the area. By analyzing the inclination angles of multiple local areas, the inclination angle of the entire text is obtained. In a specific application scenario, for example, in the processing of scientific and technological handwritten archives, if there is obvious character inclination in the document (such as irregular handwriting, paper inclination, etc.), this step can effectively extract the inclination angle of the entire text, which is convenient for subsequent standardization processing and text recognition.
[0075] Step S242: performing morphological processing on the ink diffusion feature data, thereby obtaining stroke contour data including stroke width data and stroke direction data;
[0076] The embodiment of the present invention uses morphological operations (such as corrosion, expansion, etc.) to process the diffusion range of ink according to the ink diffusion feature data obtained in step S16. Specifically, the corrosion operation is used to remove the noise points in the ink diffusion, and the ink contour is enhanced by the expansion operation to make the edge of the stroke clearer. Then, the contour of the stroke is extracted using a contour extraction algorithm (such as the findContours function in OpenCV), and the width and direction of the stroke are calculated by the minimum circumscribed rectangle of the contour. The stroke width is obtained by calculating the aspect ratio of the minimum circumscribed rectangle of the contour boundary, and the stroke direction is obtained by calculating the angle of the minimum circumscribed rectangle. These stroke width and direction data can accurately describe the writing characteristics of the characters, which is very important for subsequent character standardization and morphological repair. In the processing of scientific and technological handwritten archives, especially when processing documents such as reports or prescriptions, accurately obtaining the width and direction of the strokes helps to judge the writing clarity and position deviation of the characters.
[0077] Step S243: thinning the connection area between adjacent strokes according to the stroke outline data, thereby obtaining stroke intersection data;
[0078] The embodiment of the present invention analyzes the starting point and the ending point of each stroke based on the stroke outline data obtained in step S242, and identifies areas where there may be connections or intersections. Then, the opening and closing operations of image morphology are used to further refine these connection areas to ensure that the intersection points of adjacent strokes can be clearly distinguished. If there are slight breaks or intersections between the strokes, morphological processing can help bridge these areas. Then, connected component analysis (Connected Component Analysis) or region growing algorithm is used to further identify and locate the positions of the stroke intersections. These intersections are usually where multiple strokes intersect, which is crucial for restoring the integrity and coherence of the characters. In practical applications, especially when dealing with blurred or partially overlapping handwriting, stroke intersection data can help identify and repair breaks between characters, thereby improving the accuracy of document recognition.
[0079] Step S244: repairing possible breakpoints according to the stroke intersection data, thereby obtaining stroke connection feature data;
[0080] The embodiment of the present invention identifies potential breakpoints between adjacent strokes based on the stroke intersection data obtained in step S243. Then, a curve fitting or interpolation algorithm (such as Bezier curve fitting or spline interpolation) is used to repair the breakpoints. For each breakpoint, the algorithm infers the most likely connection method based on the contextual information of the intersection (including the stroke direction, width and other features at both ends of the break), performs interpolation calculations, and fills in the broken part. The repaired stroke connection feature data will contain information such as the starting and ending points, direction, width, etc. of each stroke, and the connectivity between characters is improved. In application scenarios, such as in the process of studying document recognition, especially when dealing with overlapping or incomplete characters, breakpoint repair can significantly improve the quality of document recognition and reduce character separation or ambiguity.
[0081] Step S245: Standardize the text in the original scanned image data according to the character inclination angle data and the stroke connection feature data, so as to obtain standardized character sequence data.
[0082] The embodiment of the present invention uses the overall text tilt angle data obtained in step S241 to perform rotation correction on the original scanned image, adjust all the text to a uniform tilt angle, and ensure the consistency of the horizontality of the text. Then, according to the stroke connection feature data obtained in step S244, the characters in the original scanned image are partially repaired and reconstructed, especially for characters with connection problems between strokes, the continuity of the characters is restored by a connection repair algorithm. Finally, through the above processing, the standardized character sequence data obtained will be a unified format, clear and coherent character sequence. These standardized character sequence data facilitate subsequent character recognition, text parsing and information extraction. In specific application scenarios, such as the processing of scientific and technological handwritten archives, this standardization process can effectively improve the recognition rate of documents, especially when processing handwritten texts with tilted characters, poor connections or blurry characters, the standardized processing can provide more accurate character and information recognition results.
[0083] The present invention can effectively identify and quantify the degree of inclination of characters in a document by calculating the inclination angle of the entire text and obtaining character inclination angle data. This correction of the inclination angle is crucial for subsequent character recognition and text processing, because it can adjust the inclined characters to a standard vertical or horizontal state, thereby improving the accuracy and consistency of character recognition. For example, in some handwritten documents, due to the writer's habits or the placement angle problem during scanning, the characters may appear inclined to varying degrees. This step can accurately correct this inclination, making the characters more regular and convenient for subsequent processing. The ink diffusion feature data is morphologically processed, and the stroke width and direction data are extracted to form stroke contour data. This processing helps to more accurately describe the shape and structure of the strokes, and provides an important basis for the connectivity analysis of the strokes. Through morphological processing, the thickness, direction and shape changes of the strokes can be clearly identified, which is very helpful for understanding the writing style and habits of the writer. At the same time, the acquisition of stroke contour data also provides an accurate reference for subsequent stroke intersection detection and breakpoint repair, further improving the refinement of text processing. According to the stroke outline data, the connection area between adjacent strokes is refined to obtain the stroke intersection data. This step can accurately identify the intersections between strokes, which is crucial for repairing breakpoints and improper connections during the writing process. Through the refinement process, the connection position of the strokes can be more accurately located, providing accurate coordinate information for subsequent breakpoint repair. This refinement process helps to improve the coherence and integrity of the text, making the connection between characters more natural and smooth. According to the stroke intersection data, the possible breakpoints are repaired to obtain the stroke connection feature data. This repair process can effectively solve the problem of stroke breakage that may occur during the writing process, making the character shape more complete and standardized. By repairing the breakpoints, the accuracy and reliability of character recognition can be improved, and the recognition errors caused by stroke breakage can be reduced. At the same time, the generation of stroke connection feature data also provides important feature support for the subsequent character standardization process, further improving the quality and effect of text processing. Finally, the text in the original scanned image data is standardized according to the character inclination angle data and the stroke connection feature data to obtain standardized character sequence data. This step is the key link in the entire processing flow. It combines the results of all previous steps and makes the final standardized adjustment to the text. Through standardization, characters with different writing styles and different degrees of inclination can be unified into a standard form and format, thereby generating high-quality standardized character sequence data. This standardized character sequence data is not only easy to store and manage, but also provides a solid foundation for subsequent research term matching, intelligent error correction, and scientific and technological research and development record index construction. It greatly improves the overall performance and effect of the digital processing of scientific and technological handwritten archives, and ensures the accurate transmission and effective use of research information.
[0084] Preferably, step S3 comprises the following steps:
[0085] Step S31: using a preset research term verification word library to perform term matching processing on the standardized character sequence data, thereby obtaining term similarity data;
[0086] The embodiment of the present invention needs to obtain a preset research term verification word library. The word library contains terms and variant writings in a specific field, and is usually established by an expert team in documents or databases in a specific field. In order to improve the accuracy of matching, the term verification word library will be pre-classified and indexed, and the terms will be grouped according to their subject categories or usage frequencies. Afterwards, a word segmentation algorithm (such as a BERT model based on deep learning) is used to segment the standardized character sequence data to extract possible term candidate words. Then, through word frequency statistical analysis, the candidate word sequence data is obtained, and the frequency of occurrence of each word and its co-occurrence probability with other phrases are calculated. This information is used to perform fuzzy matching on the candidate word sequence, and finally obtain term similarity data, indicating the similarity of each candidate word with the term in the word library. In the application scenario, assuming that a character sequence in a certain scientific paper is matched with terms, the implementation of step S31 can help extract terms closely related to the research field from the document, providing a basis for subsequent error correction and optimization.
[0087] Step S32: screening the matching results according to the term similarity data, thereby obtaining the term data to be verified, wherein the term data to be verified includes completely matched term data and suspected erroneous term data;
[0088] The embodiment of the present invention screens the matched results according to the term similarity data, with the purpose of identifying the complete matches and suspected erroneous terms. The term similarity data will reflect the degree of match between each candidate term and the standard term in the vocabulary. Candidate terms with a match higher than a preset threshold (for example, 0.9) are identified as complete matches, while those with a match lower than the threshold are considered suspected erroneous terms. For complete match term data, subsequent processing is performed directly, while for suspected erroneous terms, further analysis is required, especially considering the possible term variants and writing differences in the field. Assuming that in a certain file, "stem cell" is identified as a suspected erroneous term, step S32 will mark it as data to be verified so that it can enter the next stage of correction processing.
[0089] Step S33: performing context association analysis on the standardized character sequence data according to the suspected erroneous terminology data, thereby obtaining semantic association data;
[0090] The embodiment of the present invention performs context association analysis on the suspected erroneous term data in order to find out the context or context information that may cause the error. This analysis method is usually based on the context modeling technology in natural language processing. In the specific implementation, a model based on the long short-term memory network (LSTM) or transformer (Transformer) architecture can be used to identify the context of the term. The model analyzes the vocabulary, sentence structure and semantic information around the suspected erroneous term to infer its actual meaning in the context. For example, if "stem cell" is mistakenly marked as a suspected erroneous term, the model may analyze the context before and after it, such as "stem cell culture" or "stem cell therapy", and obtain semantic association data based on these association data. These data will help to further correct the recognition and matching of terms.
[0091] Step S34: Calculate context similarity using semantic association data to obtain term modification suggestion data;
[0092] The embodiment of the present invention further verifies the correctness of the term by calculating the context similarity. In implementation, the suspected erroneous term data and context-related data are first used to calculate the context similarity through a deep learning model. The specific method can measure the similarity between contexts by calculating the cosine similarity of word vectors. For example, for the term "stem cell", its similarity with other related terms (such as "cell culture") can be calculated by semantically parsing its context. If the context similarity is high, the system will think that the term is likely to be correct, otherwise it may be an erroneous term. Through this process, the term correction suggestion data will provide an important basis for the next step of intelligent error correction.
[0093] Step S35: intelligently correcting the suspected erroneous term data according to the term correction suggestion data, and performing credibility evaluation, so as to obtain the optimal matching term data;
[0094] The embodiment of the present invention performs intelligent error correction processing on suspected erroneous terms based on the term correction suggestion data and evaluates its credibility. During the implementation process, a machine learning algorithm (such as a random forest or a support vector machine) is used to correct suspected erroneous terms, and the model is trained based on its context similarity data and previously accumulated term data. By scoring the credibility of the candidate corrected terms (for example, using statistical methods to evaluate whether the corrected terms conform to the commonly used terms in the field), the model generates term correction suggestions. Assuming that "stem cells" are misidentified as "fetal cells", after intelligent error correction processing, the system will evaluate the credibility of the correction based on parameters such as the frequency of related terms in the field and the accuracy of contextual semantics, and ultimately determine "stem cells" as the best matching term.
[0095] Step S36: Perform normalization and reconstruction processing on the standardized character sequence data according to the complete matching term data and the optimal matching term data, so as to obtain normalized research text data.
[0096] The embodiment of the present invention performs normalization and reconstruction processing on the standardized character sequence data according to the exact matching terms and the optimal matching terms. This step regenerates the standardized research text data by inserting the exact matching terms and the corrected optimal terms into the standardized character sequence data. Specifically, step S36 replaces the original suspected erroneous terms with the re-corrected terms to ensure that the entire text data meets the standardization requirements in terms of term consistency and accuracy. During the reconstruction process, if there are multiple writings or synonyms of the terms, the system will select the most appropriate version of the term based on the aforementioned word frequency and semantic similarity data to ensure the fluency and academic nature of the text. In application scenarios, such as when processing a scientific paper, term reconstruction will make the professional terms in the document conform to the field specifications, and finally obtain standardized research text data to ensure the accuracy and usability of the document.
[0097] The present invention uses a preset research term verification word library to perform term matching processing on the standardized character sequence data, thereby obtaining term similarity data. This step can effectively identify and quantify the similarity between the research terms in the document and the standard terms in the preset word library, providing a scientific basis for subsequent term screening and error correction. Through term matching, professional terms in the document can be quickly located, the accuracy and efficiency of term recognition can be improved, and the professionalism and accuracy of the research document can be ensured. The matching results are screened according to the term similarity data to obtain the term data to be verified, including fully matched term data and suspected erroneous term data. This screening process helps to distinguish the correct terms in the document from possible erroneous terms, providing a clear direction for further analysis and processing. Through screening, unnecessary processing steps can be reduced, the efficiency of the entire processing flow can be improved, and the pertinence and effectiveness of subsequent processing can be ensured. Context association analysis is performed on the standardized character sequence data according to the suspected erroneous term data, thereby obtaining semantic association data. Context association analysis can take into account the contextual environment of the term in the document and provide richer semantic information for the correction of the term. By analyzing the context, the meaning and usage of the term can be understood more accurately, thereby improving the accuracy and reliability of the term correction and avoiding erroneous correction caused by processing the term in isolation. The context similarity is calculated using semantic association data to obtain term correction suggestion data. Calculating context similarity helps determine the best match of a term in a specific context and provides specific suggestions for the correction of the term. In this way, it can be ensured that the corrected term is not only consistent with the standard term in form, but also semantically consistent with the context of the document, further improving the quality and consistency of the document. Intelligent error correction processing is performed on suspected erroneous term data based on the term correction suggestion data, and credibility assessment is performed to obtain the optimal matching term data. Intelligent error correction processing combined with credibility assessment can automatically correct erroneous terms in the document and ensure the reliability of the correction results. This step not only improves the accuracy of the document, but also reduces the need for manual intervention, improves processing efficiency, and ensures the professionalism and credibility of the document. The standardized character sequence data is normalized and reconstructed based on the complete matching term data and the optimal matching term data to obtain the normalized research text data. The normalized reconstruction processing can reintegrate the corrected term into the document to generate high-quality normalized research text data. This standardized processing not only facilitates the storage and management of documents, but also provides a solid foundation for subsequent steps such as research term matching, intelligent error correction, and index construction of scientific and technological research and development records. It greatly improves the overall performance and effect of digital processing of scientific and technological handwritten archives, and ensures the accurate transmission and effective use of research information.In summary, these steps significantly improve the quality of digital processing of scientific and technological handwritten archives through precise term matching, screening, context analysis, intelligent error correction and standardized reconstruction, ensure the accuracy and consistency of research information, and provide strong technical support for the information management of the research industry.
[0098] Preferably, step S31 includes the following steps:
[0099] Step S311: obtaining preset research term verification word library data; performing classification and indexing processing on the research term verification word library data, thereby obtaining term feature data, wherein the term feature data includes standard term data and variant writing data;
[0100] The embodiment of the present invention needs to obtain a preset research term verification word library data. The word library is usually compiled by field experts based on literature, research reports and industry standards, and contains a large number of field terms, variants of terms and their common writing methods. In order to improve the matching accuracy, the terms in the word library will be classified and indexed according to their semantic relationship, frequency of use, and field characteristics. For example, all biomedical terms can be classified into one category, and chemical field terms can be classified into another category. During the processing, a text parsing method based on regular expressions is used to identify the terms and their variants, and mark them as standard terms and variant writing methods. For example, "cell culture" can be regarded as a standard term, while "cell cultivation" and "cell cultivation" are its variant writing methods. Finally, the term feature data obtained by the classification indexing technology will contain the standard writing method and variant writing method of each term, forming a structured data set. These term feature data will provide a basis for the processing of standardized character sequence data in subsequent matching and analysis.
[0101] Step S312: using a word segmentation algorithm to perform sequence segmentation processing on the standardized character sequence data, thereby obtaining candidate word sequence data;
[0102] The embodiment of the present invention needs to perform word segmentation processing on the standardized character sequence data, the purpose of which is to extract possible term candidate word sequences from the original text. The choice of word segmentation algorithm usually depends on the nature of the data and the characteristics of the field. Here, commonly used word segmentation methods include rule-based word segmentation algorithms, statistical word segmentation algorithms (such as maximum entropy models) and deep learning methods (such as BERT models). For example, when processing a research article in the medical field, the character sequence is first separated by spaces and punctuation marks, and then the pre-trained BERT model is used to perform a more fine-grained segmentation of the text, which can identify term-level segmentations such as "gene editing" and "protein structure". In this process, the text is segmented into multiple candidate word sequence data, and the output data includes multiple possible combinations of single words and their combinations. Finally, these candidate word sequence data are passed to subsequent steps for word frequency statistics and matching processing to ensure that each possible term variant is detected.
[0103] Step S313: Perform word frequency statistical analysis on the candidate word sequence data to obtain word frequency feature data, wherein the word frequency feature data includes word occurrence frequency data and phrase co-occurrence probability data;
[0104] The embodiment of the present invention performs a word frequency statistical analysis on the obtained candidate word sequence data to extract word frequency feature data. The core purpose of the analysis is to count the frequency of occurrence of each word and phrase and the probability of co-occurrence between them, so as to reflect the importance and connection of the vocabulary in the corpus. In the implementation process, a statistical method based on the n-gram model is usually used to calculate the word frequency. In the application scenario, if the terms in a scientific document are analyzed, the frequency of phrases such as "cancer research" and "cancer treatment" can be counted through the n-gram model, and the potential association between them can be revealed by calculating their co-occurrence probability in the text (for example, the frequency of occurrence in the same paragraph or sentence). These statistical information will be used to further analyze which words are frequently used key terms and which phrases often co-occur together, thereby providing valuable data support for subsequent fuzzy matching processing.
[0105] Step S314: perform fuzzy matching processing on the candidate word sequence data using the term feature data, thereby obtaining term similarity data.
[0106] The embodiment of the present invention uses term feature data to perform fuzzy matching on candidate word sequence data, with the purpose of calculating the similarity between each candidate word and the term in the standard term library. This process usually uses methods such as edit distance, cosine similarity, Jaccard similarity, etc. for fuzzy matching. For example, the similarity can be evaluated by calculating the edit distance between the candidate word and the standard term (i.e., the minimum number of operations for insertion, deletion, and replacement), or the cosine similarity can be used to measure the degree of similarity in the word vector space. Assuming that a candidate word "gene therapy" is similar to the standard term "gene therapy" in the word library, the edit distance is 1 and the cosine similarity is 0.85, then the candidate word can be regarded as highly similar to the standard term. The system will screen out possible terms according to a preset similarity threshold (for example, a similarity greater than 0.8 is considered a successful match) and generate term similarity data. These data will be used for subsequent term correction and verification processing to ensure that the terms in the text are accurate and meet the field specifications.
[0107] The present invention can effectively identify and quantify the similarity between the research terms in the document and the standard terms in the preset vocabulary, providing a scientific basis for subsequent term screening and error correction. Through term matching, professional terms in the document can be quickly located, the accuracy and efficiency of term identification can be improved, and the professionalism and accuracy of the research document can be ensured. This screening process helps to distinguish the correct terms in the document from possible erroneous terms, providing a clear direction for further analysis and processing. Through screening, unnecessary processing steps can be reduced, the efficiency of the entire processing flow can be improved, and the pertinence and effectiveness of subsequent processing can be ensured. Context association analysis can consider the context of the terms in the document and provide richer semantic information for the correction of the terms. By analyzing the context, the meaning and usage of the terms can be understood more accurately, thereby improving the accuracy and reliability of the term correction and avoiding erroneous corrections caused by processing the terms in isolation. Calculating context similarity helps to determine the best match of the terms in a specific context and provide specific suggestions for the correction of the terms. In this way, it can be ensured that the corrected terms are not only consistent with the standard terms in form, but also semantically consistent with the context of the document, further improving the quality and consistency of the document. Intelligent error correction processing combined with credibility assessment can automatically correct incorrect terms in documents and ensure the reliability of the correction results. This step not only improves the accuracy of the document, but also reduces the need for manual intervention, improves processing efficiency, and ensures the professionalism and credibility of the document. The standardized character sequence data is normalized and reconstructed according to the complete matching term data and the optimal matching term data to obtain the standardized research text data. The normalized reconstruction processing can reintegrate the corrected terms into the document to generate high-quality normalized research text data. This normalization processing not only facilitates the storage and management of documents, but also provides a solid foundation for subsequent research term matching, intelligent error correction, and scientific and technological research and development record index construction. It greatly improves the overall performance and effect of the digital processing of scientific and technological handwritten archives and ensures the accurate transmission and effective use of research information. In summary, these steps significantly improve the quality of digital processing of scientific and technological handwritten archives through precise term matching, screening, context analysis, intelligent error correction, and normalized reconstruction, ensure the accuracy and consistency of research information, and provide strong technical support for the information management of the research industry.
[0108] Preferably, step S4 comprises the following steps:
[0109] Step S41: Acquire research document structure template data; perform hierarchical parsing on the research document structure template data, distinguish between fixed field positions and variable field ranges, and thereby obtain field template data;
[0110] The embodiment of the present invention obtains preset research document structure template data. The structure template is usually designed by an expert team in the research field according to common document formats and structures, including multiple parts such as cover, abstract, introduction, method, result, discussion, conclusion, etc. In order to realize hierarchical parsing processing, a rule-based text parsing method or a machine learning model is used to analyze the template data. In specific operations, firstly, a regular expression, a parser based on grammatical rules, or a natural language processing technology (such as an LSTM network, a BERT model) is used to hierarchically split the document structure to identify the position of fixed fields and the range of variable fields. For example, the "Introduction" part of the document usually starts from the first paragraph of page 1 and ends at the second paragraph of page 2, and this part is a fixed field. The position of the "Method" part has a certain variability and needs to be identified by markers in the document (such as "method", "experimental design", etc.). Finally, the obtained field template data contains fixed fields (such as titles, headers, etc.) and variable fields (such as specific content blocks, chapters) in the document and their corresponding range information, forming a structured field template.
[0111] Step S42: performing layout analysis on the standardized research text data through the field template data, thereby obtaining document layout data;
[0112] The embodiment of the present invention performs layout analysis on the standardized research text data through field template data, with the purpose of determining the layout information in the document. In the specific operation, it is first necessary to identify the structural information such as each chapter, paragraph and page number of the document according to the field template data in step S41. Then, image processing technology (such as OCR recognition or layout analysis tool) is used to analyze the layout of the document, extract the information such as page margins, font size, line spacing of each page, and analyze the distribution of fixed fields and variable fields. By comparing the positions of fixed fields and variable fields, the system can automatically determine the layout mode of different areas in the document, such as directory, title, text, table, picture, etc. For example, for a medical research document, the system can identify that the "method" chapter starts from the first page, the text content is located in the center of the page, and the possible chart area is marked, and the layout data of the document is finally formed, including the layout and regional distribution of each part of the document.
[0113] Step S43: extracting regional features of the document layout data based on the basic information area and the data attribute information area, thereby obtaining content partition data; performing deep learning analysis on the data attribute information area information in the content partition data using a pre-trained text classification model, thereby obtaining semantic classification data;
[0114] The embodiment of the present invention extracts regional features from document layout data, and uses a pre-trained text classification model to perform deep learning analysis. First, the document is divided into regions to distinguish between basic information regions and data attribute information regions. For example, the basic information region usually includes the title, author, date, abstract, etc. of the document, while the data attribute information region includes research data, experimental settings, etc. in the text. These regions are semantically classified using a deep learning model (such as a convolutional neural network CNN or a BERT model) to extract content features of each region. For example, in a document about drug research, the system can identify that "experimental design" and "research results" are data attribute information regions, while "abstract" and "conclusion" belong to the basic information region. Based on these classifications, the system eventually obtains content partition data, identifying the types of different regions and their locations in the document.
[0115] Step S44: performing technical field recognition based on parameter fields and label fields on the text content according to the semantic classification data, thereby obtaining field type data;
[0116] The embodiment of the present invention performs technical field recognition on the document text according to the semantic classification data, which is mainly achieved by analysis based on parameter fields and label fields. First, according to the content partition data obtained in step S43, the system analyzes the specific text in the data attribute information area and extracts relevant parameter fields (such as "drug dosage", "research method") and label fields (such as "experimental group", "control group"). This step relies on natural language processing technology, using vocabulary matching, dependency parsing and named entity recognition (NER) and other methods to accurately identify the technical fields in the text. For example, in medical research, the parameter field may be "patient's age", "drug concentration in the experiment", and the label field may be "treatment group" or "control group". Through this process, the system can accurately identify various field types in the document and obtain field type data.
[0117] Step S45: performing spatial mapping processing on the text in the document according to the field type data and the original scanned image data, thereby obtaining coordinate mapping data;
[0118] The embodiment of the present invention performs spatial mapping processing on the text in the document according to the field type data and the original scanned image data, with the purpose of obtaining the coordinate information of the text in the document. This step involves the use of OCR technology and spatial coordinate mapping algorithm. During the specific operation, the OCR engine (such as Tesseract or Google Vision) converts the text in the scanned image into digital text and identifies the spatial position of each character (i.e., the coordinates of the character). Then, the specific position of each technical field in the document is identified according to the field type data, and this information is converted into absolute coordinates (such as the coordinate system position on A4 paper). For example, the coordinates of the title field may be (50,100) to (500,150), while a paragraph in the text may be located between (50,200) and (500,250). In this way, the obtained coordinate mapping data can accurately reflect the spatial position of each field in the document.
[0119] Step S46: constructing a text position index using the coordinate mapping data, thereby obtaining position positioning data, wherein the position positioning data includes absolute coordinate data and relative position data;
[0120] The embodiment of the present invention uses coordinate mapping data to construct a text position index. First, based on the coordinate mapping data obtained in step S45, the system creates a position index for each text or field in the document. This process can use inverted indexing technology to store the spatial position of each field or word in an index structure. The index of each word or field not only includes its absolute coordinates in the document (such as (50, 100)), but also includes its relative position in the page (such as the distance relative to the top of the page). For example, if "drug dosage" is located at the coordinates of (100, 200) in the document, the system will create an index for this field and record its specific position in the document. Finally, the mapping data of all fields and their positions are summarized into a position positioning data set, which contains the absolute coordinates and relative position data of all texts.
[0121] Step S47: extract keywords from the research text content according to the field type data and the location positioning data, thereby obtaining search feature data; perform weight calculation on the search feature data based on the field importance and term relevance, thereby obtaining index weight data;
[0122] The embodiment of the present invention performs keyword extraction on the document content according to the field type data and the position positioning data. First, according to the field type data identified in step S44, the system extracts keywords related to each field from the document. For technical fields (such as "drug concentration" and "experimental results"), the system uses the TF-IDF algorithm or the BERT model to extract keywords, thereby obtaining high-frequency keywords. In addition, the system assigns a weight to each keyword based on the term relevance and field importance. For example, "experimental design" is a key field in the document, so the weight of this keyword is higher, while "side effects" as a secondary field has a lower weight. Through these analyses, the system finally obtains the retrieval feature data, and performs weighted calculation based on the field importance and term relevance to obtain the index weight data.
[0123] Step S48: Perform multi-dimensional index construction processing on the standardized research text data according to the location positioning data and the index weight data, so as to obtain the research and development record index data.
[0124] The embodiment of the present invention constructs a multidimensional index for normalized research text data based on position positioning data and index weight data. First, the system combines the coordinate information and index weight data in the position positioning data to construct a multidimensional index structure, usually using an inverted index or a B+ tree index method. The multidimensional index can map each field or keyword in the document to its spatial position in the document and sort them according to the weight of the field. For example, for a paper on cancer research, the system will set the fields related to "cancer treatment" as high-weight indexes and indicate their specific locations in the index structure. Finally, the system constructs a complete R&D record index data based on the multidimensional index, and users can quickly locate relevant content in the document based on keywords or field types.
[0125] The present invention can ensure that in subsequent processing, the system can quickly and accurately identify and locate key information fields in the document. Hierarchical parsing processing helps to improve the efficiency and accuracy of field recognition, reduce information extraction errors caused by changes in field positions, and thus improve the accuracy and efficiency of the entire digital processing. Layout analysis processing can help the system understand the overall structure and layout of the document, and provide a basis for further information extraction and processing. Through layout analysis, different areas in the document, such as titles, text, tables, etc., can be more accurately identified, thereby improving the accuracy and completeness of information extraction. Regional feature extraction and semantic classification analysis help to more accurately identify and classify different content areas in the document, such as basic information areas and data attribute information areas. Through the application of deep learning models, the accuracy and reliability of semantic classification can be improved, providing a more accurate basis for subsequent field recognition and information extraction. It can effectively identify key research information fields in documents, such as conditions, tests, operations, etc., and provide important support for the structured processing of research information. Through accurate field recognition, the readability and usability of research information can be improved, which is convenient for subsequent data analysis and application. According to field type data and original scanned image data, the text in the document is spatially mapped to obtain coordinate mapping data. Spatial mapping processing can accurately match text content with original image data to achieve text location. Coordinate mapping can more accurately determine the specific location of text in the document, providing support for subsequent index construction and information retrieval. The construction of position index helps to quickly locate specific text or fields in the document and improve the efficiency of information retrieval. Through position positioning data, document browsing and information search can be more convenient, improving user experience. Keywords are extracted from the research text content according to field type data and position positioning data to obtain retrieval feature data, and the retrieval feature data is weighted based on field importance and term relevance to obtain index weight data. Keyword extraction and weight calculation can help the system identify key information in the document and sort it according to its importance and relevance. In this way, the accuracy and relevance of information retrieval can be improved, making it easier for users to quickly find the information they need. Multidimensional index construction can achieve efficient organization and management of research text data, facilitating rapid query and retrieval. Through R&D record index data, research information can be shared and utilized more conveniently, improving the quality and efficiency of research services. In summary, these steps significantly improve the quality of digital processing of scientific and technological handwritten archives through precise layout analysis, regional feature extraction, field recognition, spatial mapping, location index construction, keyword extraction and multidimensional index construction, ensure the accuracy and consistency of research information, and provide strong technical support for the information management of the research industry.
[0126] Preferably, the present invention further provides a digital archive management system for executing the digital archive management method described above, wherein the digital archive management system comprises:
[0127] The image preprocessing module is used to obtain the original scanned image data of the scientific and technological handwritten archives; perform regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data;
[0128] A character standardization module is used to perform local projection analysis on the original scanned image data according to the handwriting feature data, and to perform breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data;
[0129] The term processing module is used to perform term matching processing on the standardized character sequence data using the preset research term verification word library, and perform intelligent error correction processing to obtain standardized research text data;
[0130] The data structuring and indexing module is used to perform structured field recognition processing on standardized research text data to obtain field type data; locate the text in scientific and technological handwritten archives based on the field type data and the original scanned image data, and construct an index to generate R&D record index data.
[0131] The present invention also provides a storage medium storing a computer program, wherein the computer program is used to execute the above-mentioned method for managing digital archives.
[0132] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0133] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A method for managing digital archives, characterized in that: The following steps are involved: Step S1: obtaining original scanned image data of scientific and technological handwritten files; performing regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data; Step S2: performing local projection analysis on the original scanned image data according to the handwriting feature data, and performing breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data; Step S3: using a preset research term verification word library to perform term matching processing on the standardized character sequence data, and performing intelligent error correction processing, thereby obtaining standardized research text data; Step S4: performing structured field recognition processing on the standardized research text data to obtain field type data; The text is located in the scientific and technological handwritten archives according to the field type data and the original scanned image data, and an index is constructed to generate R&D record index data.
2. The method for managing digital archives according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: obtaining physical scanning parameter data of the scientific and technological handwritten archive; adjusting the resolution and contrast of the scanning device according to the physical scanning parameter data, thereby obtaining scanning configuration data; Step S12: performing multi-region scanning processing on the scientific handwritten file according to the scanning configuration data, thereby obtaining original scanned image data; Step S13: performing image quality evaluation based on clarity and noise level on the original scanned image data, thereby obtaining image quality feature data; Step S14: performing adaptive region division processing on the original scanned image data according to the clarity information in the image quality feature data, thereby obtaining image block data, wherein the image block data includes core text region data and edge background region data; Step S15: calculating the local brightness mean of each area according to the image block data, thereby obtaining background brightness distribution data; constructing a brightness compensation curve using the background brightness distribution data, thereby obtaining brightness compensation parameter data; Step S16: Adaptively adjust the brightness of the core text area data in the image block data according to the brightness compensation parameter data, thereby obtaining the handwriting pressure change characteristic data; perform edge detection on the core text area data after brightness adjustment, thereby obtaining the ink diffusion characteristic data; Step S17: performing feature fusion processing on the original scanned image data according to the handwriting pressure change feature data and the ink diffusion feature data, so as to obtain the handwriting feature data.
3. The method for managing digital archives according to claim 2, characterized in that: Step S15 includes the following steps: Step S151: Scan point by point in the core text area in the core text area data using a sliding window algorithm, extract the gray value of each pixel, and calculate the average gray value of the pixels in the window, thereby obtaining the local brightness mean data of the core area; Step S152: segmenting the edge background region data by a region growing algorithm, extracting the gray value distribution characteristics of each sub-region, and calculating the average gray value of the sub-region, thereby obtaining the local brightness mean data of the background region; Step S153: Integrate the brightness information in the local brightness mean data of the core area and the local brightness mean data of the background area to establish an overall brightness distribution model, thereby obtaining background brightness distribution data; Step S154: performing mathematical modeling based on the background brightness distribution data, and determining curve parameters by the least square method, thereby generating a brightness compensation curve that can characterize the overall brightness change trend; Step S155: discretizing the brightness compensation curve into specific numerical parameters, and calculating corresponding brightness compensation coefficients for different positions of the image, thereby obtaining brightness compensation parameter data.
4. The method for managing digital archives according to claim 3, characterized in that: Step S2 includes the following steps: Step S21: performing pixel projection processing on the handwriting pressure change characteristic data in the horizontal direction and the vertical direction, thereby obtaining character projection characteristic data; Step S22: performing statistical analysis on the blank areas between adjacent characters according to the character projection feature data, thereby obtaining character spacing feature data, wherein the character spacing feature data includes line spacing data and character spacing data; Step S23: performing sliding window analysis on the handwriting feature data using the character spacing feature data, thereby obtaining local area feature data; Step S24: performing breakpoint connectivity analysis on the local area feature data, thereby obtaining standardized character sequence data including character inclination angle data and stroke connection feature data.
5. The method for managing digitalized archives according to claim 4, characterized in that: Step S24 includes the following steps: Step S241: calculating the inclination angle of the entire text according to the local area feature data, thereby obtaining character inclination angle data; Step S242: performing morphological processing on the ink diffusion feature data, thereby obtaining stroke contour data including stroke width data and stroke direction data; Step S243: thinning the connection area between adjacent strokes according to the stroke outline data, thereby obtaining stroke intersection data; Step S244: repairing possible breakpoints according to the stroke intersection data, thereby obtaining stroke connection feature data; Step S245: Standardize the text in the original scanned image data according to the character inclination angle data and the stroke connection feature data, so as to obtain standardized character sequence data.
6. The method for managing digitalized archives according to claim 5, characterized in that: Step S3 includes the following steps: Step S31: using a preset research term verification word library to perform term matching processing on the standardized character sequence data, thereby obtaining term similarity data; Step S32: screening the matching results according to the term similarity data, thereby obtaining the term data to be verified, wherein the term data to be verified includes completely matched term data and suspected erroneous term data; Step S33: performing context association analysis on the standardized character sequence data according to the suspected erroneous terminology data, thereby obtaining semantic association data; Step S34: Calculate context similarity using semantic association data to obtain term modification suggestion data; Step S35: intelligently correcting the suspected erroneous term data according to the term correction suggestion data, and performing credibility evaluation, so as to obtain the optimal matching term data; Step S36: Perform normalization and reconstruction processing on the standardized character sequence data according to the complete matching term data and the optimal matching term data, so as to obtain normalized research text data.
7. The method for managing digitalized archives according to claim 6, characterized in that: Step S31 includes the following steps: Step S311: obtaining preset research term verification word library data; performing classification and indexing processing on the research term verification word library data, thereby obtaining term feature data, wherein the term feature data includes standard term data and variant writing data; Step S312: using a word segmentation algorithm to perform sequence segmentation processing on the standardized character sequence data, thereby obtaining candidate word sequence data; Step S313: Perform word frequency statistical analysis on the candidate word sequence data to obtain word frequency feature data, wherein the word frequency feature data includes word occurrence frequency data and phrase co-occurrence probability data; Step S314: perform fuzzy matching processing on the candidate word sequence data using the term feature data, thereby obtaining term similarity data.
8. The method for managing digital archives according to claim 7, characterized in that: Step S4 includes the following steps: Step S41: Acquire research document structure template data; perform hierarchical parsing on the research document structure template data, distinguish between fixed field positions and variable field ranges, and thereby obtain field template data; Step S42: performing layout analysis on the standardized research text data through the field template data, thereby obtaining document layout data; Step S43: extracting regional features of the document layout data based on the basic information area and the data attribute information area, thereby obtaining content partition data; performing deep learning analysis on the data attribute information area information in the content partition data using a pre-trained text classification model, thereby obtaining semantic classification data; Step S44: performing technical field recognition based on parameter fields and label fields on the text content according to the semantic classification data, thereby obtaining field type data; Step S45: performing spatial mapping processing on the text in the document according to the field type data and the original scanned image data, thereby obtaining coordinate mapping data; Step S46: constructing a text position index using the coordinate mapping data, thereby obtaining position positioning data, wherein the position positioning data includes absolute coordinate data and relative position data; Step S47: extract keywords from the research text content according to the field type data and the location positioning data, thereby obtaining search feature data; perform weight calculation on the search feature data based on the field importance and term relevance, thereby obtaining index weight data; Step S48: Perform multi-dimensional index construction processing on the standardized research text data according to the location positioning data and the index weight data, so as to obtain the research and development record index data.
9. A digital archive management system, characterized in that: Used to execute the management method of digitalized archives as claimed in claim 1, the management system of digitalized archives comprises: The image preprocessing module is used to obtain the original scanned image data of the scientific and technological handwritten archives; perform regional brightness compensation processing on the original scanned image data to obtain handwriting feature data, wherein the handwriting feature data includes handwriting pressure change feature data and ink diffusion feature data; A character standardization module is used to perform local projection analysis on the original scanned image data according to the handwriting feature data, and to perform breakpoint connectivity analysis to obtain standardized character sequence data including character tilt angle data and stroke connection feature data; The term processing module is used to perform term matching processing on the standardized character sequence data using the preset research term verification word library, and perform intelligent error correction processing to obtain standardized research text data; The data structuring and indexing module is used to perform structured field recognition processing on standardized research text data to obtain field type data; locate the text in scientific and technological handwritten archives based on the field type data and the original scanned image data, and construct an index to generate R&D record index data.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed, the method for managing digital archives as described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Digital archive intelligent processing method, storage medium and system
CN121121786A
Multi-mode intelligent dictation and correction system and method based on dynamic calibration
CN121236768A
Multi-modal intelligent dictation and correction system and method based on dynamic calibration
CN121236768B