Literature report generation evaluation method and system based on large model

By acquiring multi-channel spectral data of document images and performing dynamic bit width allocation to generate structured text descriptions, the problems of color shift and detail loss caused by material differences and printing quality in existing technologies are solved, and the accuracy and logical coherence of document image content recognition are improved.

CN120724992AActive Publication Date: 2025-09-30LUSTER LIGHTWAVE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511213131.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-30
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

When processing image data with complex spectral characteristics, existing technologies have difficulty effectively capturing color shifts and detail loss caused by material differences or printing quality, resulting in insufficient accuracy in document image content recognition. This is especially true in high-density charts or complex mathematical formula scenarios, where key information is easily missed or semantics are misinterpreted.

Method used

By acquiring multi-channel spectral data of document images, dividing the high-sensitivity and low-sensitivity bands, calculating the spectral feature differences, and performing dynamic bit width allocation, spatial continuity and frequency domain distribution coded data are generated, and these coded data are parsed using a large model to generate structured text descriptions.

Benefits of technology

It improves the recognition accuracy of non-text elements, solves the problems of missing key information and semantic misinterpretation, enhances the robustness and information integrity of document analysis, adapts to the characteristic requirements of complex spectral scenes, and ensures that the evaluation results are close to the actual content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724992A_ABST
    Figure CN120724992A_ABST
Patent Text Reader

Abstract

The invention provides a literature report generation evaluation method and system based on a large model, and the method comprises the steps: obtaining a literature image data set containing non-text elements, and the non-text elements comprise chart elements and formula symbols; the method comprises the following steps: separating different spectral band data of a literature image data set to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics, and dividing a high-sensitivity band and a low-sensitivity band of the multi-channel spectral data according to a spectral sensitivity threshold, calculating spectral characteristic difference quantity between the two wave bands and executing bit width distribution to generate spatial continuity coded data and frequency domain distribution coded data; analyzing the two types of coded data through a large model to generate a structured text description, and verifying the information coverage degree of the description relative to the literature image data set to generate an evaluation result; according to the method, the recognition precision and the structural description capability of the non-text elements in the document image are improved, and the requirements of automatic analysis and evaluation of high-quality multimedia documents are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large-model-based literature report generation and evaluation method and system. Background Art

[0002] In general multimedia document analysis scenarios, with the increasing digitization and visualization of academic resources, how to efficiently and accurately process documents containing non-text elements such as charts, images, and formulas has become a key requirement for information extraction and knowledge induction.

[0003] Currently, there are document content understanding solutions based on multimodal fusion. This solution combines convolutional neural networks with the Transformer architecture. It first detects and segments image regions in documents, then uses a pre-trained vision-language model to generate descriptive text for the image content. Furthermore, the semantic consistency between the generated text and the original image content is evaluated through text similarity calculation and keyword matching. However, existing solutions have obvious limitations when dealing with image data with complex spectral characteristics. Because they mainly rely on conventional visual information in the red, green, and blue channels, they cannot effectively capture color shifts and detail loss in document images caused by material differences or printing quality, thus affecting the accuracy of content recognition. The generated text descriptions still have significant room for improvement in terms of structure and information completeness. In particular, when faced with high-density charts or complex mathematical formulas, key information is easily omitted or semantic misinterpretation occurs, limiting their practical application in high-quality document analysis tasks. Summary of the Invention

[0004] The present invention provides a large-model-based literature report generation and evaluation method and system to address the problem of insufficient accuracy in document image content recognition in the prior art; in scenarios with high-density charts or complex mathematical formulas, problems such as omission of key information or semantic misinterpretation are prone to occur.

[0005] In a first aspect, the present invention provides a method for generating and evaluating literature reports based on a large model, comprising: Acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols; Separating different spectral band data of the document image dataset to generate multi-channel spectral data including color distribution characteristics and material reflection characteristics; Dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating a spectral feature difference between the high-sensitivity band and the low-sensitivity band; Performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; Parsing the spatial continuity coded data and the frequency domain distribution coded data by a large model to generate a structured text description; The information coverage of the structured text description relative to the document image dataset is verified to generate an evaluation result.

[0006] Optionally, different spectral band data of the document image dataset are separated to generate multi-channel spectral data containing color distribution characteristics and material reflectance characteristics, including: According to the spectral band range, the document image data set is divided into visible light band data and near infrared band data; Identifying a spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectrum data carrying three primary color intensity information; Detecting spectral absorption feature points related to the material reflection characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength; The visible light spectrum data and the near-infrared spectrum data are channel-joined to form multi-channel spectrum data including color distribution characteristics and material reflection characteristics.

[0007] Optionally, detecting spectral absorption feature points related to material reflection characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength includes: Detecting spectral absorption characteristic points in the near-infrared band data at which the reflection intensity is lower than that at adjacent wavelength positions; According to a preset material reflection characteristic library, the wavelength position of the spectral absorption characteristic point is matched to screen out the target absorption point associated with the material reflection characteristic; Extracting the reflection intensity value at the wavelength position of the target absorption point to generate reflection intensity distribution data carrying a wavelength-intensity mapping relationship; The reflection intensity distribution data is bound to the spatial position information of the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength.

[0008] Optionally, dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating the spectral feature difference between the high-sensitivity band and the low-sensitivity band, includes: Counting the spectral intensity of each spectral channel in the multi-channel spectral data to generate a spectral intensity distribution curve corresponding to each spectral channel; Marking the band interval where the spectral intensity change rate on the spectral intensity distribution curve exceeds a preset spectral sensitivity threshold as a high-sensitivity band, and marking the band interval where the spectral intensity change rate is lower than the preset spectral sensitivity threshold as a low-sensitivity band; A first spectral intensity extreme point is extracted from the high-sensitivity band, and a second spectral intensity extreme point is extracted from the low-sensitivity band. The absolute value of the intensity difference between the first spectral intensity extreme point and the second spectral intensity extreme point is calculated, and the absolute value of the intensity difference is used as the spectral feature difference amount.

[0009] Optionally, performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data includes: Calculating the intensity ratio of the spatial continuity attribute to the frequency domain distribution attribute in the spectral feature difference; Determining, according to the intensity ratio value, a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute; Calculating the spectral intensity variation between adjacent pixel positions in the color distribution feature to generate a set of spatial continuity variation; Calculating the intensity fluctuation amplitudes at different wavelength positions in the material reflection characteristics to generate a frequency domain distribution fluctuation quantity set; encoding the set of spatial continuity variation values ​​according to the first target bit width value to generate spatial continuity coded data; According to the second target bit width value, the frequency domain distribution fluctuation quantity set is coded to generate frequency domain distribution coded data.

[0010] Optionally, determining, according to the intensity ratio value, a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute includes: Taking the intensity ratio value as a weight ratio, performing a multiplication operation on the weight ratio and a preset bit width base to generate an initial bit width value of spatial continuity; According to the fluctuation amplitude range of the frequency domain distribution attribute, the bit width cardinality is proportionally adjusted to generate an initial bit width value of the frequency domain distribution; According to the intensity stability between different wavelength positions, the initial bit width value of the frequency domain distribution is corrected to generate a second target bit width value; The initial bit width value of spatial continuity is modified according to the intensity correlation degree between adjacent pixel positions to generate a first target bit width value.

[0011] In a second aspect, the present invention provides a large-scale model-based literature report generation and evaluation system, comprising: An acquisition module, configured to acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols; A separation module is used to separate different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics; a calculation module, configured to divide the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculate a spectral feature difference between the high-sensitivity band and the low-sensitivity band; an allocation module, configured to perform dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; A generating module, configured to parse the spatial continuity coded data and the frequency domain distribution coded data through a large model to generate a structured text description; A verification module is used to verify the information coverage of the structured text description relative to the document image dataset to generate an evaluation result.

[0012] In a third aspect, the present invention provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a large model-based literature report generation and evaluation method as described in the first aspect above.

[0013] In a fourth aspect, the present invention provides a computer storage medium storing a computer program, which, when executed by a computer, implements a large model-based literature report generation and evaluation method as described in the first aspect.

[0014] In the present invention, a document image dataset containing non-text elements is obtained, wherein the non-text elements include chart elements and formula symbols; different spectral band data of the document image dataset are separated to generate multi-channel spectral data containing color distribution characteristics and material reflectance characteristics; high-sensitivity bands and low-sensitivity bands of the multi-channel spectral data are divided according to a preset spectral sensitivity threshold, and the spectral feature difference between the high-sensitivity band and the low-sensitivity band is calculated; dynamic bit width allocation is performed on the spectral feature difference to generate spatial continuity coding data and frequency domain distribution coding data; the spatial continuity coding data and the frequency domain distribution coding data are parsed by a large model to generate a structured text description; and the information coverage of the structured text description relative to the document image dataset is verified to generate an evaluation result. The technical solution provided by the present invention solves the problem of information loss caused by insufficient separation of non-text elements in existing solutions, and avoids the excessive reliance of traditional methods on text areas; it breaks through the limitation of existing solutions that rely only on red, green and blue channels, and can capture the problem of detail loss in document images caused by material differences (such as paper reflection, ink density) or printing quality problems, improve the recognition accuracy of non-text elements, and provide richer spectral information for subsequent feature extraction; it can effectively screen out spectral information that plays a key role in content recognition, solves the problem of feature redundancy or omission caused by failure to distinguish the importance of bands in existing solutions, provides accurate feature input for subsequent dynamic coding, and improves the model's recognition of complex spectral scenes. Adaptability: This approach addresses the problem of insufficient structured description capabilities caused by a single encoding method in existing solutions. It can adapt to the multi-scale feature requirements of high-density graphs or complex mathematical formulas, avoiding the loss of key information (such as graph boundaries and formula symbol levels) during the encoding process. It also addresses the semantic misinterpretation or information fragmentation problems caused by existing solutions' reliance on shallow visual-language models, improving the logical coherence and information integrity of text descriptions, and demonstrating greater robustness when processing complex visual elements. It also addresses the limitations of existing solutions that rely solely on text similarity matching (such as ignoring image details or semantic deviations), ensuring that evaluation results are closer to the actual content of the document and providing a reliable feedback mechanism for high-quality document analysis tasks. Furthermore, a target bit width value is dynamically allocated based on the intensity ratio of the spatial continuity attribute and the frequency domain distribution attribute in the spectral feature difference. A set of spatial continuity variation quantities is generated by calculating the spectral intensity variation at adjacent pixel positions, while a set of frequency domain distribution fluctuation quantities is generated by calculating the intensity fluctuation amplitude at different wavelength positions. Finally, the two sets are independently encoded according to the target bit width value, outputting spatial continuity encoded data and frequency domain distribution encoded data, respectively.Through the attribute-aware dynamic bit width allocation mechanism, the intensity transition characteristics of the boundaries of chart elements are retained, solving the structural distortion caused by differences in printing quality; the reflection fluctuation characteristics of the formula symbol material are maintained to avoid blurring the key features of complex formula areas; and the coding resources are dynamically allocated according to the proportion of spectral feature differences to achieve a breakthrough balance between compression rate and information integrity.

[0015] These and other aspects of the present invention will become more readily apparent from the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A flowchart of a method for generating and evaluating a literature report based on a large model provided by the present invention is shown; Figure 2 The present invention shows a schematic structural diagram of a large-scale model-based literature report generation and evaluation system; Figure 3 A schematic structural diagram of a computing device provided by the present invention is shown. DETAILED DESCRIPTION

[0018] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0019] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] In general multimedia document analysis scenarios, traditional multimodal fusion solutions are limited to using only red, green, and blue channels of visual information when processing document images containing non-text elements such as charts and formulas. This makes it difficult to deal with color shift and detail loss caused by factors such as material and printing quality, thereby affecting the accuracy of content recognition and the integrity of the generated text. At the same time, existing solutions lack detailed modeling of spatial continuity and frequency domain distribution in feature coding, resulting in insufficient structured description capabilities, especially in high-density charts or complex mathematical formula scenarios. Key information is easily omitted or semantic misinterpretation occurs. To address the above-mentioned shortcomings, the present invention introduces a spectral data analysis mechanism to expand document images from conventional color space to a multi-channel spectral feature space, and further combines a dynamic bit width allocation strategy to extract highly expressive coding data of spatial continuity and frequency domain distribution, respectively, as input to a large model, thereby improving its depth of understanding of complex visual elements and its ability to express structured expressions, and ultimately achieving an effective evaluation of the information coverage of the generated text relative to the original image content, thereby enhancing the quality control link of automatic document parsing. Figure 1 The present invention provides a flowchart of a method for generating and evaluating literature reports based on a large model, such as Figure 1 As shown, the method includes: Step 101: Acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols; In this step, non-text elements refer to visual objects in the document that cannot be directly represented by characters, including chart elements that reflect data relationships and formula symbols that express mathematical logic. Document image datasets refer to a collection of digitized document images collected by spectral imaging equipment, containing spatial information based on pixel coordinates and spectral intensity information based on wavelength channels. Chart elements refer to two-dimensional graphic structures used to visualize data relationships, including geometric shapes composed of coordinate axes, data points, connecting lines, and their fill color attributes. Formula symbols refer to special characters with specific semantics in mathematical expressions, including operators (such as ∑, ∫), variable symbols (such as α, β), and structural markers (such as matrix brackets).

[0022] In one embodiment of the present invention, raw document image data containing chart elements and formula symbols is collected using a high-resolution spectral imaging device. This device integrates dual-band acquisition modules for visible light (400-760nm) and near-infrared (760-1100nm), employing a 1200dpi resolution, line-by-line scanning method to synchronously record the spatial coordinates (X, Y) of each pixel and the intensity values ​​of 256 spectral channels (λ1-λ256). The color information of the chart elements is captured using an array of red, green, and blue primary color sensors, generating visible light spectral data with the printed color gamut distribution of cyan, magenta, yellow, and black. The material reflectance characteristics of the formula symbols are sampled by sampling the reflection intensity at 16 characteristic wavelengths in the near-infrared band (e.g., 850nm and 940nm), generating material reflectance data containing information on ink density and paper texture. The above-mentioned visible light spectrum data and material reflectance data are stored in a three-dimensional tensor format, with the dimensions corresponding to the image width, height, and number of spectral channels, respectively, to obtain a document image dataset. For example, a standard A4 document image generates a 5952×8424×256 raw data cube, where each spatial position (i, j) corresponds to a 256-dimensional spectral intensity vector, which can accurately characterize the color gradient changes of chart borders and the metallic ink reflection characteristics of formula symbols.

[0023] Step 102: Separating different spectral band data of the document image dataset to generate multi-channel spectral data including color distribution characteristics and material reflectance characteristics; In an embodiment of the present invention, a document image dataset is divided into visible light band data and near-infrared band data according to the spectral band range; the spectral response intensity distribution related to color attributes in the visible light band data is identified to generate visible light spectrum data carrying three-primary color intensity information; spectral absorption feature points related to material reflection characteristics in the near-infrared band data are detected to generate near-infrared spectrum data carrying specific wavelength reflection intensity information; the two types of data are channel-joined to form multi-channel spectral data containing color distribution characteristics and material reflection characteristics.

[0024] Step 103: dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating a spectral feature difference between the high-sensitivity band and the low-sensitivity band; In an embodiment of the present invention, the spectral intensity of each spectral channel in the multi-channel spectral data is statistically analyzed to generate a spectral intensity distribution curve corresponding to each spectral channel; the band interval on the curve where the spectral intensity change rate exceeds a preset spectral sensitivity threshold is marked as a high-sensitivity band, and the band interval below the threshold is marked as a low-sensitivity band; the first spectral intensity extreme point is extracted from the high-sensitivity band, and the second spectral intensity extreme point is extracted from the low-sensitivity band, and the absolute value of the intensity difference between the two types of extreme points is calculated as the spectral feature difference.

[0025] Step 104: performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; In an embodiment of the present invention, a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute are determined by calculating the intensity ratio value of the spatial continuity attribute and the frequency domain distribution attribute in the spectral feature difference; the spectral intensity variation between adjacent pixel positions in the color distribution feature is calculated to generate a set of spatial continuity variation quantities; the intensity fluctuation amplitude at different wavelength positions in the material reflection feature is calculated to generate a set of frequency domain distribution fluctuation quantities; and the spatial continuity variation quantity set and the frequency domain distribution fluctuation quantity set are encoded according to the first target bit width value and the second target bit width value, respectively, to generate spatial continuity encoded data and frequency domain distribution encoded data.

[0026] Step 105: parsing the spatial continuity coded data and the frequency domain distribution coded data by a large model to generate a structured text description; In an embodiment of the present invention, the element position coordinate sequence in the spatial continuity coded data is parsed to generate a set of spatial position descriptions of the chart elements, the correspondence between the wavelength and the reflection intensity in the frequency domain distribution coded data is parsed to generate a set of material property descriptions of the formula symbols, and an element symbol association mapping table is generated based on the two types of sets, so as to retrieve the chart element position description from the spatial position description set and retrieve the formula symbol material description from the material property description set according to the table; the chart element position description and the formula symbol material description associated with the same document position are combined into a composite text paragraph containing a chart element topology description and a formula symbol attribute description, and the composite text paragraphs corresponding to all document positions are integrated to generate a structured text description.

[0027] Step 106: Verify the information coverage of the structured text description relative to the document image dataset to generate an evaluation result; In this step, information coverage refers to the ratio of correctly represented non-text element features in the generated structured text description to the total number of non-text element features in the original document, and is used to quantify content completeness. Evaluation results refer to qualitative judgments generated based on information coverage, including qualified tags for indicating completeness or a coordinate-wavelength missing feature list for locating missing features.

[0028] In this embodiment of the present invention, the spatial coordinate information of graphic elements and the wavelength characteristics of formula symbols in structured text descriptions are matched against corresponding elements in a document image dataset using a dual mapping of spatial position and spectral characteristics. The validity of the match is determined by coordinate alignment and wavelength feature comparison. The number of correctly matched non-text elements is counted and divided by the total number of non-text elements in the dataset to obtain information coverage. When this information coverage exceeds a preset completeness threshold, a qualified assessment result is generated; otherwise, a missing report (i.e., the qualified assessment result) is generated based on the coordinate position and wavelength characteristics of the unmatched elements.

[0029] The embodiments of the present invention solve the problems of color shift and detail loss caused by reliance on the red, green, and blue channels in existing document analysis methods, while improving the structured description capability of high-density charts and complex mathematical formulas; enhancing the model's adaptability to material differences and image quality fluctuations; improving the semantic integrity and logical coherence of generated text; and achieving quantitative evaluation of the quality of document content restoration. This comprehensively breaks through the technical barriers of traditional methods in non-text element recognition and structured expression, and provides an efficient and robust solution for high-quality multimedia document analysis.

[0030] The present invention provides a specific embodiment, step 102, separating different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflectance characteristics, specifically comprising the following steps: Step 201: dividing the document image dataset into visible light band data and near infrared band data according to the spectral band range; In this step, the spectral band range refers to the electromagnetic wavelength range (400-2500nm) that can be captured by the spectral imaging device, including the visible light band (400-700nm) that reflects color information and the near-infrared band (700-2500nm) that reflects material information. Visible light band data refers to the collection of spectral intensities within the wavelength range of 400-700nm and is used for color analysis based on human eye perception. Near-infrared band data refers to the collection of spectral intensities within the wavelength range of 700-2500nm and is used for material property analysis based on molecular vibrations.

[0031] In this embodiment of the present invention, an optical spectrometer system is used to physically segment the document image dataset according to spectral band ranges: visible light band data (which carries color information) is output in the wavelength range of 400-700nm, and near-infrared band data (which carries material information) is output in the wavelength range of 700-2500nm. The segmentation process achieves band isolation based on the reflection / transmission characteristics of the dichroic mirror.

[0032] Step 202: Identify the spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectrum data carrying three primary color intensity information; In this step, color attributes refer to the visual characteristics of light reflected from an object's surface in the visible light band, including hue, saturation, and brightness. Spectral response intensity distribution refers to the set of intensity response values ​​for light waves of different wavelengths on the sensor, reflecting the color energy distribution. Three-primary color intensity information refers to the standardized intensity values ​​(range 0-1) of the three primary colors red, green, and blue at the pixel location, used to accurately characterize color. Visible light spectrum data refers to a data structure that integrates spatial coordinates and three-primary color intensity information, in the format of [row coordinate, column coordinate, R intensity, G intensity, B intensity].

[0033] In an embodiment of the present invention, the spectral intensity values ​​of each pixel in the red-sensitive band, the green-sensitive band, and the blue-sensitive band are measured in the visible light band data; intensity distribution data of the red primary color at each pixel is generated based on the spectral intensity value of the red-sensitive band; intensity distribution data of the green primary color at each pixel is generated based on the spectral intensity value of the green-sensitive band; and intensity distribution data of the blue primary color at each pixel is generated based on the spectral intensity value of the blue-sensitive band. The above three types of data are integrated according to spatial position to form visible light spectrum data carrying information on the intensity of the three primary colors.

[0034] Step 203: Detecting spectral absorption feature points related to material reflection characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength; In this step, material reflectance characteristics refer to the characteristic reflection / absorption patterns of an object in the near-infrared band due to its molecular structure. Spectral absorption signatures are wavelengths where the reflection intensity drops suddenly (by >10%), corresponding to the material's molecular resonance absorption peaks. Wavelength-specific reflection intensity information refers to the mapping between the wavelength of the target absorption point and its reflection intensity value (e.g., {1200nm: 0.32, 1550nm: 0.18}). Near-infrared spectral data refers to a data set that stores wavelength-reflection intensity mappings and their corresponding spatial coordinates.

[0035] In an embodiment of the present invention, spectral absorption feature points in near-infrared band data whose reflection intensity is lower than that of adjacent wavelength positions are detected; combined with a preset material reflection characteristic library (which stores the absorption wavelengths of common printed materials), target absorption points that match the material reflection characteristics are screened out; the reflection intensity value at the wavelength position where the target absorption point is located is extracted, and reflection intensity distribution data carrying a wavelength-intensity mapping relationship is generated. This data is then bound to the spatial position information of the near-infrared band data to generate near-infrared spectral data.

[0036] Step 204: performing channel splicing on the visible light spectrum data and the near-infrared spectrum data to form multi-channel spectrum data including color distribution characteristics and material reflection characteristics; In this step, color distribution refers to the spatial distribution of pixel-level color patterns described by visible light spectral data. Material reflectance refers specifically to the spatial distribution of material absorption characteristics described by near-infrared spectral data. Multi-channel spectral data refers to the four-dimensional data matrix formed by combining three visible light channels and a single near-infrared channel, with dimensions [height, width, 4].

[0037] In an embodiment of the present invention, the three primary color channels of visible light spectral data and the single channel of near-infrared spectral data are aligned according to the data dimension; the four-channel data are merged into a unified multi-channel matrix through a spatial coordinate matching circuit, where the first three channels store color distribution characteristics (i.e., red, green, and blue intensities), and the fourth channel stores material reflection characteristics (i.e., reflection intensity at a specific wavelength), and finally output multi-channel spectral data containing both types of characteristics.

[0038] The embodiment of the present invention captures near-infrared features through physical spectrometry, effectively identifying the absorption characteristics of printing ink molecules (such as the characteristic absorption of carbon-based ink at 1200nm), and solving the color distortion problem caused by the fading of old documents; the dual-band independent processing mechanism improves the accuracy of chart color restoration, while achieving precise distinction between the material properties of formula symbols (such as laser printing toner vs. inkjet ink), laying a data foundation for subsequent cross-modal correlation analysis.

[0039] The present invention provides a specific embodiment, step 203, detecting spectral absorption feature points related to the material reflectance characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength, specifically comprising the following steps: Step 211: detecting a spectral absorption feature point in the near-infrared band data where the reflection intensity is lower than that of an adjacent wavelength position; In this step, reflection intensity refers to the near-infrared light energy reflected from the surface of an object, ranging from 0 to 1 (0 = total absorption, 1 = total reflection), reflecting the molecular vibration absorption characteristics of the material. Spectral absorption characteristic points are wavelengths where the reflection intensity suddenly drops. The criterion for determining whether a point's intensity is ≤ 85% of the average intensity of the three adjacent points corresponds to a molecular resonance absorption peak in the material.

[0040] In an embodiment of the present invention, a sliding window method is used to traverse the wavelength sequence of near-infrared band data, and the average reflection intensity of each wavelength position and its three adjacent wavelength positions on the left and right is calculated; when the reflection intensity at the wavelength position is lower than 85% of the average value, that is, the reflection intensity ≤ average value × 0.85, the corresponding wavelength position is marked as a spectral absorption characteristic point.

[0041] Step 212: matching the wavelength positions of the spectral absorption characteristic points according to a preset material reflection characteristic library to screen out target absorption points associated with the material reflection characteristics; In this step, the preset material reflectance characteristic library refers to a pre-stored database of characteristic absorption wavelengths for printing materials, including standard absorption wavelengths and fluctuation ranges for materials such as toner and coated paper, based on laboratory measurements. The target absorption point is a valid absorption point selected through library matching. Its wavelength deviates from the standard value in the library by ≤±5nm and is used to identify a specific printing material.

[0042] In an embodiment of the present invention, the wavelength position of the spectral absorption characteristic point is matched with a preset material reflection characteristic library. Specifically, the material reflection characteristic library stores the characteristic absorption wavelengths of printed materials (such as 1550nm for toner and 1200nm for coated paper). When the wavelength of a certain spectral absorption characteristic point differs from the characteristic wavelength in the library within the range of ±5nm, it is determined to be a target absorption point associated with the material reflection characteristic.

[0043] Step 213: extracting the reflection intensity value at the wavelength position of the target absorption point, and generating reflection intensity distribution data carrying a wavelength-intensity mapping relationship; In this step, the reflection intensity value refers to the measured reflected light energy at the target absorption point's wavelength, expressed as a floating-point number, with lower values ​​indicating stronger absorption. The wavelength-intensity mapping is a key-value pairing of wavelength and corresponding reflection intensity values ​​(e.g., {1550: 0.23, 1700: 0.41}), representing the material's characteristic absorption spectrum. The reflection intensity distribution data is a data structure that integrates the wavelength-intensity mappings for the target absorption points, formatted as [wavelength1:intensity1, wavelength2:intensity2,...].

[0044] In an embodiment of the present invention, the reflection intensity corresponding to each target absorption point in the original near-infrared band data is extracted; a mapping relationship is constructed with wavelength as the key and reflection intensity as the value, such as a wavelength of 1550nm corresponding to a reflection intensity of 0.23, to generate reflection intensity distribution data consisting of key-value pairs.

[0045] Step 214: Binding the reflection intensity distribution data with the spatial position information of the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength; In this step, the spatial position information refers to the two-dimensional coordinates (row number, column number) of the pixel in the document image, which is used to locate the physical position of the non-text element.

[0046] In an embodiment of the present invention, each key-value pair in the reflection intensity distribution data is associated with the spatial coordinates of the near-infrared band data. Specifically, the pixel area corresponding to each wavelength is located through the spatial index matrix, and the third-dimensional coordinate information (row number, column number) is added to the reflection intensity distribution data to form a three-dimensional structure of near-infrared spectrum data, whose format is [row coordinate, column coordinate, wavelength-intensity mapping relationship].

[0047] The embodiment of the present invention effectively identifies weak characteristic peaks through neighborhood contrast absorption point detection; the precise matching of the material reflection characteristic library solves the pain point that traditional solutions cannot distinguish similar materials, thereby reducing the material misjudgment rate; the wavelength-space dual-dimensional binding ensures that the material properties and positions of the formula symbols correspond accurately, providing reliable input for subsequent cross-modal analysis.

[0048] For example, let's analyze a document containing formulas for toner printing on yellowing copperplate paper. Based on the aforementioned steps, near-infrared data (with a wavelength range of 700-2500 nm and a resolution of 5 nm) has been acquired. First, the system detects spectral absorption feature points within the near-infrared data where the reflection intensity is lower than that of adjacent wavelengths. Using a sliding window approach, the system detects a reflection intensity of 0.19 at a wavelength of 1550 nm, and the average intensity of its three adjacent points is 0.83. Since 0.19 ≤ 0.83 × 0.85 (i.e., 0.7055) is determined, this wavelength is marked as a spectral absorption feature point. Subsequently, based on a pre-defined material reflectance database, this spectral absorption feature point (1550 nm) is compared with a standard wavelength. The standard absorption peak for toner is 1552 nm ± 3 nm. The calculated result is |1550–1552| = 2 nm < 3 nm. This point is determined to meet the reflectance characteristics of the toner material and is designated as a target absorption point associated with the material's reflectance characteristics. Next, the reflection intensity value of 0.19 at the wavelength of 1550 nm, where the target absorption point is located, is extracted, generating the reflection intensity distribution data {1550:0.19}, which carries the wavelength-intensity mapping relationship. Finally, this reflection intensity distribution data is bound to the spatial location information in the near-infrared band data, locating the coordinates of the region where the formula symbol in the image is located at (120, 45)-(130, 55), and outputting the corresponding near-infrared spectrum data [rows 120-130, columns 45-55, {1550:0.19}]. This method has been proven to successfully identify formula symbols printed with toner.

[0049] The present invention provides a specific embodiment, step 103, dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating the spectral feature difference between the high-sensitivity band and the low-sensitivity band, specifically includes the following steps: Step 301: Counting the spectral intensity of each spectral channel in the multi-channel spectral data to generate a spectral intensity distribution curve corresponding to each spectral channel; In this step, spectral intensity refers to the energy value of light reflected from an object at a specific wavelength, ranging from 0 to 1 (0 means no reflection, 1 means full reflection), reflecting the optical response characteristics of the material. The spectral intensity distribution curve is a continuous function curve with wavelength (in nanometers) as the horizontal axis and average spectral intensity as the vertical axis. It is used to visualize the spectral energy distribution pattern.

[0050] In an embodiment of the present invention, wavelength-by-wavelength scanning is performed on each spectral channel of multi-channel spectral data (such as visible light red, green, blue channels, and near-infrared channel): the spectral intensity values ​​of all pixels at each wavelength position are accumulated and divided by the total number of pixels to generate the average spectral intensity at that wavelength position; a continuous curve is drawn with the wavelength as the horizontal axis and the average intensity as the vertical axis, and a spectral intensity distribution curve containing peak points and valley points is output.

[0051] Step 302: Marking the band intervals on the spectral intensity distribution curve where the spectral intensity change rate exceeds a preset spectral sensitivity threshold as high-sensitivity bands, and marking the band intervals where the spectral intensity change rate is lower than the preset spectral sensitivity threshold as low-sensitivity bands; In this step, the spectral intensity change rate refers to the amount of change in spectral intensity per unit wavelength interval, reflecting the dynamic sensitivity of the spectral response. The preset spectral sensitivity threshold refers to a pre-set critical value for the rate of change, which is used to distinguish between highly dynamic response areas and stable response areas. The highly sensitive band refers to the continuous wavelength range where the spectral intensity change rate exceeds the threshold, corresponding to areas of characteristic mutation such as the edges of graphic elements and the boundaries of formula symbols in the literature. The low-sensitivity band refers to the continuous wavelength range where the spectral intensity change rate is below the threshold, corresponding to the literature background or uniformly filled areas.

[0052] In an embodiment of the present invention, a first-order derivative of the spectral intensity distribution curve is calculated (the current wavelength intensity value minus the previous wavelength intensity value, divided by the wavelength interval) to obtain a spectral intensity change rate sequence; continuous wavelength intervals where the spectral intensity change rate exceeds a preset spectral sensitivity threshold (e.g., 0.25 intensity units / nanometer) are marked as high-sensitivity bands; and continuous wavelength intervals where the change rate is below the threshold are marked as low-sensitivity bands.

[0053] Step 303: extracting a first spectral intensity extreme point from the high-sensitivity band and a second spectral intensity extreme point from the low-sensitivity band, and calculating the absolute value of the intensity difference between the first spectral intensity extreme point and the second spectral intensity extreme point, and using the absolute value of the intensity difference as the spectral feature difference; In this step, the first spectral intensity extreme point refers to the local maximum point of the spectral intensity distribution curve within the high-sensitivity band, representing the intensity peak of the key feature area. The second spectral intensity extreme point refers to the local minimum point of the spectral intensity distribution curve within the low-sensitivity band, representing the intensity valley of the background reference area. The spectral feature difference refers to the absolute value of the difference between the intensity of the first extreme point and the second extreme point, quantifying the contrast difference between the key feature and the background.

[0054] In an embodiment of the present invention, the local maximum point of the spectral intensity distribution curve is located within the high-sensitivity band interval as the first spectral intensity extreme point; the local minimum point is located within the low-sensitivity band interval as the second spectral intensity extreme point; the intensity value of the first spectral intensity extreme point is subtracted from the intensity value of the second spectral intensity extreme point, and the absolute value is taken to obtain the absolute value of the intensity difference, which is used as the spectral feature difference.

[0055] The embodiment of the present invention effectively distinguishes the edges of chart elements from the paper background through band division driven by the rate of change; the quantification of extreme point differences significantly enhances the boundary characteristics of formula symbols, thereby improving the recognition rate of key features of fuzzy documents; and the dynamic threshold judgment adaptively processes documents of different printing qualities, solving the pain point of poor generalization of traditional fixed threshold solutions.

[0056] The present invention provides a specific embodiment, step 104, performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data, specifically comprising the following steps: Step 401: Calculating the intensity ratio of the spatial continuity attribute and the frequency domain distribution attribute in the spectral feature difference; In this step, the spatial continuity attribute refers to the component of the spectral feature difference that reflects the stability of the spatial structure of the chart elements. It is calculated based on the intensity of the chart edge extreme points within the highly sensitive band. The frequency domain distribution attribute refers to the component of the spectral feature difference that reflects the change in material reflectance of the formula symbol. It is calculated based on the intensity of the extreme points of the material absorption feature points. The intensity ratio value is the quotient of the spatial continuity attribute intensity divided by the frequency domain distribution attribute intensity, which is used to determine the bit width resource allocation weight.

[0057] In this embodiment of the present invention, the intensity value of the spatial continuity attribute (the intensity of the extreme points at the edges of the chart elements in the high-sensitivity band) and the intensity value of the frequency domain distribution attribute (the intensity of the extreme points at the absorption peaks of the material reflection characteristics) are extracted from the spectral feature difference. The intensity value of the spatial continuity attribute is divided by the intensity value of the frequency domain distribution attribute to generate an intensity ratio value that represents the energy proportion relationship between the two.

[0058] Step 402: Determine a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute according to the intensity ratio value; In this step, the first target bit width value refers to the number of bits allocated to the spatial continuity attribute. Its value is positively correlated with the intensity ratio value and is used to control the encoding accuracy of the chart structure. The second target bit width value refers to the number of bits allocated to the frequency domain distribution attribute. Its value is negatively correlated with the intensity ratio value and is used to control the encoding accuracy of the material characteristics.

[0059] In an embodiment of the present invention, the intensity ratio value is used as the weight ratio, and the weight ratio is multiplied by a preset bit width base to generate an initial bit width value for spatial continuity; the bit width base is adjusted according to the fluctuation amplitude range of the frequency domain distribution attribute to generate an initial bit width value for the frequency domain distribution; the initial bit width value for the frequency domain distribution is corrected according to the intensity stability between different wavelength positions to generate a second target bit width value; the initial bit width value for spatial continuity is corrected according to the intensity correlation between adjacent pixel positions to generate a first target bit width value.

[0060] Step 403: Calculating the spectral intensity variation between adjacent pixel positions in the color distribution feature to generate a set of spatial continuity variation; In this step, the spectral intensity variation refers to the absolute value of the intensity difference between the same pixel position in adjacent spatial directions (up / down / left / right), quantifying the degree of local structural abruptness. The spatial continuity variation set is a two-dimensional matrix that stores the intensity differences between all pixel positions and their four neighbors, with dimensions equal to the document image size.

[0061] In an embodiment of the present invention, each pixel position in the color distribution feature is traversed, and the spectral intensity difference between it and the upper, lower, left, and right adjacent pixels is calculated (the current pixel intensity minus the adjacent pixel intensity); the absolute values ​​of the differences between all adjacent positions are stored as a two-dimensional matrix according to the pixel coordinates to generate a set of spatial continuity change quantities.

[0062] Step 404: Calculate the intensity fluctuation amplitudes at different wavelength positions in the material reflection characteristics to generate a frequency domain distribution fluctuation quantity set; In this step, the intensity fluctuation amplitude refers to the difference between the maximum and minimum reflection intensities at the same wavelength in the frequency domain, representing the material's reflective stability. The frequency domain distribution fluctuation quantity set is a one-dimensional array that stores the fluctuation amplitude at each position in a wavelength sequence, with a length equal to the number of wavelength sampling points.

[0063] In an embodiment of the present invention, each wavelength position in the material reflection characteristics is traversed, and the intensity difference between its maximum reflection intensity and minimum reflection intensity is calculated as the intensity fluctuation amplitude; the intensity fluctuation amplitude values ​​at different wavelength positions (intervals of 10nm) are stored as a one-dimensional array according to the wavelength sequence to generate a set of frequency domain distribution fluctuation quantities.

[0064] Step 405: encoding the set of spatial continuity variation values ​​according to the first target bit width value to generate spatial continuity encoding data; In this step, the spatial continuity encoding data refers to the graph structure feature data compressed by bit width constraint, and includes a binary stream of pixel position-intensity change mapping.

[0065] In an embodiment of the present invention, an adaptive arithmetic encoder is used to process a set of spatial continuity variation quantities: the encoding accuracy is determined according to the first target bit width value, such as the first target bit width value 8 corresponds to an integer range of 0-255, the intensity difference between adjacent pixels is quantized into an integer and then compressed and encoded to obtain spatial continuity encoding data in binary format.

[0066] Step 406: encoding the frequency domain distribution fluctuation quantity set according to the second target bit width value to generate frequency domain distribution coded data; In this step, the frequency domain distribution coded data refers to the material feature data compressed by the bit width constraint, and includes a binary stream of wavelength-fluctuation amplitude mapping.

[0067] In an embodiment of the present invention, differential pulse coding is performed on a set of frequency domain distributed fluctuation quantities: the fluctuation amplitude quantization order is limited according to the second target bit width value, such as the second target bit width value 4 corresponds to 16 levels, the fluctuation amplitude sequence is differentially calculated and encoded to obtain frequency domain distributed coded data.

[0068] The embodiments of the present invention dynamically allocate resources through intensity ratio values ​​through attribute-aware bit width allocation (such as allocating high bits to chart structures and low bits to material features), thereby improving the retention rate of key information; spatial neighborhood differential coding effectively captures the boundary mutation characteristics of chart elements (such as histogram outlines), solving the problem of structural distortion in fuzzy documents; frequency domain fluctuation quantization compression maintains the material reflection characteristics of formula symbols, avoiding traditional solutions from misjudging the material of complex formulas.

[0069] The present invention provides a specific embodiment, step 402, determining a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute based on the intensity ratio value, specifically comprising the following steps: Step 411: using the intensity ratio value as a weight ratio, performing a multiplication operation on the weight ratio and a preset bit width base to generate an initial bit width value of spatial continuity; In this step, the weight ratio refers to the weighted contribution of the two features to the spectral difference. The preset bit width cardinality refers to the initial bit width baseline value set based on system resources and is used to calculate the initial bit width allocation for each attribute. The initial spatial continuity bit width value refers to the initial number of bits for the spatial attribute calculated based on the weight ratio and the bit width cardinality, without considering local correlation optimization.

[0070] In an embodiment of the present invention, a multiplication operation is performed on the weight ratio and the preset bit width cardinality. Specifically: when the weight ratio is greater than 1, the initial bit width value of spatial continuity = bit width cardinality × weight ratio; when the weight ratio is less than 1, the initial bit width value of spatial continuity = bit width cardinality × (1 ÷ weight ratio), ensuring that spatial features are prioritized.

[0071] Step 412: scaling the bit width base according to the fluctuation range of the frequency domain distribution attribute to generate an initial bit width value of the frequency domain distribution; In this step, the fluctuation amplitude range refers to the difference between the maximum and minimum reflection intensities of the frequency domain distribution attribute at all wavelengths, reflecting the overall fluctuation of the material's reflectance. The initial bit width value of the frequency domain distribution refers to the initial number of bits of the frequency domain attribute obtained by scaling the bit width base based on the fluctuation amplitude range.

[0072] In an embodiment of the present invention, the difference between the maximum reflection intensity and the minimum reflection intensity of the frequency domain distribution attribute at all wavelength positions is calculated to obtain the fluctuation amplitude range; the initial bit width value of the frequency domain distribution = the bit width base × the fluctuation amplitude range scaling factor, where the fluctuation amplitude range scaling factor = 1 - the fluctuation amplitude range ÷ the maximum possible fluctuation value.

[0073] Step 413: Modify the initial bit width value of the frequency domain distribution according to the intensity stability between different wavelength positions to generate a second target bit width value; In this step, intensity stability refers to the variance of reflection intensities at different wavelengths, quantifying the wavelength-dimensional consistency of the material's reflections. Intensity correlation, calculated over a 3×3 pixel neighborhood (covariance divided by the product of standard deviations), measures the color continuity of local regions within the chart element.

[0074] In an embodiment of the present invention, the variance value of the reflection intensity at adjacent wavelength positions is calculated as the intensity stability, that is, the intensity stability = the sum of the squares of the differences between the reflection intensity at each wavelength position and the mean divided by the number of wavelengths; when the intensity stability exceeds a preset threshold, the second target bit width value = the initial bit width value of the frequency domain distribution × 1.2 is used to enhance the stability characterization; otherwise, the original value is maintained as the second target bit width value.

[0075] Step 414: Modify the initial bit width value of spatial continuity according to the intensity correlation between adjacent pixel positions to generate a first target bit width value; In this embodiment of the present invention, a 3×3 neighborhood of pixels is selected, and the Pearson correlation coefficient between the intensity of the central pixel and the surrounding pixels is calculated as the degree of intensity correlation. When the Pearson correlation coefficient is greater than 0.7, it indicates that the intensity of the neighborhood pixels is highly correlated with the central pixel, corresponding to smooth areas in the document image (such as chart backgrounds and evenly filled color blocks). Such areas have strong spatial continuity and high information redundancy, and the data volume can be compressed by reducing the bit width. In this case, the first target bit width value = the initial bit width value for spatial continuity × 0.8. When the Pearson correlation coefficient is ≤0.7, it indicates that the intensity difference between the neighborhood and the central pixel is significant, corresponding to edges, textures, or boundaries of chart elements in the document image (such as the outline of formula symbols or lines in line graphs). Such areas have weak spatial continuity and contain key structural information, and the bit width needs to be increased to preserve details. In this case, the first target bit width value = the initial bit width value for spatial continuity × 1.1.

[0076] The embodiment of the present invention improves the bit width resources in dense areas of the chart through dynamic allocation of weight ratios, thereby solving the problem of blurred edges in high-density curve charts; the fluctuation amplitude range is scaled to automatically adjust the frequency domain bit width according to the reflection stability of the material; the intensity stability correction avoids over-allocation of stable materials, thereby improving the frequency domain data compression rate; the intensity correlation correction strengthens the bit width allocation of broken chart elements, thereby improving the restoration of key structures.

[0077] The present invention provides a specific embodiment, step 105, parsing the spatial continuity coded data and the frequency domain distribution coded data by a large model to generate a structured text description, specifically comprising the following steps: Step 501: Parsing the element position coordinate sequence in the spatial continuity coded data to generate a spatial position description set of the chart element, wherein the spatial position description set includes a first identifier and a first document position tag; In this step, the element position coordinate sequence refers to the set of geometric vertex coordinates of the chart elements extracted from the spatially encoded data (e.g., the line chart point sequence [x1, y1, x2, y2, ...]), reflecting the spatial topology of the chart. The spatial position description set refers to a key-value pair that stores the chart element identifier and its document location label, used to locate the physical location of the chart. The first identifier refers to the unique alphanumeric code assigned to the chart element, used for cross-module references. The first document location label describes the spatial location of the chart element in the document, in the following format: page number, geometry type (coordinates). For example, "Page 2, center [50, 60, 15]" indicates page 2, a circle with a center of (50, 60) and a radius of 15.

[0078] In an embodiment of the present invention, the spatial continuity coded data is decoded by the geometric parsing module of the large model to extract the stored chart element vertex coordinate sequence (such as the coordinates of the line chart data points); a unique first identifier is assigned to each chart element, such as icon 1, and the coordinates of the document page area it occupies are marked as the first document location label, such as the rectangular area [120, 50, 180, 80] on page 3, and finally a spatial location description set containing the key-value pairs of {first identifier: first document location label} is output.

[0079] Step 502: parsing the correspondence between wavelength and reflection intensity in the frequency domain distribution coded data to generate a material property description set of formula symbols, wherein the material property description set includes a second identifier and a second document location tag; In this step, the correspondence between wavelength and reflection intensity refers to material characteristic data parsed from frequency-domain encoded data, such as {1200nm: 0.35, 1550nm: 0.18}, which characterizes the reflectance spectrum characteristics of the formula symbol. The material characteristic description set refers to a key-value pair storing the formula symbol identifier and its material characteristics. The second identifier is a unique code assigned to the formula symbol, including mathematical semantic tags. The second document location tag is the coordinates of the center point of the formula symbol, in the format: page number, point position (x, y).

[0080] In an embodiment of the present invention, the frequency domain analysis module of the large model is used to decode the frequency domain distribution coded data and extract the wavelength-reflection intensity mapping relationship, such as the reflection intensity corresponding to 1550nm is 0.23; a second identifier is assigned to each formula symbol, such as the formula ∫, and its center point coordinates are marked as the second document location label, such as the point coordinates [88,72] on page 5, and a material characteristic description set containing the key-value pair {second identifier: second document location label} is output.

[0081] Step 503: Based on the consistency of the document location tags, cross-modal matching is performed on the first identifier and the second identifier to generate an element symbol association mapping table; In this step, document location tag consistency refers to the matching condition that the spatial distance between two document location tags is ≤5 pixels, ensuring the physical proximity of the associated elements. The element symbol association mapping table refers to a table that records the mapping relationship between diagrams and formula identifiers, used for cross-modal association.

[0082] In an embodiment of the present invention, the spatial position description set and the material characteristic description set are coordinate matched according to the document position tag. Specifically: when the spatial distance between the position tags corresponding to the two identifiers is less than a preset threshold (such as 5 pixels), it is determined that the document position tag consistency is satisfied; a mapping relationship between the first identifier and the second identifier is established, and an element symbol association mapping table is generated.

[0083] Step 504: According to the element symbol association mapping table, retrieve the chart element position description from the spatial position description set, and retrieve the formula symbol material description from the material property description set; In this step, the chart element position description refers to the natural language description of the spatial position of the chart element, such as a curve spanning the left half of the page and a horizontal axis range of 0-100. The formula symbol material description refers to the physical properties of the formula symbol, such as the integral symbol ∫ uses coated paper ink with a characteristic absorption peak at 1200nm.

[0084] In an embodiment of the present invention, according to the association relationship of the element symbol association mapping table, the chart element position description is retrieved from the spatial position description set; the formula symbol material description is retrieved from the material property description set, such as the formula ∑: toner printing, 1550nm absorption intensity 0.18.

[0085] Step 505: combining the diagram element location descriptions and the formula symbol material descriptions associated with the same document location into a composite text paragraph containing the diagram element topology description and the formula symbol attribute description, integrating the composite text paragraphs corresponding to all document locations to generate a structured text description; In this step, the chart element topology description refers to a detailed description of the chart element. For example, a bar chart contains three data sets with a height ratio of 2:5:3. The formula symbol attribute description refers to the material and semantic description of the formula symbol. For example, the symbol α represents an angle variable and is printed on coated paper. The composite text paragraph is a related description paragraph that integrates the chart topology and formula attributes. The format is: In [position]: Chart description...; Related formula description... The structured text description is a sequence of composite text paragraphs sorted by document location, forming a complete report.

[0086] In an embodiment of the present invention, for each chart element and formula symbol group associated through an element symbol association mapping table (based on document location label consistency matching, such as spatial distance ≤ 5 pixels), the chart element location description is converted into a chart element topology description including spatial location (such as the upper right quadrant of page 4), geometric type, and data relationship (such as a bar chart containing 5 data columns with a height ratio of 1:3:2:4:5). The formula symbol material description is converted into a formula symbol attribute description covering printing material (such as laser toner), spectral characteristics (1550nm reflection peak), and mathematical semantics (such as ∑ represents data accumulation operation, and then follows the format: on [page 4, coordinates (88,72), surrounding area]: the bar chart shows the data trend from 2020 to 2024; the associated formula ∑ is printed using toner and is used to calculate the sum of growth rates). The chart topology and formula attributes of the same document location are merged into a composite text paragraph. Finally, all composite text paragraphs are sorted by page order and spatial location to generate a structured text description, thereby achieving semantic integration of non-text elements based on location association and solving the problem of semantic fragmentation of cross-modal elements in traditional solutions.

[0087] The embodiment of the present invention improves the accuracy of the association recognition between formula symbols and adjacent bar graphs through cross-modal position matching, solving the semantic fragmentation problem of traditional solutions; material-enhanced semantic analysis combined with printing properties improves the recognition rate of formula symbols, especially improves the discrimination of handwritten symbols; structured paragraph generation improves the completeness of the description of the association between charts and formulas in academic literature to meet the journal abstract standards.

[0088] Figure 2 The present invention provides a structural diagram of a large model-based literature report generation and evaluation system. Figure 2 As shown, the system includes: An acquisition module 21 is configured to acquire a document image dataset containing non-text elements, wherein the non-text elements include graphic elements and formula symbols; A separation module 22 is used to separate different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics; A calculation module 23 is configured to divide the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculate a spectral feature difference between the high-sensitivity band and the low-sensitivity band; an allocation module 24 for performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; A generating module 25, configured to parse the spatial continuity coded data and the frequency domain distribution coded data through a large model to generate a structured text description; The verification module 26 is used to verify the information coverage of the structured text description relative to the document image dataset to generate an evaluation result.

[0089] Figure 2 The literature report generation and evaluation system based on the large model can be executed Figure 1 The implementation principle and technical effects of the large-scale model-based literature report generation and evaluation method described in the illustrated embodiment will not be elaborated on here. The specific manner in which each module and unit performs operations in the large-scale model-based literature report generation and evaluation system in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.

[0090] In one possible design, Figure 2 The literature report generation and evaluation system based on a large model of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32; The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0091] The processing component 32 is used for the above Figure 1 The embodiment provides a method for evaluating the generation of literature reports based on a large model.

[0092] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0093] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0094] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0095] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0096] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0097] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0098] The embodiment of the present invention further provides a computer storage medium storing a computer program, which can achieve the above-mentioned Figure 1 The illustrated embodiment is a method for evaluating the generation of literature reports based on a large model.

[0099] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0101] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating and evaluating literature reports based on a large model, characterized in that: include: Acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols; Separating different spectral band data of the document image dataset to generate multi-channel spectral data including color distribution characteristics and material reflection characteristics; Dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating a spectral feature difference between the high-sensitivity band and the low-sensitivity band; Performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; Parsing the spatial continuity coded data and the frequency domain distribution coded data by a large model to generate a structured text description; The information coverage of the structured text description relative to the document image dataset is verified to generate an evaluation result.

2. The method according to claim 1, characterized in that Performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data, including: Calculating the intensity ratio of the spatial continuity attribute to the frequency domain distribution attribute in the spectral feature difference; Determining, according to the intensity ratio value, a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute; Calculating the spectral intensity variation between adjacent pixel positions in the color distribution feature to generate a set of spatial continuity variation; Calculating the intensity fluctuation amplitudes at different wavelength positions in the material reflection characteristics to generate a frequency domain distribution fluctuation quantity set; encoding the set of spatial continuity variation values ​​according to the first target bit width value to generate spatial continuity coded data; According to the second target bit width value, the frequency domain distribution fluctuation quantity set is coded to generate frequency domain distribution coded data.

3. The method according to claim 2, characterized in that Determining, according to the intensity ratio value, a first target bit width value corresponding to the spatial continuity attribute and a second target bit width value corresponding to the frequency domain distribution attribute, including: Taking the intensity ratio value as a weight ratio, performing a multiplication operation on the weight ratio and a preset bit width base to generate an initial bit width value of spatial continuity; According to the fluctuation amplitude range of the frequency domain distribution attribute, the bit width cardinality is proportionally adjusted to generate an initial bit width value of the frequency domain distribution; According to the intensity stability between different wavelength positions, the initial bit width value of the frequency domain distribution is corrected to generate a second target bit width value; The initial bit width value of spatial continuity is modified according to the intensity correlation degree between adjacent pixel positions to generate a first target bit width value.

4. The method according to claim 1, wherein Separate the different spectral bands of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflectance characteristics, including: According to the spectral band range, the document image data set is divided into visible light band data and near infrared band data; Identifying a spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectrum data carrying three primary color intensity information; Detecting spectral absorption feature points related to the material reflection characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength; The visible light spectrum data and the near-infrared spectrum data are channel-joined to form multi-channel spectrum data including color distribution characteristics and material reflection characteristics.

5. The method according to claim 4, characterized in that Detecting spectral absorption feature points related to the material reflectance characteristics in the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength, including: Detecting spectral absorption characteristic points in the near-infrared band data at which the reflection intensity is lower than that at adjacent wavelength positions; According to a preset material reflection characteristic library, the wavelength position of the spectral absorption characteristic point is matched to screen out the target absorption point associated with the material reflection characteristic; Extracting the reflection intensity value at the wavelength position of the target absorption point to generate reflection intensity distribution data carrying a wavelength-intensity mapping relationship; The reflection intensity distribution data is bound to the spatial position information of the near-infrared band data to generate near-infrared spectrum data carrying reflection intensity information of a specific wavelength.

6. The method according to claim 1, characterized in that Dividing the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculating the spectral feature difference between the high-sensitivity band and the low-sensitivity band, including: Counting the spectral intensity of each spectral channel in the multi-channel spectral data to generate a spectral intensity distribution curve corresponding to each spectral channel; Marking the band interval where the spectral intensity change rate on the spectral intensity distribution curve exceeds a preset spectral sensitivity threshold as a high-sensitivity band, and marking the band interval where the spectral intensity change rate is lower than the preset spectral sensitivity threshold as a low-sensitivity band; A first spectral intensity extreme point is extracted from the high-sensitivity band, and a second spectral intensity extreme point is extracted from the low-sensitivity band. The absolute value of the intensity difference between the first spectral intensity extreme point and the second spectral intensity extreme point is calculated, and the absolute value of the intensity difference is used as the spectral feature difference amount.

7. The method according to claim 1, characterized in that Parsing the spatial continuity coded data and the frequency domain distribution coded data by a large model to generate a structured text description includes: Parsing the element position coordinate sequence in the spatial continuity coded data to generate a spatial position description set of the chart element, the spatial position description set including a first identifier and a first document position tag; parsing the correspondence between the wavelength and the reflection intensity in the frequency domain distribution coded data to generate a material property description set of formula symbols, wherein the material property description set includes a second identifier and a second document location tag; Based on the consistency of the document location tags, cross-modally matching the first identifier with the second identifier to generate an element symbol association mapping table; According to the element symbol association mapping table, the chart element position description is retrieved from the spatial position description set, and the formula symbol material description is retrieved from the material property description set; The chart element position descriptions and formula symbol material descriptions associated with the same document position are combined into a compound text paragraph containing a chart element topology description and a formula symbol attribute description. The compound text paragraphs corresponding to all document positions are integrated to generate a structured text description.

8. A large-scale model-based literature report generation and evaluation system, characterized in that: include: An acquisition module, configured to acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols; A separation module is used to separate different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics; a calculation module, configured to divide the multi-channel spectral data into a high-sensitivity band and a low-sensitivity band according to a preset spectral sensitivity threshold, and calculate a spectral feature difference between the high-sensitivity band and the low-sensitivity band; an allocation module, configured to perform dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data; A generating module, configured to parse the spatial continuity coded data and the frequency domain distribution coded data through a large model to generate a structured text description; A verification module is used to verify the information coverage of the structured text description relative to the document image dataset to generate an evaluation result.

9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a large model-based literature report generation and evaluation method as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the method for generating and evaluating a literature report based on a large model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Multi-modal literature data extraction method and device and medium

    CN119311880A

  • Open set cross-domain hyperspectral classification method and system based on multiple modes and optimal transmission

    CN119832325A

  • Archive data classification method and device based on artificial intelligence and storage medium

    CN119917662A

  • Method for processing cultural relic data by using three-dimensional scanner

    CN120125945A

  • Image processing device, image processing method, and image sensor

    JP2019215676A