A large model-based literature report generation evaluation method and system

By acquiring multi-channel spectral data of document images and generating structured text descriptions, the problem of color shift and loss of detail caused by material differences and printing quality in existing technologies has been solved, enabling accurate identification and evaluation of high-density charts and complex mathematical formulas.

CN120724992BActive Publication Date: 2025-11-18LUSTER LIGHTWAVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511213131.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-18
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture color shifts and detail loss caused by material differences or printing quality when processing image data with complex spectral characteristics. This results in insufficient accuracy in recognizing document image content, especially in scenarios with high-density charts or complex mathematical formulas, where key information may be missed or semantic misinterpretations may occur.

Method used

By acquiring multi-channel spectral data of document images, dividing them into high-sensitivity and low-sensitivity bands, calculating the spectral feature differences, and performing dynamic bit-width allocation, spatial continuity and frequency domain distribution coded data are generated. These coded data are then analyzed using a large model to generate structured text descriptions.

Benefits of technology

It improves the accuracy of identifying non-textual elements, solves the problems of missing key information and semantic misinterpretation, enhances the logical coherence and information integrity of document analysis, adapts to complex spectral scenarios, and provides reliable evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724992B_ABST
    Figure CN120724992B_ABST
Patent Text Reader

Abstract

The application provides a literature report generation evaluation method and system based on a large model, wherein a literature image dataset containing non-text elements is obtained, the non-text elements including chart elements and formula symbols; different spectral band data of the literature image dataset is separated to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics, high-sensitivity bands and low-sensitivity bands of the multi-channel spectral data are divided according to a spectral sensitivity threshold, a spectral feature difference between the two bands is calculated, and bit width allocation is performed to generate spatial continuity encoding data and frequency domain distribution encoding data; the two types of encoding data are analyzed by a large model to generate structured text description, and the information coverage of the description relative to the literature image dataset is verified to generate an evaluation result; the application improves the recognition accuracy and structured description capability of non-text elements in the literature image, and meets the demand for automatic analysis and evaluation of high-quality multimedia literature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for generating and evaluating literature reports based on large models. Background Technology

[0002] In general multimedia literature analysis scenarios, with the increasing digitization and visualization of academic resources, how to efficiently and accurately process literature containing non-text elements such as charts, images, and formulas has become a key requirement for information extraction and knowledge summarization.

[0003] Currently, there are existing solutions for document content understanding based on multimodal fusion. These solutions combine convolutional neural networks and the Transformer architecture. First, they detect and segment image regions in the document, and then use a pre-trained vision-language model to generate descriptive text for the image content. Furthermore, they evaluate the semantic consistency between the generated text and the original image content through text similarity calculation and keyword matching. However, existing solutions have significant limitations when dealing with image data with complex spectral characteristics. Because they mainly rely on conventional visual information from the red, green, and blue channels, they struggle to effectively capture color shifts and detail loss caused by material differences or printing quality in document images, thus affecting the accuracy of content recognition. The generated text descriptions also have considerable room for improvement in terms of structure and information completeness, especially in scenarios involving high-density charts or complex mathematical formulas, where key information may be missed or semantic misinterpretations may occur, limiting their practical application in high-quality document analysis tasks. Summary of the Invention

[0004] This invention provides a method and system for generating and evaluating literature reports based on a large model, which addresses the shortcomings in the accuracy of document image content recognition in existing technologies; and the potential for omissions of key information or semantic misinterpretations in scenarios with high-density charts or complex mathematical formulas.

[0005] In a first aspect, the present invention provides a method for generating and evaluating literature reports based on a large model, comprising:

[0006] Obtain a dataset of document images containing non-text elements, including chart elements and formula symbols;

[0007] The different spectral band data of the document image dataset are separated to generate multi-channel spectral data containing color distribution features and material reflection features;

[0008] The multi-channel spectral data is divided into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands is calculated.

[0009] Dynamic bit width allocation is performed on the spectral feature differences to generate spatially continuous coded data and frequency domain distributed coded data;

[0010] The spatial continuity coded data and the frequency domain distribution coded data are analyzed using a large model to generate a structured text description;

[0011] Verify the information coverage of the structured text description relative to the document image dataset to generate evaluation results.

[0012] Optionally, the different spectral band data of the document image dataset are separated to generate multi-channel spectral data containing color distribution features and material reflectance features, including:

[0013] Based on the spectral band range, the document image dataset is divided into visible light band data and near-infrared band data;

[0014] Identify the spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectral data carrying three primary color intensity information;

[0015] Detect spectral absorption feature points related to the material's reflectivity in the near-infrared band data to generate near-infrared spectral data carrying reflectivity information at a specific wavelength;

[0016] The visible light spectral data and the near-infrared spectral data are spliced ​​together to form multi-channel spectral data that includes color distribution characteristics and material reflection characteristics.

[0017] Optionally, detecting spectral absorption feature points in the near-infrared band data that are related to the material's reflectivity to generate near-infrared spectral data carrying reflectivity information at a specific wavelength includes:

[0018] Detect spectral absorption feature points in the near-infrared band data where the reflection intensity is lower than that of adjacent wavelengths;

[0019] Based on a preset library of material reflection characteristics, the wavelength positions of the spectral absorption feature points are matched to filter out target absorption points associated with the material's reflection characteristics.

[0020] Extract the reflection intensity value at the wavelength position of the target absorption point to generate reflection intensity distribution data carrying wavelength intensity mapping relationship;

[0021] The reflection intensity distribution data is bound to the spatial location information of the near-infrared band data to generate near-infrared spectral data carrying reflection intensity information of a specific wavelength.

[0022] Optionally, the multi-channel spectral data is divided into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands is calculated, including:

[0023] The spectral intensity of each spectral channel in the multi-channel spectral data is statistically analyzed to generate the spectral intensity distribution curve corresponding to each spectral channel;

[0024] The bands on the spectral intensity distribution curve whose rate of change of spectral intensity exceeds a preset spectral sensitivity threshold are marked as high-sensitivity bands, while the bands whose rate of change of spectral intensity is lower than the preset spectral sensitivity threshold are marked as low-sensitivity bands.

[0025] The first spectral intensity extreme point is extracted from the high-sensitivity band, and the second spectral intensity extreme point is extracted from the low-sensitivity band. The absolute value of the intensity difference between the first spectral intensity extreme point and the second spectral intensity extreme point is calculated, and the absolute value of the intensity difference is used as the spectral feature difference quantity.

[0026] Optionally, dynamic bit-width allocation is performed on the spectral feature differences to generate spatially continuous coded data and frequency-domain distributed coded data, including:

[0027] Calculate the intensity ratio of the spatial continuity attribute to the frequency domain distribution attribute in the spectral feature difference;

[0028] Based on the intensity ratio value, determine the first target bit width value corresponding to the spatial continuity attribute and the second target bit width value corresponding to the frequency domain distribution attribute;

[0029] Calculate the spectral intensity variation between adjacent pixel positions in the color distribution features to generate a set of spatially continuous variation quantities;

[0030] Calculate the intensity fluctuation amplitude at different wavelength positions in the material's reflection characteristics to generate a set of frequency domain distribution fluctuations;

[0031] Based on the first target bit width value, the set of spatial continuity changes is encoded to generate spatial continuity encoded data;

[0032] Based on the second target bit width value, the frequency domain distribution fluctuation set is encoded to generate frequency domain distribution encoded data.

[0033] Optionally, determining the first target bit width value corresponding to the spatial continuity attribute and the second target bit width value corresponding to the frequency domain distribution attribute based on the intensity ratio value includes:

[0034] The intensity ratio is used as the weight ratio, and the weight ratio is multiplied by the preset bit width base to generate the initial bit width value for spatial continuity.

[0035] Based on the fluctuation range of the frequency domain distribution attribute, the bit width base is proportionally scaled and adjusted to generate an initial bit width value for the frequency domain distribution.

[0036] Based on the intensity stability between different wavelength positions, the initial bit width value of the frequency domain distribution is corrected to generate a second target bit width value;

[0037] Based on the strength correlation between adjacent pixel positions, the initial bit width value of spatial continuity is corrected to generate a first target bit width value.

[0038] Secondly, this invention provides a literature report generation and evaluation system based on a large model, comprising:

[0039] The acquisition module is used to acquire a dataset of document images containing non-text elements, including chart elements and formula symbols;

[0040] The separation module is used to separate the different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution features and material reflection features;

[0041] The calculation module is used to divide the multi-channel spectral data into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and to calculate the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands.

[0042] The allocation module is used to perform dynamic bit width allocation on the spectral feature difference to generate spatially continuous coded data and frequency domain distributed coded data.

[0043] The generation module is used to parse the spatial continuity coded data and the frequency domain distribution coded data through a large model to generate a structured text description;

[0044] The verification module is used to verify the information coverage of the structured text description relative to the document image dataset in order to generate evaluation results.

[0045] Thirdly, the present invention provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement a literature report generation and evaluation method based on a large model as described in the first aspect above.

[0046] Fourthly, the present invention provides a computer storage medium storing a computer program, which, when executed by a computer, implements a literature report generation and evaluation method based on a large model as described in the first aspect.

[0047] In this invention, a document image dataset containing non-text elements, including chart elements and formula symbols, is acquired. Different spectral bands of the document image dataset are separated to generate multi-channel spectral data containing color distribution features and material reflectance features. The multi-channel spectral data is divided into high-sensitivity and low-sensitivity bands according to a preset spectral sensitivity threshold, and the spectral feature difference between the high-sensitivity and low-sensitivity bands is calculated. Dynamic bit-width allocation is performed on the spectral feature difference to generate spatial continuity encoded data and frequency domain distribution encoded data. The spatial continuity encoded data and the frequency domain distribution encoded data are analyzed using a large model to generate a structured text description. The information coverage of the structured text description relative to the document image dataset is verified to generate an evaluation result. The technical solution provided by this invention solves the problem of information loss caused by insufficient separation of non-text elements in existing solutions, avoiding the over-reliance on text regions in traditional methods. It overcomes the limitations of existing solutions that rely solely on the red, green, and blue channels, enabling the capture of details lost in document images due to material differences (such as paper reflectivity and ink density) or printing quality issues, thus improving the accuracy of non-text element recognition and providing richer spectral information for subsequent feature extraction. Furthermore, it effectively filters out spectral information crucial for content recognition, solving the problem of feature redundancy or omission caused by the failure to distinguish the importance of bands in existing solutions, providing accurate feature input for subsequent dynamic coding, and improving the model's performance in complex spectral scenarios. Adaptability: It addresses the insufficient structured description capability caused by the single encoding method in existing schemes, enabling it to adapt to the multi-scale feature requirements of high-density charts or complex mathematical formulas, and avoiding the loss of key information (such as chart boundaries and formula symbol hierarchy) during the encoding process. It also solves the semantic misreading or information fragmentation problems caused by the reliance on shallow visual-language models in existing schemes, improving the logical coherence and information integrity of text descriptions, especially showing stronger robustness when dealing with complex visual elements. Furthermore, it overcomes the limitations of existing schemes that rely solely on text similarity matching (such as ignoring image details or semantic deviations), ensuring that the evaluation results are closer to the actual content of the literature, and providing a reliable feedback mechanism for high-quality literature analysis tasks. Further, based on the intensity ratio of spatial continuity attributes and frequency domain distribution attributes in the spectral feature difference, a target bit width value is dynamically allocated; a set of spatial continuity variations is generated by calculating the spectral intensity changes at adjacent pixel positions, and a set of frequency domain distribution fluctuations is generated by calculating the intensity fluctuation amplitude at different wavelength positions; finally, the two sets are independently encoded according to the target bit width value, and spatial continuity encoded data and frequency domain distribution encoded data are output separately.By employing an attribute-aware dynamic bit-width allocation mechanism, the boundary intensity transition characteristics of chart elements are preserved, thus resolving structural distortion caused by differences in print quality. The material reflection fluctuation characteristics of formula symbols are maintained to avoid blurring of key features in complex formula areas. Encoding resources are dynamically allocated according to the proportion of spectral feature differences, achieving a breakthrough balance between compression ratio and information integrity.

[0048] These or other aspects of the invention will become more apparent from the following description of the embodiments. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 A flowchart of a literature report generation and evaluation method based on a large model provided by the present invention is shown;

[0051] Figure 2 A schematic diagram of the structure of a literature report generation and evaluation system based on a large model provided by the present invention is shown;

[0052] Figure 3 A schematic diagram of the structure of a computing device provided by the present invention is shown. Detailed Implementation

[0053] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0054] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] In general multimedia document analysis scenarios, traditional multimodal fusion schemes, when processing document images containing non-textual elements such as charts and formulas, are limited to using only red, green, and blue visual information. This makes it difficult to address color shifts and detail loss caused by factors such as material and printing quality, thus affecting the accuracy of content recognition and the completeness of the generated text. Furthermore, existing schemes lack fine-grained modeling of spatial continuity and frequency domain distribution in feature encoding, resulting in insufficient structured description capabilities, especially in high-density charts or complex mathematical formula scenarios, which can easily lead to the omission of key information or semantic misinterpretation. To address these shortcomings, this invention introduces a spectral data analysis mechanism, expanding the document image from a conventional color space to a multi-channel spectral feature space. Further, by combining a dynamic bit-width allocation strategy, highly expressive encoded data on spatial continuity and frequency domain distribution are extracted. This data is used as input to a large model, improving its understanding of complex visual elements and its structured expression capabilities. Ultimately, this allows for effective evaluation of the information coverage of the generated text relative to the original image content, enhancing the quality control of automatic document parsing. Figure 1 A flowchart of a literature report generation and evaluation method based on a large model is provided as an embodiment of the present invention, such as... Figure 1 As shown, the method includes:

[0057] Step 101: Obtain a dataset of document images containing non-text elements, including chart elements and formula symbols;

[0058] In this step, non-text elements refer to visual objects in the literature that cannot be directly represented by characters, including chart elements reflecting data relationships and formula symbols expressing mathematical logic. The literature image dataset refers to a collection of digitized literature images acquired through spectral imaging equipment, containing spatial information based on pixel coordinates and spectral intensity information based on wavelength channels. Chart elements refer to two-dimensional graphical structures used to visualize data relationships, including geometric shapes composed of coordinate axes, data points, connecting lines, and their fill color attributes. Formula symbols refer to special characters with specific semantics in mathematical expressions, including operators (such as ∑, ∫), variable symbols (such as α, β), and structural markers (such as matrix brackets).

[0059] In this embodiment of the invention, raw image data of documents containing chart elements and formula symbols are acquired using a high-resolution spectral imaging device. This device integrates a dual-band acquisition module for visible light (400-760nm) and near-infrared (760-1100nm), employing a 1200dpi resolution progressive scan method to simultaneously record the spatial coordinates (X,Y) of each pixel and the intensity values ​​(λ1-λ256) of 256 spectral channels. Specifically, the color information of the chart elements is captured by a red, green, and blue primary color sensor array, forming visible light spectral data carrying the cyan, magenta, yellow, and black printing color gamut distribution; the material reflectance characteristics of the formula symbols are sampled using the reflectance intensity of 16 characteristic wavelengths (e.g., 850nm, 940nm) in the near-infrared band, generating material reflectance data containing ink density and paper texture information. The visible light spectral data and material reflectance data mentioned above are stored in a three-dimensional tensor format, with the dimensions corresponding to the image width, height and number of spectral channels, respectively, to obtain a literature image dataset. For example, a standard A4 literature image generates a 5952×8424×256 original data cube, where each spatial location (i,j) corresponds to a 256-dimensional spectral intensity vector, which can accurately characterize the color gradient changes of the chart borders and the metallic ink reflectance characteristics of the formula symbols.

[0060] Step 102: Separate the different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution features and material reflection features;

[0061] In this embodiment of the invention, the document image dataset is divided into visible light band data and near-infrared band data according to the spectral band range; the spectral response intensity distribution related to color attributes in the visible light band data is identified to generate visible light spectral data carrying the intensity information of the three primary colors; spectral absorption feature points related to material reflection characteristics in the near-infrared band data are detected to generate near-infrared spectral data carrying the reflection intensity information of a specific wavelength; the two types of data are spliced ​​together to form multi-channel spectral data containing color distribution features and material reflection features.

[0062] Step 103: Divide the multi-channel spectral data into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and calculate the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands;

[0063] In this embodiment of the invention, the spectral intensity of each spectral channel in the multi-channel spectral data is statistically analyzed to generate a spectral intensity distribution curve corresponding to each spectral channel; the band intervals on the curve where the rate of change of spectral intensity exceeds a preset spectral sensitivity threshold are marked as high-sensitivity bands, and the band intervals below the threshold are marked as low-sensitivity bands; a first spectral intensity extreme point is extracted from the high-sensitivity band, a second spectral intensity extreme point is extracted from the low-sensitivity band, and the absolute value of the intensity difference between the two types of extreme points is calculated as the spectral feature difference quantity.

[0064] Step 104: Perform dynamic bit width allocation on the spectral feature difference to generate spatial continuity coded data and frequency domain distribution coded data;

[0065] In this embodiment of the invention, the intensity ratio of spatial continuity attribute to frequency distribution attribute in the spectral feature difference is calculated to determine the first target bit width value corresponding to the spatial continuity attribute and the second target bit width value corresponding to the frequency distribution attribute; the spectral intensity change between adjacent pixel positions in the color distribution feature is calculated to generate a set of spatial continuity changes; the intensity fluctuation amplitude at different wavelength positions in the material reflection feature is calculated to generate a set of frequency distribution fluctuations; and the set of spatial continuity changes and the set of frequency distribution fluctuations are encoded according to the first target bit width value and the second target bit width value to generate spatial continuity encoded data and frequency distribution encoded data.

[0066] Step 105: Analyze the spatial continuity coding data and the frequency domain distribution coding data using a large model to generate a structured text description;

[0067] In this embodiment of the invention, the element position coordinate sequence in the spatial continuity coding data is parsed to generate a set of spatial position descriptions of chart elements. The correspondence between wavelength and reflection intensity in the frequency domain distribution coding data is parsed to generate a set of material property descriptions of formula symbols. Based on the two sets, an element symbol association mapping table is generated to retrieve the chart element position description from the spatial position description set and the formula symbol material description from the material property description set. The chart element position descriptions associated with the same document position and the formula symbol material descriptions are combined into a composite text paragraph containing chart element topology descriptions and formula symbol attribute descriptions. All composite text paragraphs corresponding to document positions are integrated to generate a structured text description.

[0068] Step 106: Verify the information coverage of the structured text description relative to the document image dataset to generate evaluation results;

[0069] In this step, information coverage refers to the proportion of non-textual element features correctly represented in the generated structured text description to the total number of non-textual element features in the original document, used to quantify content completeness. The evaluation result refers to the qualitative judgment conclusion generated based on information coverage, including qualified labels used to identify completeness or a list of coordinate-wavelength missing features used to locate deficiencies.

[0070] In this embodiment of the invention, the spatial coordinate information of the chart elements and the wavelength features of the formula symbols in the structured text description are matched with the corresponding elements in the document image dataset using a dual mapping of spatial location and spectral features. The matching validity is determined by comparing coordinate alignment and wavelength features. The number of correctly matched non-text elements is counted and divided by the total number of non-text elements in the dataset to obtain the information coverage. When the information coverage exceeds a preset integrity threshold, a qualified assessment result is generated; otherwise, a missing report containing coordinate location and wavelength features (i.e., a qualified assessment result) is generated based on the unmatched elements.

[0071] This invention addresses the color shift and detail loss issues caused by reliance on red, green, and blue channels in existing document analysis methods. It also enhances the ability to structurally describe high-density charts and complex mathematical formulas; improves the model's adaptability to material differences and image quality fluctuations; enhances the semantic integrity and logical coherence of the generated text; and enables quantitative evaluation of the quality of document content restoration. This invention comprehensively overcomes the technical barriers of traditional methods in non-textual element recognition and structured expression, providing an efficient and robust solution for high-quality multimedia document analysis.

[0072] This invention provides a specific embodiment. Step 102 involves separating the different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution features and material reflectance features. Specifically, this includes the following steps:

[0073] Step 201: Divide the document image dataset into visible light band data and near-infrared band data according to the spectral band range;

[0074] In this step, the spectral band range refers to the electromagnetic wave wavelength range (400-2500nm) that the spectral imaging device can capture, including the visible light band (400-700nm) reflecting color information and the near-infrared band (700-2500nm) reflecting material information. Visible light band data refers to the set of spectral intensities within the wavelength range of 400-700nm, used for color analysis based on human visual perception. Near-infrared band data refers to the set of spectral intensities within the wavelength range of 700-2500nm, used for material property analysis based on molecular vibrations.

[0075] In this embodiment of the invention, the document image dataset is physically segmented according to the spectral band range by an optical beam splitting system: visible light band data (carrying color information) is output in the wavelength range of 400-700nm, and near-infrared band data (carrying material information) is output in the wavelength range of 700-2500nm. The segmentation process achieves band isolation based on the reflection / transmission characteristics of a dichroic mirror.

[0076] Step 202: Identify the spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectral data carrying the intensity information of the three primary colors;

[0077] In this step, color attributes refer to the visual characteristics formed by light reflected from an object's surface in the visible light band, including hue, saturation, and brightness. Spectral response intensity distribution refers to the set of intensity response values ​​of different wavelengths of light on the sensor, reflecting the distribution of color energy. Primary color intensity information refers to the standardized intensity values ​​(range 0-1) of the three primary colors (red, green, and blue) at pixel locations, used to accurately characterize color. Visible light spectral data refers to a data structure that integrates spatial coordinates and primary color intensity information, in the format [row coordinate, column coordinate, R intensity, G intensity, B intensity].

[0078] In this embodiment of the invention, the spectral intensity values ​​of each pixel in the visible light band are measured in the red-sensitive, green-sensitive, and blue-sensitive bands; based on the spectral intensity values ​​of the red-sensitive band, intensity distribution data of the red primary color in each pixel is generated; based on the spectral intensity values ​​of the green-sensitive band, intensity distribution data of the green primary color in each pixel is generated; based on the spectral intensity values ​​of the blue-sensitive band, intensity distribution data of the blue primary color in each pixel is generated; the above three types of data are integrated according to spatial location to form visible light spectral data carrying the intensity information of the three primary colors.

[0079] Step 203: Detect spectral absorption feature points related to the material's reflectivity in the near-infrared band data to generate near-infrared spectral data carrying specific wavelength reflectivity information;

[0080] In this step, material reflectivity refers to the characteristic reflection / absorption patterns of an object in the near-infrared band due to its molecular structure. Spectral absorption characteristic points refer to the wavelength positions where the reflection intensity suddenly drops (a decrease of >10%), corresponding to the material's molecular resonance absorption peak. Specific wavelength reflection intensity information refers to the mapping pair between the target absorption point wavelength and its reflection intensity value (e.g., {1200nm: 0.32, 1550nm: 0.18}). Near-infrared spectral data refers to the dataset storing the wavelength-reflection intensity mapping relationship and its corresponding spatial coordinates.

[0081] In this embodiment of the invention, spectral absorption feature points in near-infrared band data with reflection intensity lower than adjacent wavelength positions are detected; target absorption points matching the material reflection characteristics are selected by combining a preset material reflection characteristic library (which stores the absorption wavelengths of common printing materials); the reflection intensity value at the wavelength position of the target absorption point is extracted, and reflection intensity distribution data carrying wavelength intensity mapping relationship is generated, and it is bound to the spatial location information of near-infrared band data to generate near-infrared spectral data.

[0082] Step 204: Perform channel splicing of the visible light spectral data and the near-infrared spectral data to form multi-channel spectral data that includes color distribution characteristics and material reflectance characteristics;

[0083] In this step, color distribution characteristics refer to the pixel-level color space distribution pattern described by visible light spectral data. Material reflection specifically refers to the spatial distribution pattern of material absorption characteristics described by near-infrared spectral data. Multi-channel spectral data refers to a four-dimensional data matrix formed by merging the three visible light channels and the single near-infrared channel, with dimensions [height, width, 4].

[0084] In this embodiment of the invention, the three primary color channels of visible light spectral data and the single channel of near-infrared spectral data are aligned according to the data dimension; the four channels are merged into a unified multi-channel matrix through a spatial coordinate matching circuit, wherein the first three channels store color distribution characteristics (i.e., red, green and blue intensities), the fourth channel stores material reflection characteristics (i.e., reflection intensity at a specific wavelength), and the final output is multi-channel spectral data that simultaneously contains both types of features.

[0085] This invention captures near-infrared features through physical spectroscopy, effectively identifying the absorption characteristics of printing ink molecules (such as the characteristic absorption of carbon-based inks at 1200nm), solving the color distortion problem caused by fading of old documents; the dual-band independent processing mechanism improves the accuracy of color restoration of charts and graphs, while also achieving accurate differentiation of the material properties of formula symbols (such as laser printing toner vs. inkjet ink), laying a data foundation for subsequent cross-modal correlation analysis.

[0086] This invention provides a specific embodiment, step 203, detecting spectral absorption feature points related to the material's reflectivity in the near-infrared band data to generate near-infrared spectral data carrying specific wavelength reflectance intensity information, specifically including the following steps:

[0087] Step 211: Detect spectral absorption feature points in the near-infrared band data where the reflection intensity is lower than that of adjacent wavelengths;

[0088] In this step, reflection intensity refers to the near-infrared light energy value reflected by the object's surface, ranging from 0 to 1 (0 for total absorption, 1 for total reflection), reflecting the vibrational absorption characteristics of the material's molecules. The spectral absorption characteristic point refers to the wavelength position where the reflection intensity suddenly drops; the criterion is that the intensity at this point is ≤ 85% of the average intensity of the three adjacent points, corresponding to the material's molecular resonance absorption peak.

[0089] In this embodiment of the invention, the sliding window method is used to traverse the wavelength sequence of near-infrared band data. For each wavelength position, the average value of its reflection intensity and that of the three adjacent wavelength positions is calculated. When the reflection intensity of the wavelength position is less than 85% of the average value, that is, the reflection intensity ≤ average value × 0.85, the corresponding wavelength position is marked as a spectral absorption feature point.

[0090] Step 212: Match the wavelength positions of the spectral absorption feature points according to the preset material reflection characteristic library to screen out target absorption points associated with the material reflection characteristics;

[0091] In this step, the preset material reflectivity library refers to a pre-stored database of characteristic absorption wavelengths of printing materials, including standard absorption wavelengths and fluctuation ranges of materials such as toner and coated paper based on laboratory measurements. The target absorption point refers to the effective absorption point selected through matching with the characteristic library, whose wavelength deviates from the standard value in the library by ≤±5nm, and is used to identify specific printing materials.

[0092] In this embodiment of the invention, the wavelength position of the spectral absorption feature point is matched with a preset material reflectance characteristic library. Specifically, the material reflectance characteristic library stores the characteristic absorption wavelengths of printing materials (such as 1550nm for toner and 1200nm for coated paper). When the wavelength of a certain spectral absorption feature point differs from the characteristic wavelength in the library within ±5nm, it is determined to be a target absorption point associated with the material reflectance characteristics.

[0093] Step 213: Extract the reflection intensity value at the wavelength position of the target absorption point and generate reflection intensity distribution data carrying wavelength intensity mapping relationship;

[0094] In this step, the reflection intensity value refers to the measured reflected light energy at the wavelength position of the target absorption point, expressed as a floating-point number; the lower the value, the stronger the absorption. The wavelength-intensity mapping relationship refers to the set of key-value pairs between wavelength values ​​and corresponding reflection intensity values ​​(e.g., {1550: 0.23, 1700: 0.41}), characterizing the material's characteristic absorption spectrum. The reflection intensity distribution data refers to the data structure that integrates the wavelength-intensity mapping relationship of the target absorption point, in the format [wavelength 1: intensity 1, wavelength 2: intensity 2,...].

[0095] In this embodiment of the invention, the reflection intensity corresponding to each target absorption point in the original near-infrared band data is extracted; a mapping relationship with wavelength as the key and reflection intensity as the value is constructed, such as a wavelength of 1550nm corresponding to a reflection intensity of 0.23, and reflection intensity distribution data composed of key-value pairs is generated.

[0096] Step 214: Bind the reflection intensity distribution data with the spatial location information of the near-infrared band data to generate near-infrared spectral data carrying reflection intensity information of a specific wavelength;

[0097] In this step, spatial location information refers to the two-dimensional coordinates (row number, column number) of a pixel in the document image, which is used to locate the physical location of non-text elements.

[0098] In this embodiment of the invention, each key-value pair in the reflectance intensity distribution data is associated with the spatial coordinates of the near-infrared band data. Specifically, the pixel region corresponding to each wavelength is located by using a spatial index matrix, and a third-dimensional coordinate information (row number, column number) is added to the reflectance intensity distribution data to form a three-dimensional near-infrared spectral data, the format of which is [row coordinate, column coordinate, wavelength-intensity mapping relationship].

[0099] The embodiments of this invention effectively identify weak feature peaks through neighborhood contrast absorption point detection; the material reflection characteristic library accurately matches and solves the pain point of traditional solutions being unable to distinguish similar materials, thus reducing the material misclassification rate; wavelength-space dual-dimensional binding ensures that the material properties and positions of the formula symbols correspond precisely, providing reliable input for subsequent cross-modal analysis.

[0100] Taking a specific literature analysis as an example, this literature contains a formula for toner printing on brass paper. Following the aforementioned steps, near-infrared data (wavelength range 700-2500nm, resolution 5nm) has been obtained. First, the system detects spectral absorption feature points in the near-infrared data where the reflection intensity is lower than adjacent wavelengths. Specifically, using a sliding window method, a reflection intensity of 0.19 is detected at wavelength 1550nm, with an average intensity of 0.83 for its three adjacent points. It is determined that 0.19 ≤ 0.83 × 0.85 (i.e., 0.7055), therefore this wavelength position is marked as a spectral absorption feature point. Subsequently, based on a preset material reflection characteristic library, this spectral absorption feature point (1550nm) is compared with a standard wavelength. The standard absorption peak of toner is 1552nm ± 3nm. Calculation shows that |1550–1552| = 2nm < 3nm, indicating that this point conforms to the reflection characteristics of the toner material and is listed as a target absorption point associated with the material's reflection characteristics. Next, the reflection intensity value of 0.19 at the wavelength position of 1550nm, where the target absorption point is located, is extracted to generate reflection intensity distribution data {1550:0.19} carrying the wavelength intensity mapping relationship. Finally, this reflection intensity distribution data is bound to the spatial location information in the near-infrared band data, locating the region where the formula symbol is located in the image as (120,45)-(130,55), and the corresponding near-infrared spectral data [rows 120-130, columns 45-55, {1550:0.19}]. The method has been verified to successfully identify the formula symbol printed with toner.

[0101] This invention provides a specific embodiment. Step 103 involves dividing the multi-channel spectral data into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and calculating the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands. This specifically includes the following steps:

[0102] Step 301: Calculate the spectral intensity of each spectral channel in the multi-channel spectral data to generate the spectral intensity distribution curve corresponding to each spectral channel;

[0103] In this step, spectral intensity refers to the energy value of light reflected by an object at a specific wavelength, ranging from 0 to 1 (0 for no reflection, 1 for total internal reflection), reflecting the optical response characteristics of the material. The spectral intensity distribution curve is a continuous function curve with wavelength (unit: nanometers) as the abscissa and average spectral intensity as the ordinate, used to visualize the spectral energy distribution pattern.

[0104] In this embodiment of the invention, a wavelength-by-wavelength scan is performed on each spectral channel (such as the visible red, green, blue, and near-infrared channels) of the multi-channel spectral data: the spectral intensity values ​​of all pixels are accumulated at each wavelength position and divided by the total number of pixels to generate the average spectral intensity at that wavelength position; a continuous curve is plotted with wavelength as the horizontal axis and average intensity as the vertical axis to output the spectral intensity distribution curve containing peak points and valley points.

[0105] Step 302: Mark the band intervals on the spectral intensity distribution curve where the rate of change of spectral intensity exceeds a preset spectral sensitivity threshold as high-sensitivity bands, and mark the band intervals where the rate of change of spectral intensity is lower than the preset spectral sensitivity threshold as low-sensitivity bands.

[0106] In this step, the rate of change of spectral intensity refers to the amount of change in spectral intensity within a unit wavelength interval, reflecting the dynamic sensitivity of the spectral response. The preset spectral sensitivity threshold is a pre-defined critical value for the rate of change, used to distinguish between high dynamic response regions and stable response regions. The high-sensitivity band refers to the continuous wavelength range where the rate of change of spectral intensity exceeds the threshold, corresponding to abrupt changes in features such as the edges of figures and tables, and the boundaries of formula symbols in the literature. The low-sensitivity band refers to the continuous wavelength range where the rate of change of spectral intensity is below the threshold, corresponding to the background or uniformly filled areas in the literature.

[0107] In this embodiment of the invention, the first derivative of the spectral intensity distribution curve is calculated (the current wavelength intensity value minus the previous wavelength intensity value, and then divided by the wavelength interval) to obtain a sequence of spectral intensity change rates; continuous wavelength intervals with spectral intensity change rates exceeding a preset spectral sensitivity threshold (e.g., 0.25 intensity units / nanometer) are marked as high-sensitivity bands; continuous wavelength intervals with change rates below the threshold are marked as low-sensitivity bands.

[0108] Step 303: Extract the first spectral intensity extreme point from the high-sensitivity band, and simultaneously extract the second spectral intensity extreme point from the low-sensitivity band. Calculate the absolute value of the intensity difference between the first and second spectral intensity extreme points, and use the absolute value of the intensity difference as the spectral feature difference quantity.

[0109] In this step, the first spectral intensity extreme point refers to the local maximum value of the spectral intensity distribution curve within the high-sensitivity band, representing the intensity peak of the key feature region. The second spectral intensity extreme point refers to the local minimum value of the spectral intensity distribution curve within the low-sensitivity band, representing the intensity valley of the background reference region. The spectral feature difference refers to the absolute value of the difference between the intensity of the first extreme point and the intensity of the second extreme point, quantifying the contrast difference between the key feature and the background.

[0110] In this embodiment of the invention, the local maximum point of the spectral intensity distribution curve in the high-sensitivity band interval is located as the first spectral intensity extreme point; the local minimum point in the low-sensitivity band interval is located as the second spectral intensity extreme point; the absolute value of the intensity difference is obtained by subtracting the intensity value of the second spectral intensity extreme point from the intensity value of the first spectral intensity extreme point, and this absolute value is used as the spectral feature difference quantity.

[0111] This invention effectively distinguishes the edges of chart elements from the paper background through band division driven by the rate of change; the extreme point difference quantification significantly enhances the boundary features of formula symbols, thereby improving the key feature recognition rate of fuzzy documents; and the dynamic threshold determination adaptively processes documents of different printing quality, solving the problem of poor generalization of traditional fixed threshold schemes.

[0112] This invention provides a specific embodiment. Step 104 involves performing dynamic bit-width allocation on the spectral feature difference to generate spatially continuous coded data and frequency domain distributed coded data, specifically including the following steps:

[0113] Step 401: Calculate the intensity ratio of the spatial continuity attribute to the frequency domain distribution attribute in the spectral feature difference quantity;

[0114] In this step, the spatial continuity attribute refers to the component in the spectral feature difference that reflects the spatial structural stability of the chart elements, calculated based on the intensity of the extreme points at the chart edges within the high-sensitivity band. The frequency domain distribution attribute refers to the component in the spectral feature difference that reflects the material reflection changes of the formula symbols, calculated based on the extreme intensity of the material absorption feature points. The intensity ratio value is the quotient of the spatial continuity attribute intensity divided by the frequency domain distribution attribute intensity, used to determine the bit width resource allocation weight.

[0115] In this embodiment of the invention, the intensity values ​​of the spatial continuity attribute (from the extreme point intensity of the edge of the chart element in the high-sensitivity band) and the intensity values ​​of the frequency domain distribution attribute (from the extreme point intensity of the absorption peak of the material reflection feature) are extracted from the spectral feature difference quantity; the intensity value of the spatial continuity attribute is divided by the intensity value of the frequency domain distribution attribute to generate an intensity ratio value that characterizes the energy ratio relationship between the two.

[0116] Step 402: Determine the first target bit width value corresponding to the spatial continuity attribute and the second target bit width value corresponding to the frequency domain distribution attribute based on the intensity ratio value;

[0117] In this step, the first target bit width value refers to the number of bits allocated to the spatial continuity attribute, and its value is positively correlated with the intensity ratio value, used to control the encoding accuracy of the chart structure. The second target bit width value refers to the number of bits allocated to the frequency domain distribution attribute, and its value is negatively correlated with the intensity ratio value, used to control the encoding accuracy of the material features.

[0118] In this embodiment of the invention, the intensity ratio is used as a weighting ratio, and the weighting ratio is multiplied by a preset bit width base to generate an initial bit width value for spatial continuity; the bit width base is adjusted according to the fluctuation range of the frequency domain distribution attribute to generate an initial bit width value for frequency domain distribution; the initial bit width value for frequency domain distribution is corrected according to the intensity stability between different wavelength positions to generate a second target bit width value; and the initial bit width value for spatial continuity is corrected according to the intensity correlation between adjacent pixel positions to generate a first target bit width value.

[0119] Step 403: Calculate the spectral intensity variation between adjacent pixel positions in the color distribution feature to generate a set of spatial continuity variation.

[0120] In this step, the spectral intensity change refers to the absolute value of the intensity difference between adjacent spatial directions (up / down / left / right) at the same pixel location, quantifying the degree of abrupt changes in local structure. The spatial continuity change set refers to a two-dimensional matrix storing the intensity differences between all pixel locations and their four neighbors, with dimensions equal to the size of the document image.

[0121] In this embodiment of the invention, each pixel position in the color distribution feature is traversed, and the spectral intensity difference between it and its upper, lower, left, and right adjacent pixels (current pixel intensity minus adjacent pixel intensity) is calculated; the absolute values ​​of the differences between all adjacent positions are stored as a two-dimensional matrix according to pixel coordinates to generate a set of spatial continuity changes.

[0122] Step 404: Calculate the intensity fluctuation amplitude at different wavelength positions in the material reflection characteristics to generate a set of frequency domain distribution fluctuations;

[0123] In this step, the intensity fluctuation amplitude refers to the difference between the maximum and minimum reflection intensity at the same wavelength location in the frequency domain, characterizing the material's reflection stability. The frequency domain distribution fluctuation set refers to a one-dimensional array that stores the fluctuation amplitude at each location according to the wavelength sequence, with a length equal to the number of wavelength sampling points.

[0124] In this embodiment of the invention, each wavelength position in the material's reflection characteristics is traversed, and the intensity difference between its maximum and minimum reflection intensity is calculated as the intensity fluctuation amplitude; the intensity fluctuation amplitude values ​​at different wavelength positions (interval of 10nm) are stored as a one-dimensional array according to the wavelength sequence to generate a set of frequency domain distribution fluctuations.

[0125] Step 405: Based on the first target bit width value, encode the set of spatial continuity changes to generate spatial continuity encoded data;

[0126] In this step, spatial continuity encoded data refers to the chart structure feature data after bit-width constraint compression, which contains a binary stream of pixel position-intensity change mapping.

[0127] In this embodiment of the invention, an adaptive arithmetic encoder is used to process the set of spatial continuity variations: the encoding precision is determined according to the first target bit width value, such as the first target bit width value of 8 corresponding to the integer range of 0-255. The intensity difference between adjacent pixels is quantized into integers and then compressed and encoded to obtain spatial continuity encoded data in binary format.

[0128] Step 406: Based on the second target bit width value, encode the set of frequency domain distribution fluctuations to generate frequency domain distribution coded data;

[0129] In this step, the frequency domain distributed coding data refers to the material feature data after bit-width constraint compression, which includes a binary stream with wavelength-fluctuation amplitude mapping.

[0130] In this embodiment of the invention, differential pulse coding is performed on the frequency domain distribution fluctuation set: the fluctuation amplitude quantization order is limited according to the second target bit width value, such as the second target bit width value of 4 corresponding to 16 levels, the fluctuation amplitude sequence is differentially calculated and encoded to obtain frequency domain distribution coded data.

[0131] This invention improves the retention rate of key information by dynamically allocating resources through attribute-aware bit-width allocation and intensity ratio values ​​(e.g., allocating high bits to chart structures and low bits to material features); spatial neighborhood differential encoding effectively captures the boundary abrupt features of chart elements (e.g., the outline of a bar chart), solving the problem of structural distortion in fuzzy documents; frequency domain fluctuation quantization compression preserves the material reflection characteristics of formula symbols, avoiding misjudgment of materials in complex formulas by traditional schemes.

[0132] This invention provides a specific embodiment. Step 402 involves determining the first target bit width value corresponding to the spatial continuity attribute and the second target bit width value corresponding to the frequency domain distribution attribute based on the intensity ratio value. This specifically includes the following steps:

[0133] Step 411: Using the intensity ratio value as the weight ratio value, multiply the weight ratio value with the preset bit width base to generate the initial bit width value for spatial continuity.

[0134] In this step, the weight ratio refers to the weight proportion of the two types of features in the spectral difference. The preset bit width base value refers to the initial bit width base value set according to system resources, used to calculate the initial bit width allocation of each attribute. The initial bit width value for spatial continuity refers to the initial number of bits for spatial attributes calculated based on the weight ratio and the bit width base value, without considering local correlation optimization.

[0135] In this embodiment of the invention, the weight ratio is multiplied by a preset bit width base. Specifically, when the weight ratio is greater than 1, the initial bit width value for spatial continuity is equal to the bit width base × the weight ratio; when the weight ratio is less than 1, the initial bit width value for spatial continuity is equal to the bit width base × (1 ÷ weight ratio), ensuring that spatial features are given priority protection.

[0136] Step 412: Based on the fluctuation range of the frequency domain distribution attribute, the bit width base is proportionally scaled and adjusted to generate the initial bit width value of the frequency domain distribution;

[0137] In this step, the fluctuation range refers to the difference between the maximum and minimum reflection intensity of the frequency domain distribution attribute across all wavelength positions, reflecting the overall fluctuation of the material's reflection. The initial bit width value of the frequency domain distribution refers to the initial number of bits of the frequency domain attribute obtained by scaling the bit width base based on the fluctuation range.

[0138] In this embodiment of the invention, the difference between the maximum and minimum reflection intensity of the frequency domain distribution attribute at all wavelength positions is calculated to obtain the fluctuation range; the initial bit width value of the frequency domain distribution = bit width base × fluctuation range scaling factor, wherein the fluctuation range scaling factor = 1 - fluctuation range ÷ maximum possible fluctuation value.

[0139] Step 413: Based on the intensity stability between different wavelength positions, the initial bit width value of the frequency domain distribution is corrected to generate a second target bit width value;

[0140] In this step, intensity stability refers to the variance of reflection intensity at different wavelength positions, quantifying the consistency of the wavelength dimension of material reflection. Intensity correlation refers to the Pearson correlation coefficient (covariance divided by standard deviation) calculated using a 3×3 pixel neighborhood, measuring the color continuity of local areas of chart elements.

[0141] In this embodiment of the invention, the variance of the reflection intensity at adjacent wavelength positions is calculated as the intensity stability, that is, the intensity stability = the sum of the squares of the differences between the reflection intensity and the mean at each wavelength position divided by the number of wavelengths; when the intensity stability exceeds a preset threshold, the second target bit width value = the initial bit width value of the frequency domain distribution × 1.2 to enhance the stability characterization; otherwise, the original value is maintained as the second target bit width value.

[0142] Step 414: Based on the strength correlation between adjacent pixel positions, the initial bit width value of spatial continuity is corrected to generate the first target bit width value;

[0143] In this embodiment of the invention, a 3×3 neighborhood of a pixel is selected, and the Pearson correlation coefficient between the intensity of the central pixel and the surrounding pixels is calculated as the degree of intensity correlation. When the Pearson correlation coefficient is greater than 0.7, it indicates that the intensity of the neighboring pixels and the central pixel is highly correlated, corresponding to smooth areas in the document image (such as chart backgrounds or uniformly filled color blocks). Such areas have strong spatial continuity and high information redundancy, and the amount of data can be compressed by reducing the bit width. In this case, the first target bit width value = initial bit width value of spatial continuity × 0.8. When the Pearson correlation coefficient ≤ 0.7, it indicates that the intensity difference between the neighboring pixels and the central pixel is significant, corresponding to edges, textures, or boundaries of chart elements in the document image (such as formula symbol outlines or line graph lines). Such areas have weak spatial continuity but contain key structural information, and the bit width needs to be increased to retain details. In this case, the first target bit width value = initial bit width value of spatial continuity × 1.1.

[0144] This invention addresses the issue of blurred edges in high-density curve charts by dynamically allocating weight ratios to increase bit width resources in densely populated areas of the chart; automatically adjusting the frequency domain bit width based on material reflection stability; correcting the intensity stability to avoid over-allocation of stable materials, thus improving the compression rate of frequency domain data; and correcting the intensity correlation to strengthen the bit width allocation of broken chart elements, thereby improving the fidelity of key structures.

[0145] This invention provides a specific embodiment. Step 105 involves parsing the spatial continuity coding data and the frequency domain distribution coding data using a large model to generate a structured text description. This specifically includes the following steps:

[0146] Step 501: Parse the element position coordinate sequence in the spatial continuity coding data to generate a spatial position description set of the chart elements, wherein the spatial position description set includes a first identifier and a first document position label;

[0147] In this step, the element position coordinate sequence refers to the set of geometric vertex coordinates of the chart elements extracted from the spatially encoded data (such as the line chart point sequence [x1,y1,x2,y2,...]), reflecting the spatial topology of the chart. The spatial location description set refers to the set of key-value pairs storing the chart element identifier and its document location label, used to locate the physical location of the chart. The first identifier refers to the unique alphanumeric code assigned to the chart element, used for cross-module citation. The first document location label describes the spatial location information of the chart element in the document, in the format: page number, geometric type (coordinates), such as page 2, where the center [50,60,15] represents page 2, and the circle with center (50,60) and radius 15.

[0148] In this embodiment of the invention, the spatial continuity encoded data is decoded by the geometric analysis module of the large model, and the vertex coordinate sequence of the chart elements (such as the coordinates of data points in a line chart) stored therein is extracted. A unique first identifier is assigned to each chart element, such as icon 1, and the coordinates of the document page area it occupies are marked as the first document location label, such as the rectangular area [120,50,180,80] on page 3. Finally, a set of spatial location descriptions containing key-value pairs of {first identifier: first document location label} is output.

[0149] Step 502: Analyze the correspondence between wavelength and reflection intensity in the frequency domain distributed coding data to generate a set of material characteristic descriptions of formula symbols. The set of material characteristic descriptions includes a second identifier and a second document location label.

[0150] In this step, the correspondence between wavelength and reflection intensity refers to the material characteristic data parsed from the frequency domain encoded data, such as {1200nm: 0.35, 1550nm: 0.18}, characterizing the reflection spectrum properties of the formula symbol. The material characteristic description set refers to the set of key-value pairs storing the formula symbol identifier and its material characteristics. The second identifier refers to the unique code assigned to the formula symbol, containing mathematical semantic markers. The second document location label refers to the center point coordinates of the formula symbol, in the format: page number, point position (x, y).

[0151] In this embodiment of the invention, the frequency domain analysis module of the large model is used to decode the frequency domain distribution coding data and extract the wavelength-reflection intensity mapping relationship, such as the reflection intensity corresponding to 1550nm being 0.23; a second identifier is assigned to each formula symbol, such as formula ∫, and its center point coordinates are marked as the second literature location label, such as the point coordinates [88,72] on page 5, and the output is a set of material property descriptions containing the key-value pairs {second identifier: second literature location label}.

[0152] Step 503: Based on the consistency of document location tags, perform cross-modal matching between the first identifier and the second identifier to generate an element symbol association mapping table;

[0153] In this step, document location label consistency refers to the matching condition that the spatial distance between two document location labels is ≤5 pixels, ensuring that the physical locations of associated elements are close. The element symbol association mapping table is a table that records the mapping relationship between charts and formula identifiers, used for cross-modal association.

[0154] In this embodiment of the invention, the spatial location description set and the material characteristic description set are matched by coordinates according to the document location labels. Specifically, when the spatial distance between the location labels corresponding to the two identifiers is less than a preset threshold (e.g., 5 pixels), it is determined that the document location label consistency is satisfied; a mapping relationship between the first identifier and the second identifier is established, and an element symbol association mapping table is generated.

[0155] Step 504: Based on the element symbol association mapping table, retrieve the chart element position description from the spatial position description set, and retrieve the formula symbol material description from the material property description set;

[0156] In this step, the chart element location description refers to the natural language description of the spatial location of the chart element, such as a line chart spanning the left half of the page with a horizontal axis range of 0-100. The formula symbol material description refers to the physical property description of the formula symbol; for example, the integral symbol ∫ uses coated paper ink with a characteristic absorption peak of 1200nm.

[0157] In this embodiment of the invention, the location description of the chart element is retrieved from the spatial location description set according to the association relationship of the element symbol association mapping table; the material description of the formula symbol is retrieved from the material characteristic description set, such as formula ∑: toner printing, 1550nm absorption intensity 0.18.

[0158] Step 505: Combine the location descriptions of chart elements associated with the same document location with the material descriptions of formula symbols into a composite text paragraph containing topological descriptions of chart elements and attribute descriptions of formula symbols. Integrate the composite text paragraphs corresponding to all document locations to generate a structured text description.

[0159] In this step, the chart element topology description refers to a detailed description of the chart element; for example, a bar chart contains three sets of data with a height ratio of 2:5:3. The formula symbol attribute description refers to the material and semantic description of the formula symbols, such as the α symbol representing an angle variable, printed on coated paper. The compound text paragraph refers to a descriptive paragraph integrating the chart topology and formula attributes, formatted as: [Location]: Chart Description...; Related Formula Description.... The structured text description refers to a sequence of compound text paragraphs ordered by document location, forming the complete report.

[0160] In this embodiment of the invention, for each chart element and formula symbol group associated through the element symbol association mapping table (based on document location label consistency matching, such as spatial distance ≤ 5 pixels), the position description of the chart element is converted into a chart element topology description that includes spatial location (such as the upper right quadrant of page 4), geometric type, and data relationship (such as a bar chart containing 5 data bars with a height ratio of 1:3:2:4:5). The material description of the formula symbol is converted into a formula symbol attribute description that covers printing material (such as laser toner), spectral characteristics (1550nm reflection peak), and mathematical semantics (such as ∑ representing data accumulation operation, and then in the format: on [page 4, coordinates (88,72), surrounding area]: the bar chart shows the data trend from 2020 to 2024; the associated formula ∑ uses toner printing to calculate the sum of growth rates). The chart topology and formula attributes of the same document location are merged into a composite text paragraph. Finally, all composite text paragraphs are sorted according to page order and spatial location to generate a structured text description, realizing the semantic integration of non-text elements based on location association, and solving the problem of semantic fragmentation of cross-modal elements in traditional solutions.

[0161] This invention improves the accuracy of identifying the association between formula symbols and adjacent bar charts through cross-modal position matching, solving the semantic fragmentation problem of traditional solutions; material-enhanced semantic parsing combined with printing attributes improves the recognition rate of formula symbols, especially improving the discrimination of handwritten symbols; structured paragraph generation improves the completeness of the description of the association between charts and formulas in academic literature, meeting the standards of journal abstracts.

[0162] Figure 2 This invention provides a schematic diagram of the structure of a literature report generation and evaluation system based on a large model, as shown in the embodiment of the invention. Figure 2 As shown, the system includes:

[0163] The acquisition module 21 is used to acquire a dataset of document images containing non-text elements, including chart elements and formula symbols;

[0164] The separation module 22 is used to separate the different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution features and material reflection features;

[0165] The calculation module 23 is used to divide the multi-channel spectral data into high-sensitivity bands and low-sensitivity bands according to a preset spectral sensitivity threshold, and to calculate the spectral feature difference between the high-sensitivity bands and the low-sensitivity bands.

[0166] Allocation module 24 is used to perform dynamic bit width allocation on the spectral feature difference to generate spatially continuous coded data and frequency domain distributed coded data;

[0167] The generation module 25 is used to parse the spatial continuity coded data and the frequency domain distribution coded data through a large model to generate a structured text description;

[0168] The verification module 26 is used to verify the information coverage of the structured text description relative to the document image dataset in order to generate evaluation results.

[0169] Figure 2 The aforementioned literature report generation and evaluation system based on a large model can perform... Figure 1 The implementation principle and technical effects of the literature report generation and evaluation method based on a large model, as described in the above embodiments, will not be repeated here. The specific methods by which each module and unit of the literature report generation and evaluation system based on a large model performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0170] In one possible design, Figure 2 The illustrated embodiment of a literature report generation and evaluation system based on a large model can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0171] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.

[0172] The processing component 32 is used for the above Figure 1 The embodiment describes a method for evaluating the generation of literature reports based on a large model.

[0173] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0174] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0175] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.

[0176] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.

[0177] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0178] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.

[0179] This invention also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The illustrated embodiment presents a literature report generation and evaluation method based on a large model.

[0180] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model-based literature report generation evaluation method, characterized in that, The method comprises: obtaining a literature image dataset containing non-text elements, the non-text elements containing chart elements and formula symbols; separating different spectral band data of the literature image dataset to generate multi-channel spectral data containing color distribution features and material reflection features; dividing high-sensitivity bands and low-sensitivity bands of the multi-channel spectral data according to a preset spectral sensitivity threshold, and calculating a spectral feature difference between the high-sensitivity bands and the low-sensitivity bands; performing dynamic bit width allocation on the spectral feature difference to generate spatial continuity encoding data and frequency domain distribution encoding data; analyzing the spatial continuity encoding data and the frequency domain distribution encoding data by a large model to generate structured text description; verifying information coverage of the structured text description relative to the literature image dataset to generate evaluation results.

2. The method of claim 1, wherein, The method comprises: calculating an intensity ratio value of spatial continuity attributes and frequency domain distribution attributes in the spectral feature difference; determining a first target bit width value corresponding to the spatial continuity attributes and a second target bit width value corresponding to the frequency domain distribution attributes according to the intensity ratio value; calculating spectral intensity variation between adjacent pixel positions in the color distribution features to generate a spatial continuity variation set; calculating intensity fluctuation amplitude of different wavelength positions in the material reflection features to generate a frequency domain distribution fluctuation set; encoding the spatial continuity variation set according to the first target bit width value to generate spatial continuity encoding data; encoding the frequency domain distribution fluctuation set according to the second target bit width value to generate frequency domain distribution encoding data.

3. The method of claim 2, wherein, The method comprises: taking the intensity ratio value as a weight ratio value, multiplying the weight ratio value by a preset bit width base to generate a spatial continuity initial bit width value; scaling and adjusting the bit width base according to a fluctuation amplitude range of the frequency domain distribution attributes to generate a frequency domain distribution initial bit width value; correcting the frequency domain distribution initial bit width value according to intensity stability between different wavelength positions to generate a second target bit width value; correcting the spatial continuity initial bit width value according to intensity correlation between adjacent pixel positions to generate a first target bit width value.

4. The method of claim 1, wherein, The method comprises: dividing the literature image dataset into visible light band data and near-infrared band data according to spectral band range; identifying spectral response intensity distribution related to color attributes in the visible light band data to generate visible light spectral data carrying three primary color intensity information; Detecting a spectral absorption feature point related to material reflection characteristics in the near-infrared band data to generate near-infrared spectral data carrying specific wavelength reflection intensity information; Channel splicing the visible light spectral data and the near-infrared spectral data to form multi-channel spectral data containing color distribution characteristics and material reflection characteristics.

5. The method of claim 4, wherein, Detecting a spectral absorption feature point related to material reflection characteristics in the near-infrared band data to generate near-infrared spectral data carrying specific wavelength reflection intensity information, including: Detecting a spectral absorption feature point in the near-infrared band data whose reflection intensity is lower than that of adjacent wavelength positions; According to a pre-set material reflection characteristic library, matching the wavelength positions of the spectral absorption feature points to filter out target absorption points associated with material reflection characteristics; Extracting the reflection intensity values of the wavelength positions where the target absorption points are located to generate reflection intensity distribution data carrying wavelength intensity mapping relationships; Binding the reflection intensity distribution data with the spatial position information of the near-infrared band data to generate near-infrared spectral data carrying specific wavelength reflection intensity information.

6. The method of claim 1, wherein, According to a pre-set spectral sensitivity threshold, dividing the high-sensitivity band and the low-sensitivity band of the multi-channel spectral data, and calculating the spectral feature difference between the high-sensitivity band and the low-sensitivity band, including: Statistically analyzing the spectral intensity of each spectral channel in the multi-channel spectral data to generate a spectral intensity distribution curve corresponding to each spectral channel; Marking the band interval on the spectral intensity distribution curve where the spectral intensity change rate exceeds the pre-set spectral sensitivity threshold as the high-sensitivity band, and marking the band interval where the spectral intensity change rate is lower than the pre-set spectral sensitivity threshold as the low-sensitivity band; Extracting a first spectral intensity extreme point from the high-sensitivity band and a second spectral intensity extreme point from the low-sensitivity band, and calculating the absolute value of the intensity difference between the first spectral intensity extreme point and the second spectral intensity extreme point as the spectral feature difference.

7. The method of claim 1, wherein, Analyzing the spatial continuity encoding data and the frequency domain distribution encoding data through a large model to generate a structured text description, including: Analyzing the element position coordinate sequence in the spatial continuity encoding data to generate a spatial position description set of chart elements, which contains a first identifier and a first literature position label; Analyzing the correspondence between wavelength and reflection intensity in the frequency domain distribution encoding data to generate a material characteristic description set of formula symbols, which contains a second identifier and a second literature position label; Based on the consistency of the literature position labels, cross-modal matching the first identifier with the second identifier to generate an element symbol correlation mapping table; According to the element symbol correlation mapping table, retrieving a chart element position description from the spatial position description set and a formula symbol material description from the material characteristic description set; Combine the chart element location description associated with the same document location and the formula symbol material description into a composite text paragraph containing chart element topology description and formula symbol attribute description, integrate all composite text paragraphs corresponding to the document location, and generate a structured text description.

8. A large model-based literature report generation evaluation system, characterized by, The method comprises the following steps: An acquisition module is configured to acquire a document image dataset containing non-text elements, wherein the non-text elements include chart elements and formula symbols. A separation module is configured to separate different spectral band data of the document image dataset to generate multi-channel spectral data containing color distribution characteristics and material reflection characteristics. A calculation module is configured to divide high-sensitivity bands and low-sensitivity bands of the multi-channel spectral data according to a preset spectral sensitivity threshold, and calculate a spectral feature difference between the high-sensitivity bands and the low-sensitivity bands. An allocation module is configured to perform dynamic bit width allocation on the spectral feature difference to generate spatial continuity encoding data and frequency domain distribution encoding data. A generation module is configured to analyze the spatial continuity encoding data and the frequency domain distribution encoding data by using a large model to generate a structured text description. A verification module is configured to verify information coverage of the structured text description with respect to the document image dataset to generate an evaluation result.

9. A computing device, comprising: The method comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the method for generating and evaluating a document report based on a large model according to any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer program is stored in a computer and is executed by the computer to implement the method for generating and evaluating a document report based on a large model according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal literature data extraction method and device and medium

    CN119311880A

  • Open set cross-domain hyperspectral classification method and system based on multiple modes and optimal transmission

    CN119832325A