A retrieval enhancement generation method for terahertz absorption spectroscopy diagnosis

By constructing a terahertz absorption spectral structure database and a text vector library, and combining them with a large language model to generate answers, the efficiency and accuracy issues in terahertz absorption spectral diagnostic analysis were resolved, achieving more efficient diagnostic analysis.

CN122132506APending Publication Date: 2026-06-02PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIV
Filing Date
2026-02-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately pinpoint the specific influencing factors when terahertz absorption spectroscopy measurement results do not meet expectations. Traditional RAG methods fail to effectively utilize structured data, resulting in low efficiency in diagnostic analysis.

Method used

We construct a terahertz absorption spectral structure database and a text vector library, generate answers through a large language model, and combine structured data for query matching and reasoning to improve diagnostic efficiency.

Benefits of technology

By using structured data for querying, matching, and reasoning, and with strong logical connections, the efficiency of terahertz absorption spectroscopy diagnostic analysis is improved, content redundancy and matching bias are reduced, and diagnostic accuracy is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132506A_ABST
    Figure CN122132506A_ABST
Patent Text Reader

Abstract

This invention discloses a retrieval enhancement generation method for terahertz absorption spectroscopy diagnostics. The method includes: collecting terahertz absorption spectral data of biochemical substances; establishing a terahertz absorption spectral text structure database; establishing a terahertz absorption spectral image structure database; establishing a terahertz absorption spectral structure database; constructing a text vector library in the field of terahertz spectroscopy detection; determining the user question type and extracting structured information; if the user question is a terahertz absorption spectroscopy diagnostic question, then analyzing influencing factors and generating a query question, rewriting the user question into a query question; retrieving the text vector library in the field of terahertz spectroscopy detection based on the query question; generating an answer through a large language model and returning it. This method analyzes and calculates the influencing factors of terahertz absorption spectroscopy based on a structured database and rewrites the query question into a query question more suitable for RAG retrieval, improving the efficiency of researchers in diagnosing absorption spectra and retrieving knowledge in the field of terahertz spectroscopy detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of diagnostic analysis technology for terahertz spectroscopy detection, specifically to a search enhancement generation method for terahertz absorption spectroscopy diagnosis. Background Technology

[0002] Terahertz time-domain spectroscopy, with its unique physicochemical properties such as non-contact, non-destructive, fast response, and excellent "fingerprint" identification, exhibits significant advantages in the detection of biochemical substances. In particular, terahertz absorption spectroscopy can directly reflect the low-frequency vibrational and rotational modes of molecules, providing unique "fingerprint" information for the structural analysis and state diagnosis of biochemical substances. Research on terahertz absorption spectroscopy is crucial in numerous fields, including biomedicine, environmental monitoring, and security inspection. However, terahertz absorption spectroscopy measurements are easily affected by various experimental measurement conditions. When the actual measured absorption spectrum results do not meet expectations, researchers find it difficult to quickly and accurately pinpoint the specific influencing factors. Directly collecting data or reasoning through large linguistic models typically only lists various possible influencing factors, lacking strong logical connections and thus limited practical reference value, failing to meet the needs of terahertz absorption spectroscopy diagnostic analysis of specific substances under specific measurement conditions.

[0003] In recent years, as large language models have demonstrated powerful generation and reasoning capabilities in general natural language processing tasks, they also possess potential application value in knowledge-based question answering within specialized fields. Retrieval-enhanced generation (RAG) technology, incorporating expertise from the terahertz spectroscopy detection field, can address the aforementioned issues to some extent. However, for diagnostic analysis of terahertz absorption spectra of specific substances under specific measurement conditions, users often raise highly targeted questions. Traditional RAGs, primarily based on semantic vectors for similarity-based recall, fail to utilize the structured data information already present in resource data for precise conditional filtering, leading to redundant or mismatched returned content and the introduction of irrelevant or low-quality textual knowledge, thus limiting their practical application in terahertz absorption spectroscopy diagnostics.

[0004] In summary, there is an urgent need for a retrieval enhancement generation method for terahertz absorption spectroscopy diagnosis, which can realize the analysis of key factors for diagnostic problems in terahertz absorption spectroscopy, thereby improving the efficiency of researchers in diagnosing absorption spectral differences and retrieving knowledge in the field of terahertz spectroscopy detection, and promoting the in-depth application of artificial intelligence in the field of terahertz spectroscopy detection. Summary of the Invention

[0005] To address the problems existing in the prior art, this invention proposes a retrieval enhancement generation method for terahertz absorption spectroscopy diagnosis. This method constructs a terahertz absorption spectral structure database and a text vector library in the field of terahertz spectroscopy detection. It queries and matches diagnostic questions related to terahertz absorption spectroscopy with data from the terahertz absorption spectral structure database, calculates and infers the key experimental measurement factors affecting terahertz absorption spectra, and rewrites the diagnostic questions into query questions more suitable for RAG retrieval. Based on the query questions, it retrieves information from the text vector library in the field of terahertz spectroscopy detection, generates answers using a large language model, and returns them. This method can improve the efficiency of researchers in diagnosing absorption spectral differences and retrieving knowledge in the field of terahertz spectroscopy detection.

[0006] The technical solution of the present invention is as follows:

[0007] A retrieval enhancement generation method for terahertz absorption spectroscopy diagnostics, characterized by comprising the following steps:

[0008] Step S1: Collect relevant data and literature on terahertz absorption spectra of different biochemical substances, analyze the text elements in the data, construct a terahertz absorption spectrum structured information extraction prompt word template, and obtain the "substance name-measurement conditions-absorption spectrum" relationship pair through large language model reasoning to establish a terahertz absorption spectrum text structure database; the "absorption spectrum" in the database includes multiple absorption peak positions within the preset absorption peak frequency threshold range, and the "measurement conditions" include at least the measurement equipment, sample state, temperature, humidity and environment, and the data source literature is marked;

[0009] Step S2: Analyze the absorption spectra in the data and calculate the "substance name-absorption spectrum" relationship pair to establish a terahertz absorption spectrum image structure database;

[0010] Step S3: Enhance the verification of the data in the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database to obtain the "substance name-measurement conditions-absorption spectrum" relationship pair and establish the terahertz absorption spectrum structure database.

[0011] Step S4: Collect and process relevant data in the field of terahertz spectroscopy detection, and establish a text vector library in the field of terahertz spectroscopy detection.

[0012] Step S5: Receive the user's input natural language question and determine whether the user's question is a terahertz absorption spectroscopy diagnostic question. If so, extract the "substance name-measurement conditions-absorption spectrum" field and value from the user's question as structured data by combining the terahertz absorption spectroscopy structured information extraction prompt word template. If not, skip step S6 and further process the user's question in step S7.

[0013] Step S6: Match the data obtained in step S5 with the terahertz absorption spectral structure database obtained in step S3, perform measurement condition comparison and absorption peak frequency difference calculation, and generate a terahertz absorption spectral diagnostic query question containing key difference measurement conditions based on the calculation results.

[0014] Step S7: Based on the user's original question that is determined to be a non-terahertz absorption spectroscopy diagnostic question in step S5, or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S6, the user's original question is retrieved from the terahertz spectroscopy detection field text vector library obtained in step S4. The results and the question are then input into the large language model to generate an answer and return it to the user.

[0015] Furthermore, step S1 specifically includes:

[0016] Step S11: Collect data related to the terahertz absorption spectrum of biochemical substances, and parse the file using a fast parsing strategy based on the internal text object tree and layout features to obtain the processed text elements.

[0017] Step S12: Construct a prompt word template for extracting structured information from terahertz absorption spectra;

[0018] Step S13: Input the text elements and prompt word templates obtained in steps S11 and S12 into the locally deployed large language model to generate structured fields and corresponding values.

[0019] Step S14: The structured data obtained in step S13 is cleaned and verified. The absorption peak frequency field and value are converted and range verified. A threshold is set to extract the absorption peak frequency within a specific frequency range and the number of absorption peaks within a specific upper limit threshold.

[0020] Step S15: Map the data obtained in step S14 to the database, store the data according to "substance name-measurement conditions-absorption spectrum", and mark the source literature of the data to establish a terahertz absorption spectrum text structure database.

[0021] Furthermore, step S2 specifically includes:

[0022] Step S21: For the terahertz absorption spectrum related data collected in step S1, obtain the image elements in the data and extract the nearby text as the image title.

[0023] Step S22: For the image and corresponding title obtained in step S21, determine whether the title contains a substance name by using a hard-coded list of common biochemical substance names and a set regular expression, and determine whether the title contains an absorption spectrum keyword by using keyword matching. If and only if the title contains both a substance name and an absorption spectrum keyword, the image continues to be processed; otherwise, the processing of the image is skipped.

[0024] Step S23: For images that can be further processed in step S22, convert the images to grayscale images, perform adaptive threshold binarization, remove noise through morphological operations, and define the target region of the absorption spectrum.

[0025] Step S24: For the target region of the absorption spectrum map defined in step S23, the terahertz absorption spectrum curve contained in the region is found by converting the region image to the HSV color space and creating a color mask.

[0026] Step S25: Calculate the slope of the frequency points of the curve obtained in step S24 to obtain the absorption peak frequency data within the absorption peak frequency threshold range set in step S1.

[0027] Step S26: Integrate the data obtained in steps S21-S25, map it to the database, store the data according to "substance name-absorption spectrum", and establish a terahertz absorption spectrum image structure database.

[0028] Furthermore, step S3 specifically includes:

[0029] Step S31: Establish a terahertz absorption spectral structure database, storing fields in the format of "substance name-measurement conditions-absorption spectrum";

[0030] Step S32: Traverse the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database with the same data source literature, and write the data into the terahertz absorption spectrum structure database according to the writing rules. The writing rules are as follows: when the data source literature is the same and the substance name exists in both the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database, compare the absorption peak frequencies of the substance in the two databases through a double loop, calculate the difference for each absorption peak, and if all absorption peak frequencies of the substance in the terahertz absorption spectrum text structure database are empty, or if at least one pair of absorption peak frequencies have a difference less than the set tolerance threshold, then the image processing data enhancement verification is considered successful, and the data units of the substance in the two databases are merged and written; if the verification fails, neither data is written; repeat this process until the traversal is completed, and the writing of the terahertz absorption spectrum structure database is completed.

[0031] Furthermore, step S4 specifically includes:

[0032] Step S41: Collect data in the field of terahertz spectroscopy detection, parse the file using a high-resolution parsing strategy based on the internal text object tree and layout features, call the target detection model to identify the title, paragraph, table and image areas in the text, split the page elements into a set of structured elements, and obtain the continuous text, title, table content and other text and corresponding metadata.

[0033] Step S42: The data obtained in step S41 is divided into blocks using a regular expression-based Chinese recursive segmentation algorithm, and each text block is encoded using a pre-trained Chinese vector model. Each text block is mapped to a high-dimensional dense vector to form a structured text vector library for the terahertz spectroscopy detection field.

[0034] Furthermore, step S5 specifically includes:

[0035] Step S51: Receive the natural language question input by the user, and determine whether the user's question is a diagnostic question for terahertz absorption spectroscopy by keyword matching; if not, it is treated as the user's original question and further processed in step S7.

[0036] Step S52: For user questions that are determined to be diagnostic problems of terahertz absorption spectroscopy, the user questions and the terahertz absorption spectroscopy structured information extraction prompt word template constructed in step S1 are sent to the locally deployed large language model to extract the structured "substance name-measurement conditions-absorption spectrum" fields and their corresponding values.

[0037] Furthermore, step S6 specifically includes:

[0038] Step S61: Connect to the database via an asynchronous session and query the terahertz absorption spectral structure database obtained in step S3 for all data unit information that is the same as the substance name data obtained in step S5;

[0039] Step S62: For the queried material data unit information, the measurement condition data in the user structured information obtained in step S5 is compared with the measurement condition data of the material data unit. When the corresponding values ​​of all measurement condition fields of the user are exactly the same as the corresponding values ​​of all measurement condition fields of a certain data of the material data unit, or when the corresponding values ​​of all measurement condition fields of the user are only different from the corresponding values ​​of all measurement condition fields of a certain data of the material data unit, the difference between the user's absorption peak frequency and the absorption peak frequency in the data is calculated. The calculation method is as described in step S63.

[0040] Step S63: Iterate through the values ​​corresponding to the absorption peak frequency field of each data in the terahertz absorption spectral structure database that needs to be calculated in step S62, collect them into a list, calculate the absolute difference between the user's absorption peak frequency data and the closest absorption peak frequency in the list, calculate the average value and use it as the difference degree between the record and the user's absorption peak frequency.

[0041] Step S64: Sort the differences obtained in step S63 from largest to smallest, and take the field names that are different from the corresponding values ​​of the user's measurement condition field in the first n data as influencing factors; n is a custom setting.

[0042] Step S65: Based on the user structured information data obtained in step S5 and the influencing factors obtained in step S64, rewrite the user question and construct a query question.

[0043] Furthermore, step S7 specifically includes:

[0044] Step S71: Call the pre-trained Chinese vector model to encode the user's original question, which was determined to be a non-terahertz absorption spectroscopy diagnostic question in step S5, or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S6, into a vector, and filter the Top-K text segments as candidate text blocks through similarity calculation.

[0045] Step S72: Input the user's original question / query question and the candidate text block obtained in step S71 into the constructed terahertz spectroscopy detection domain prompt word template, and send the question, candidate text block and prompt word template into the locally deployed large language model to generate the answer and return it.

[0046] The technical effects of this invention are as follows:

[0047] This invention constructs a terahertz absorption spectral structure database, storing various experimental measurement factors affecting the terahertz absorption spectra of biochemical substances in structured data form. By extracting structured information from terahertz absorption spectroscopy diagnostic questions and querying and matching it with data in the terahertz absorption spectral structure database, the key experimental measurement factors affecting terahertz absorption spectra can be inferred by calculating the differences between absorption spectral frequencies. Compared to directly collecting data or reasoning through large language models, the influencing factors inferred by this invention have stronger logical connections and are more valuable for practical reference. Furthermore, after obtaining the key experimental measurement factors, this invention rewrites the terahertz absorption spectroscopy diagnostic questions into query questions more suitable for RAG retrieval. Compared to traditional RAGs, this effectively avoids the problem of introducing irrelevant or low-quality text knowledge due to content redundancy or matching deviations, thereby improving the efficiency of researchers in diagnosing absorption spectral differences and retrieving knowledge in the field of terahertz spectroscopy detection. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the overall framework and process provided in the embodiments of the present invention;

[0049] Figure 2 This is a schematic diagram of the terahertz absorption spectrum structured information extraction prompt word template provided in an embodiment of the present invention;

[0050] Figure 3 This is a prompt word template for the field of terahertz spectroscopy detection provided in this embodiment of the invention. Detailed Implementation

[0051] The present invention will be further clearly and completely described below with reference to the accompanying drawings and specific embodiments.

[0052] This invention proposes a retrieval enhancement generation method for terahertz absorption spectroscopy diagnostics, specifically, as follows: Figure 1 As shown, it illustrates the overall framework and process diagram provided by an embodiment of the present invention, including the following steps:

[0053] Step S1: Collect relevant data and literature on terahertz absorption spectra of different biochemical substances, analyze the text in the data, and obtain the relationship pairs of "substance name-measurement conditions-absorption spectrum" through large language model reasoning, and establish a terahertz absorption spectrum text structure database.

[0054] Step S11: Collect data related to the terahertz absorption spectrum of biochemical substances. This mainly involves processing Chinese PDF documents. A fast parsing strategy based on the text object tree and layout features within the PDF is adopted. The PDF file is parsed by calling the partition_pdf() function of the unstructured.partition.pdf library, which divides the PDF into pages and text blocks and returns the text elements.

[0055] Step S12, construct a prompt word template for extracting structured information from terahertz absorption spectra, such as... Figure 2 As shown, the prompt first defines the role as an information extraction assistant in the field of terahertz spectroscopy detection. The seven fields to be extracted are: substance name, measuring equipment, sample status, temperature, humidity, environment, and absorption peak frequency. Specific extraction rules are marked for each field. Finally, it emphasizes that the output results should be in JSON format.

[0056] Step S13: The text elements and prompt word templates obtained in steps S11 and S12 are sent to the locally deployed ChatGLM3-6B model (using FastChat's OpenAI-compatible interface to interact with the model). After the model returns a response, the object is located by finding the first "{" symbol and the last "}" symbol in the returned string, thereby forcibly extracting the JSON content and parsing it to obtain the structured fields and their corresponding values.

[0057] Step S14: Perform data cleaning and verification on the structured data parsed in step S13. Integrate the structured data obtained from each document. For the fields of measurement equipment, sample status, temperature, humidity, and environment, if all fields have no corresponding values ​​or any field has different values ​​(this method only considers the case where each document maintains consistent measurement conditions), then skip that document (it will not be included in the subsequent terahertz absorption spectrum text structure database). For the absorption peak frequency field and value, perform numerical conversion and range verification, retain only the frequency values ​​in the range of 0-3.5THz, and check whether the structured data fragment contains both the substance name and the absorption peak frequency. Only retain it if it contains both (same as above). After merging, deduplicating, and sorting the absorption peak frequencies of the same substance, extract a maximum of the first 15 frequency values ​​(other values ​​can be used; here, 15 is selected as the extraction threshold example based on the number of terahertz absorption peaks of most biochemical substances).

[0058] Step S15: The data obtained in S14 is mapped to the terahertz absorption spectroscopy text structure database content_thz_spectrum using the SQLAlchemy ORM framework. The stored fields are: substance name, measuring device, sample state, temperature, humidity, environment, 0-3.5THz absorption peak position 1, 0-3.5THz absorption peak position 2, ..., 0-3.5THz absorption peak position 15. The substance name field cannot be empty, and the measuring device, sample state, temperature, humidity, and environment fields cannot all be empty at the same time. The 0-3.5THz absorption peak position x (1-15) can all be empty at the same time. If the number of absorption peaks is less than 15, the subsequent columns are set to empty. In addition, the source literature of the data is marked. Finally, the data obtained from text processing is directly and persistently stored in the MySQL database.

[0059] Step S2: For the data collected in step S1, analyze the absorption spectra in the data and calculate the "substance name-absorption spectrum" relationship pair to establish a terahertz absorption spectrum image structure database.

[0060] Step S21: For the data (Chinese PDF format documents) collected in step S1, use the PyMuPDF library to traverse each page to obtain all image elements. For each image, extract the text within 50 characters nearby as the title of the image.

[0061] Step S22: For the title obtained in step S21, the title is matched to see if it contains a substance name by using a hard-coded list of 1583 common biochemical substance names and a set regular expression (such as "XX acid", "XX amino acid", "XX sugar", "XX molecule"). Absorption spectrum keywords ("absorption spectrum", "absorption coefficient", "absorption") are set. The image is processed only if it contains both a substance name and an absorption spectrum keyword. Otherwise, the processing of the image is skipped.

[0062] Step S23: For images that can be further processed in step S22, use OpenCV for image processing, convert the image to grayscale, use adaptive threshold binarization and remove noise through morphological operations, find the contour and select the largest rectangular area as the target area of ​​the absorption spectrum, and filter out areas with a width or height of less than 100 pixels.

[0063] Step S24: For the defined target area of ​​the absorption spectrum, convert the area image to the HSV color space and create a color mask. Scan the mask column by column to find the curve points in each column (here we only consider the case where the area contains only one curve). Convert the pixel coordinates to 0-3.5THz frequency coordinates and 0-1 absorption coefficient coordinates.

[0064] Step S25: Perform Savitzky-Golay smoothing on the curve, calculate the slope of each curve point in units of 0.01THz, find the point where the slope changes from positive to negative and exceeds the threshold (0.1thz), verify whether it is a local maximum, retain the absorption peak in the range of 0-3.5THz, and retain a maximum of 15 peaks after deduplication.

[0065] Step S26: The data obtained in steps S21-S25 is mapped to the terahertz absorption spectrum image structure database picture_thz_spectrum using the SQLAlchemy ORM framework. The storage fields are: substance name, 0-3.5THz absorption peak position 1, 0-3.5THz absorption peak position 2, ..., 0-3.5THz absorption peak position 15. The substance name field cannot be empty, and the 0-3.5THz absorption peak positions x (1-15) cannot all be empty at the same time. If the number of absorption peaks is less than 15, the subsequent columns are set to empty. In addition, the source literature of the data is marked. Finally, the data obtained from image processing is directly and persistently stored in the MySQL database.

[0066] Step S3: Enhance and verify the database data obtained in steps S1 and S2 to obtain the "substance name-measurement conditions-absorption spectrum" relationship pair, and establish a terahertz absorption spectrum structure database.

[0067] Step S31: Create a new terahertz absorption spectral structure database thz_spectrum, whose stored fields are: substance name, measuring device, sample state, temperature, humidity, environment, 0-3.5THz absorption peak position 1, 0-3.5THz absorption peak position 2, ..., 0-3.5THz absorption peak position 15;

[0068] Step S32: When the data source literature is the same, if the corresponding value of the substance name field exists in both the content_thz_spectrum and picture_thz_spectrum databases, a double loop is used to compare the absorption peak frequencies of the substance in the content_thz_spectrum and picture_thz_spectrum databases. For each absorption peak (not indexed by position), the difference is calculated. If the absorption peak frequencies of the substance in the content_thz_spectrum database are all empty or the difference between at least one pair of absorption peaks is less than the tolerance threshold (0.1 THz), then the graph processing data enhancement verification is considered successful, and the data units of the substance in the content_thz_spectrum and picture_thz_spectrum databases are updated. All information is written to the thz_spectrum database; otherwise, the image processing data enhancement verification is considered to have failed. The information of the data unit of the substance in the content_thz_spectrum and picture_thz_spectrum databases is not written to the thz_spectrum database. When the data unit information of the picture_thz_spectrum database is written to the thz_spectrum database, the field values ​​of measuring equipment, sample status, temperature, humidity, environment, etc. are filled with the field values ​​of the corresponding data unit information of the content_thz_spectrum database, which are the same as the data source literature and the corresponding substance name (for the reasons described in step S14). This continues until all data source literature is processed and the information in the thz_spectrum database is written.

[0069] Step S4: Collect and process relevant data in the field of terahertz spectroscopy detection, and establish a text vector library in the field of terahertz spectroscopy detection.

[0070] Step S41: Collect Chinese PDF documents in the field of terahertz spectroscopy detection, call the unstructured.partition.pdf library to perform layout analysis and content extraction on the original documents, adopt a high-resolution (hi-res) visual segmentation strategy, call the object detection model (yolox) to identify the title, paragraph, table and image areas in the text, split the layout elements into a set of structured elements, and obtain the continuous text, title, table content and other text and corresponding metadata;

[0071] Step S42: The data obtained in step S41 is divided into blocks using a Chinese recursive segmentation algorithm based on regular expressions. Each text block is encoded using a pre-trained Chinese vector model bge-large-zh-v1.5, and each text block is mapped to a high-dimensional dense vector to form structured vectorized knowledge entries. A corresponding vector library instance is constructed using a vector index structure based on FAISS, and the "text content - embedded vector - metadata" triples are added to the index in batches.

[0072] Step S5: Receive the natural language question input by the user, determine whether the user question is a diagnostic question of terahertz absorption spectroscopy, and if so, extract the "substance name-measurement conditions-absorption spectrum" field and value from the user question as structured data.

[0073] Step S51: Receive the user's input in natural language and define two sets of keywords: [Terahertz absorption spectroscopy keywords: "terahertz", "THz", "thz", "absorption spectrum", "absorption peak", "spectral peak", "spectral line". Diagnostic keywords: "cause", "impact", "diagnosis", "abnormal", "high", "low", "different", "difference"]. If the user's question contains both terahertz absorption spectroscopy keywords and diagnostic keywords, it is determined to be a terahertz absorption spectroscopy diagnostic question; otherwise, proceed directly to step S7, where the user's original question is further processed.

[0074] Step S52: For user questions that are determined to be diagnostic problems of terahertz absorption spectroscopy, the user question and the terahertz absorption spectroscopy structured information extraction prompt template constructed in step S1 are sent to the locally deployed ChatGLM3-6B model to extract the following structured fields and values: substance name, measuring device, sample state, temperature, humidity, environment, and absorption peak frequency. This method only considers the case where all field values ​​are not empty. When all field values ​​are not empty, step S6 is performed.

[0075] Step S6: Match the data obtained in step S5 with the terahertz absorption spectral structure database obtained in step S3, perform measurement condition comparison and absorption peak frequency difference calculation, and generate a diagnostic query question containing key difference measurement conditions based on the calculation results.

[0076] Step S61: Connect to the database via AsyncSessionLocal asynchronous session, and query the terahertz absorption spectral structure database thz_spectrum obtained in step S3 for all data unit information with the same substance name as obtained in step S52;

[0077] Step S62: For each piece of data in the queried material data unit, perform the following operations: compare the measurement conditions in the user structured information obtained in step S52 with the measurement conditions of the material data unit. When all the corresponding values ​​of the user's measurement condition fields are exactly the same as all the corresponding values ​​of the measurement condition fields of a certain piece of data in the material data unit, or when all the corresponding values ​​of the user's measurement condition fields are different from only one of the corresponding values ​​of the measurement condition fields of a certain piece of data in the material data unit, calculate the difference between the user's absorption peak frequency and the absorption peak frequency in that piece of data (the method is as described in step S63).

[0078] Step S63: Iterate through the values ​​corresponding to the 0-3.5THz absorption peak 1, 0-3.5THz absorption peak 2, ..., 0-3.5THz absorption peak 15 fields of each data in the thz_spectrum database that needs to be calculated in step S62, collect them into a list, find the closest peak in the list for the user's absorption peak frequency, calculate the absolute difference, and calculate the average of these absolute differences as the difference between the record and the user's absorption peak frequency peak_distance;

[0079] Step S64: Sort the peak_distance obtained in step S63 from largest to smallest (the greater the difference, the more significant the impact of this condition change may be). Take the field names that are different from the corresponding values ​​of the measurement condition fields of the first 5 data points as influencing factors (when all the corresponding values ​​of the measurement condition fields of a certain data point are exactly the same as all the corresponding values ​​of the user's measurement condition fields, then the influencing factors are factors other than measurement equipment, sample state, temperature, humidity, and environmental factors, and deduplication is performed).

[0080] Step S65: Based on the user structured information obtained in step S52 and the influencing factors obtained in step S64, rewrite the user question and construct a query question: [When {influencing factors} are different / high / low / large / small / change, what impact will the terahertz absorption spectrum of {substance name} be affected?]

[0081] Step S7: Based on the user's original question that is determined to be a non-terahertz absorption spectroscopy diagnostic question in step S5, or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S6, the user's original question is searched in the terahertz spectroscopy detection field text vector library obtained in step S4, and the results and questions are input into the large language model to generate an answer and return it to the user.

[0082] Step S71: Call the same Chinese vector model bge-large-zh-v1.5 as in step S42 to encode the user's original question that was determined to be a non-terahertz absorption spectroscopy diagnostic question in step S51 or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S65 into a vector. Input the FAISS vector index obtained in step S4 to perform cosine similarity calculation, and select the Top-K text segments with the highest similarity as candidate text blocks.

[0083] Step S72: Input the user's original question / query question and the candidate text block obtained in step S71 into the constructed terahertz spectroscopy detection domain prompt word template, such as... Figure 3 As shown, the template sets the respondent role as an "expert in the field of terahertz spectroscopy detection," requiring a professional, concise, and unembellished response based on the given context. The question, candidate text blocks, and prompt word template are fed into the locally deployed ChatGLM3-6B model to generate the answer and return it.

[0084] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art will understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection of the present invention is defined by the scope of the claims.

Claims

1. A method for search enhancement generation for terahertz absorption spectroscopy diagnostics, characterized in that, Includes the following steps: Step S1: Collect relevant data and literature on terahertz absorption spectra of different biochemical substances, analyze the text elements in the data, construct a terahertz absorption spectrum structured information extraction prompt word template, and obtain the "substance name-measurement conditions-absorption spectrum" relationship pair through large language model reasoning to establish a terahertz absorption spectrum text structure database; the "absorption spectrum" in the database includes multiple absorption peak positions within the preset absorption peak frequency threshold range, and the "measurement conditions" include at least the measurement equipment, sample state, temperature, humidity and environment, and the data source literature is marked; Step S2: Analyze the absorption spectra in the data and calculate the "substance name-absorption spectrum" relationship pair to establish a terahertz absorption spectrum image structure database; Step S3: Enhance the verification of the data in the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database to obtain the "substance name-measurement conditions-absorption spectrum" relationship pair and establish the terahertz absorption spectrum structure database. Step S4: Collect and process relevant data in the field of terahertz spectroscopy detection, and establish a text vector library in the field of terahertz spectroscopy detection. Step S5: Receive the user's input natural language question and determine whether the user's question is a diagnostic question for terahertz absorption spectroscopy. If so, extract the "substance name-measurement conditions-absorption spectrum" field and value from the user's question as structured data by combining the terahertz absorption spectroscopy structured information extraction prompt word template. If not, skip step S6 and further process the user's question in step S7. Step S6: Match the data obtained in step S5 with the terahertz absorption spectral structure database obtained in step S3, perform measurement condition comparison and absorption peak frequency difference calculation, and generate a terahertz absorption spectral diagnostic query question containing key difference measurement conditions based on the calculation results. Step S7: Based on the user's original question that is determined to be a non-terahertz absorption spectroscopy diagnostic question in step S5, or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S6, the user's original question is retrieved from the terahertz spectroscopy detection field text vector library obtained in step S4. The results and the question are then input into the large language model to generate an answer and return it to the user.

2. The method as described in claim 1, characterized in that, Step S1 specifically includes: Step S11: Collect data related to the terahertz absorption spectrum of biochemical substances, and parse the file using a fast parsing strategy based on the internal text object tree and layout features to obtain the processed text elements. Step S12: Construct a prompt word template for extracting structured information from terahertz absorption spectra; Step S13: Input the text elements and prompt word templates obtained in steps S11 and S12 into the locally deployed large language model to generate structured fields and corresponding values. Step S14: The structured data obtained in step S13 is cleaned and verified. The absorption peak frequency field and value are converted and range verified. A threshold is set to extract the absorption peak frequency within a specific frequency range and the number of absorption peaks within a specific upper limit threshold. Step S15: Map the data obtained in step S14 to the database, store the data according to "substance name-measurement conditions-absorption spectrum", and mark the source literature of the data to establish a terahertz absorption spectrum text structure database.

3. The method as described in claim 1, characterized in that, Step S2 specifically includes: Step S21: For the terahertz absorption spectrum related data collected in step S1, obtain the image elements in the data and extract the nearby text as the image title. Step S22: For the image and corresponding title obtained in step S21, determine whether the title contains a substance name by using a hard-coded list of common biochemical substance names and a set regular expression, and determine whether the title contains an absorption spectrum keyword by using keyword matching. If and only if the title contains both a substance name and an absorption spectrum keyword, the image continues to be processed; otherwise, the processing of the image is skipped. Step S23: For images that can be further processed in step S22, convert the images to grayscale images, perform adaptive threshold binarization, remove noise through morphological operations, and define the target region of the absorption spectrum. Step S24: For the target region of the absorption spectrum map defined in step S23, the terahertz absorption spectrum curve contained in the region is found by converting the region image to the HSV color space and creating a color mask. Step S25: Calculate the slope of the frequency points of the curve obtained in step S24 to obtain the absorption peak frequency data within the absorption peak frequency threshold range set in step S1. Step S26: Integrate the data obtained in steps S21-S25, map it to the database, store the data according to "substance name-absorption spectrum", and establish a terahertz absorption spectrum image structure database.

4. The method as described in claim 1, characterized in that, Step S3 specifically includes: Step S31: Establish a terahertz absorption spectrum structure database, storing fields in the format of "substance name-measurement conditions-absorption spectrum"; Step S32: Traverse the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database with the same data source literature, and write the data into the terahertz absorption spectrum structure database according to the writing rules. The writing rules are as follows: when the data source literature is the same and the substance name exists in both the terahertz absorption spectrum text structure database and the terahertz absorption spectrum image structure database, compare the absorption peak frequencies of the substance in the two databases through a double loop, calculate the difference for each absorption peak, and if all absorption peak frequencies of the substance in the terahertz absorption spectrum text structure database are empty, or if at least one pair of absorption peak frequencies have a difference less than the set tolerance threshold, then the image processing data enhancement verification is considered successful, and the data units of the substance in the two databases are merged and written. If the verification fails, neither of the two sets of data will be written; repeat this process until the traversal is complete, thus completing the writing of the terahertz absorption spectral structure database.

5. The method as described in claim 1, characterized in that, Step S4 specifically includes: Step S41: Collect data in the field of terahertz spectroscopy detection, parse the file using a high-resolution parsing strategy based on the internal text object tree and layout features, call the target detection model to identify the title, paragraph, table and image areas in the text, split the page elements into a set of structured elements, and obtain the text and corresponding metadata of continuous body text, title, table content; Step S42: The data obtained in step S41 is divided into blocks using a regular expression-based Chinese recursive segmentation algorithm, and each text block is encoded using a pre-trained Chinese vector model. Each text block is mapped to a high-dimensional dense vector to form a structured text vector library for the terahertz spectroscopy detection field.

6. The method as described in claim 1, characterized in that, Step S5 specifically includes: Step S51: Receive the natural language question input by the user, and determine whether the user's question is a diagnostic question for terahertz absorption spectroscopy by keyword matching; if not, it is treated as the user's original question and further processed in step S7. Step S52: For user questions that are determined to be diagnostic problems of terahertz absorption spectroscopy, the user questions and the terahertz absorption spectroscopy structured information extraction prompt word template constructed in step S1 are sent to the locally deployed large language model to extract the structured "substance name-measurement conditions-absorption spectrum" fields and their corresponding values.

7. The method as described in claim 1, characterized in that, Step S6 specifically includes: Step S61: Connect to the database via an asynchronous session and query the terahertz absorption spectral structure database obtained in step S3 for all data unit information that is the same as the substance name data obtained in step S5; Step S62: For the queried material data unit information, the measurement condition data in the user structured information obtained in step S5 is compared with the measurement condition data of the material data unit. When the corresponding values ​​of all measurement condition fields of the user are exactly the same as the corresponding values ​​of all measurement condition fields of a certain data of the material data unit, or when the corresponding values ​​of all measurement condition fields of the user are only different from the corresponding values ​​of all measurement condition fields of a certain data of the material data unit, the difference between the user's absorption peak frequency and the absorption peak frequency in the data is calculated. The calculation method is as described in step S63. Step S63: Iterate through the values ​​corresponding to the absorption peak frequency field of each data in the terahertz absorption spectral structure database that needs to be calculated in step S62, collect them into a list, calculate the absolute difference between the user's absorption peak frequency data and the closest absorption peak frequency in the list, calculate the average value and use it as the difference degree between the record and the user's absorption peak frequency. Step S64: Sort the differences obtained in step S63 from largest to smallest, and take the field names that are different from the corresponding values ​​of the user's measurement condition field in the first n data as influencing factors; n is a custom setting. Step S65: Based on the user structured information data obtained in step S5 and the influencing factors obtained in step S64, rewrite the user question and construct a query question.

8. The method as described in claim 1, characterized in that, Step S7 specifically includes: Step S71: Call the pre-trained Chinese vector model to encode the user's original question, which was determined to be a non-terahertz absorption spectroscopy diagnostic question in step S5, or the newly generated terahertz absorption spectroscopy diagnostic query question obtained in step S6, into a vector, and filter the Top-K text segments as candidate text blocks through similarity calculation. Step S72: Input the user's original question / query question and the candidate text block obtained in step S71 into the constructed terahertz spectroscopy detection domain prompt word template, and send the question, candidate text block and prompt word template into the locally deployed large language model to generate the answer and return it.