Water ecology literature content analysis method and system based on multi-source data fusion and large language model

Through the water ecological literature content analysis method based on multi-source data fusion and large language model, the problems of low efficiency and poor effectiveness of ecological literature analysis in the existing technology are solved, and efficient and accurate analysis of ecological related literature is achieved.

CN120163141APending Publication Date: 2025-06-17GUANGDONG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510327520.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-17
Filing Date
2025-03-19
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to conduct efficient and accurate analysis of literature content in ecology-related fields, resulting in low analysis efficiency and poor analysis effect.

Method used

The water ecological literature content analysis method based on multi-source data fusion and large language model is adopted. By obtaining the literature to be analyzed for pre-processing, the large language model is input to obtain the domain label and confidence score, the literature quality score is calculated based on the citation amount, data integrity and authoritative source of the literature, and species name and coordinate information are extracted.

Benefits of technology

Accurate analysis of ecology-related literature has been achieved, analysis efficiency has been improved, and detailed content of ecological concerns can be extracted in a targeted manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163141A_ABST
    Figure CN120163141A_ABST
Patent Text Reader

Abstract

The invention discloses a water ecology literature content analysis method and system based on multi-source data fusion and a large language model, and the method comprises the steps: obtaining a to-be-analyzed literature, and carrying out the preprocessing of the to-be-analyzed literature, and obtaining a preprocessed text; inputting the preprocessed text into a first large language model to obtain a domain label and a confidence score; if the domain label is a preset domain, the confidence score is greater than a preset confidence score; if so, calculating a literature quality score according to the quoted quantity, data integrity and authoritative originality of the literature to be analyzed; judging whether the literature quality score is greater than a preset literature quality score or not; and if the literature quality score is greater than a preset literature quality score, extracting species names and corresponding coordinate information and GIS information according to the literature to be analyzed. The literature content analysis method has the characteristics of accurate analysis and high analysis efficiency aiming at ecology-related literatures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data processing and natural language processing, and more specifically, to a method and system for analyzing the content of water ecological literature based on multi-source data fusion and large language models. Background Art

[0002] Current ecological research relies on the integration of multi-source heterogeneous data, including academic literature, news, and geospatial data.

[0003] Existing technologies often use general-purpose models for analysis, making it difficult to specifically analyze the content of literature in the field of ecology. Ecology pays more attention to the study of species distribution and geographical coordinates, and existing technologies are difficult to specifically analyze the detailed content of concern in ecology. The efficiency of literature analysis in related fields of ecology is low, and the analysis effect is poor.

[0004] The existing technology discloses a method and device for analyzing the theme content of literature. The method includes: obtaining a plurality of documents to be analyzed in a target field; obtaining the theme words under each theme output by the theme word extraction model, the extended phrases of the theme words under each theme, and the speech act annotation information of the abstracts of each document to be analyzed output by the speech act annotation model; generating analysis texts of the documents to be analyzed under each theme based on the extended phrases of the theme words under each theme and the speech act annotation information of the abstracts of each document to be analyzed. This method does not specifically analyze the content of ecological literature. Summary of the Invention

[0005] In view of the defects of the existing literature content analysis method in the related art, such as low efficiency and poor effect in analyzing the literature in the field of ecology, the present invention provides a method and system for analyzing the content of water ecological literature based on multi-source data fusion and large language models. This literature content analysis method has the characteristics of accurate analysis and high analysis efficiency for ecological related literature.

[0006] The primary object of the present invention is to solve the above technical problems, and the technical solution of the present invention is as follows:

[0007] A method for analyzing the content of water ecological literature based on multi-source data fusion and large language models includes:

[0008] S1: Obtain the documents to be analyzed and perform preprocessing to obtain the preprocessed text;

[0009] S2: Input the preprocessed text into the first large language model to obtain domain labels and confidence scores;

[0010] S3: If the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis;

[0011] S4: Calculate the literature quality score according to the citation volume, data integrity, and authoritative source of the literature to be analyzed.

[0012] S5: Determine whether the literature quality score is greater than the preset literature quality score; if the literature quality score is greater than the preset literature quality score, then execute step S6; otherwise, terminate the analysis.

[0013] S6: Extract the species name and the corresponding coordinate information and GIS information according to the literature to be analyzed.

[0014] Furthermore, the literature quality score formula in step S4 is as follows:

[0015] Score = α·Citations + β·Completeness + γ·Authority

[0016] α, β, and γ represent adjustable weight coefficients, Citations represents the citation volume, Completeness represents the data integrity, and Authority represents the authoritative source.

[0017] Furthermore, the calculation formula for the data integrity is as follows:

[0018]

[0019] S represents whether the species name exists, G represents whether the geographical coordinates exist, and T represents whether the time information exists.

[0020] Furthermore, the calculation formula for the citation volume is as follows:

[0021]

[0022] C represents the actual citation times of the current literature, and Cmax represents the highest citation times of the literature in the target field.

[0023] Furthermore, the formula for the authoritative source is as follows:

[0024]

[0025] IF represents the journal impact factor, IFmax represents the highest journal impact factor in the field, and MediaScore represents the preset credibility score table.

[0026] Furthermore, after step S6, add the adjustable weight coefficients α, β, γ and the preset credibility score set by the user to the setting database; before step S1, according to the setting database, use the clustering algorithm and / or the random forest regression model to predict the optimal adjustable weight coefficients α, β, γ and the preset credibility score.

[0027] Further, the first large language model includes: a position encoding layer and a language processing unit; the output end of the position encoding layer is connected to the input end of the language processing unit;

[0028] The language processing unit includes a plurality of sequentially connected language processing modules.

[0029] Further, the language processing module includes: a multi-head self-attention layer, a first adder, a feed-forward neural network layer, an adapter, and a second adder;

[0030] The output end of the multi-head self-attention layer, the input end of the multi-head self-attention layer are connected to the input end of the first adder, the output end of the first adder is connected to the input end of the feed-forward neural network layer, the output end of the feed-forward neural network layer, the output end of the adapter are connected to the input end of the adapter, and the output end of the adapter, the output end of the first adder are connected to the input end of the second adder.

[0031] Further, the formula of the adapter is as follows:

[0032] Adapter(x) = W up ·ReLU(W down ·x) + x

[0033] W down represents a dimensionality reduction layer, W up represents a dimensionality increase layer, ReLU represents an activation function, and x represents an input.

[0034] A water ecological literature content analysis system based on multi-source data fusion and a large language model includes:

[0035] A preprocessing module: obtaining the literature to be analyzed and performing preprocessing to obtain the preprocessed text;

[0036] A confidence module: inputting the preprocessed text into the first large language model to obtain a domain label and a confidence score;

[0037] A first judgment module: if the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis;

[0038] A scoring module: calculating a literature quality score according to the citation volume, data integrity, and authoritative source of the literature to be analyzed;

[0039] A second judgment module: judging whether the literature quality score is greater than the preset literature quality score; if the literature quality score is greater than the preset literature quality score, then execute step S6; otherwise, terminate the analysis;

[0040] Name extraction module: Extract the species name, corresponding coordinate information, and GIS information according to the literature to be analyzed.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] A method for analyzing the content of water ecological literature based on multi-source data fusion and large language models of the present invention obtains the literature to be analyzed and performs preprocessing to obtain the preprocessed text; inputs the preprocessed text into the first large language model to obtain domain labels and confidence scores; if the domain label is a preset domain and the confidence score is greater than the preset confidence score; then calculate the literature quality score according to the citation volume, data integrity, and authoritative source of the literature to be analyzed; determine whether the literature quality score is greater than the preset literature quality score; if the literature quality score is greater than the preset literature quality score, extract the species name, corresponding coordinate information, and GIS information according to the literature to be analyzed. In summary, the literature content analysis method has the characteristics of accurate analysis and high analysis efficiency for ecology-related literature. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a structural diagram of a method for analyzing the content of water ecological literature based on multi-source data fusion and large language models provided in Embodiment 1.

[0044] Figure 2 It is a structural diagram of the first large language model provided in Embodiment 1.

[0045] Figure 3 It is a structural diagram of the language processing module provided in Embodiment 1.

[0046] Figure 4 It is a schematic diagram of the principle of the first large language model provided in Embodiment 1.

[0047] Figure 5 It is an application flowchart of a system for analyzing the content of water ecological literature based on multi-source data fusion and large language models provided in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0049] To better illustrate this embodiment, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of the actual product;

[0050] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0051] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0052] Example 1

[0053] As Figure 1 shown, a method for analyzing the content of water ecology literature based on multi-source data fusion and large language models includes:

[0054] S1: Obtain the literature to be analyzed and perform preprocessing to obtain the preprocessed text;

[0055] S2: Input the preprocessed text into the first large language model to obtain domain labels and confidence scores;

[0056] S3: If the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis;

[0057] S4: Calculate the literature quality score according to the citation volume, data integrity, and authoritative source of the literature to be analyzed;

[0058] S5: Judge whether the literature quality score is greater than the preset literature quality score; if the literature quality score is greater than the preset literature quality score, then execute step S6; otherwise, terminate the analysis;

[0059] S6: Extract the species name and the corresponding coordinate information and GIS information according to the literature to be analyzed.

[0060] In a specific embodiment, in step S3, if the confidence score is less than the preset confidence score, after correctly annotating the literature, perform incremental training on the parameters of the first large language model and update the first large language model.

[0061] It should be noted that catastrophic forgetting suppression is used in incremental training, that is, elastic weight consolidation (EWC) is adopted to constrain the change of key parameters. The specific loss function is as follows:

[0062]

[0063] Fi represents parameter importance, and λ represents the balance coefficient.

[0064] In a specific embodiment, in step S6, regular expressions and / or the second large language model are used to extract the species name and the corresponding coordinate information. If the parsing fails, new regular expressions are added to match the newly discovered coordinate format, or the fuzzy description is associated with the geofence. For example, "the upper reaches of the xx River" can be associated with the coordinate range [-2.5° to -2.2°, -54.8° to -54.5°].

[0065] In a specific embodiment, the preprocessing includes: using regular expressions to remove special symbols and spaces in the literature to be analyzed, and uniformly converting the coordinate information into decimal.

[0066] Furthermore, the literature quality scoring formula in step S4 is as follows:

[0067] Score = α·Citations + β·Completeness + γ·Authority

[0068] α, β, and γ represent adjustable weight coefficients, Citations represents the number of citations, Completeness represents data integrity, and Authority represents the authority of the source.

[0069] Furthermore, the calculation formula for the data integrity is as follows:

[0070]

[0071] S represents whether the species name exists, G represents whether the geographical coordinates exist, and T represents whether the time information exists.

[0072] Furthermore, the calculation formula for the number of citations is as follows:

[0073]

[0074] C represents the actual number of citations of the current literature, and Cmax represents the highest number of citations of the literature in the target field.

[0075] It should be noted that this formula uses a logarithmic function to compress the citation range to avoid excessive weight of highly cited literature. When the number of citations exceeds Cmax, the score upper limit is 1.

[0076] Furthermore, the formula for the authority of the source is as follows:

[0077]

[0078] IF represents the journal impact factor, IFmax represents the highest journal impact factor in the field, and MediaScore represents a preset credibility scoring table.

[0079] Furthermore, after step S6, the adjustable weight coefficients α, β, γ and the preset credibility score set by the user are added to the setting database; before step S1, according to the setting database, the optimal adjustable weight coefficients α, β, γ and the preset credibility score are predicted using a clustering algorithm and / or a random forest regression model.

[0080] Furthermore, as Figure 2 shown, the first large language model includes: a position encoding layer and a language processing unit; the output end of the position encoding layer is connected to the input end of the language processing unit;

[0081] The language processing unit includes a plurality of sequentially connected language processing modules.

[0082] In a specific embodiment, the number of language processing modules is preferably 6 to 9.

[0083] Furthermore, as Figure 3 shown, the language processing module includes: a multi-head self-attention layer, a first adder, a feed-forward neural network layer, an adapter, and a second adder;

[0084] The output end of the multi-head self-attention layer, the input end of the multi-head self-attention layer are connected to the input end of the first adder, the output end of the first adder is connected to the input end of the feed-forward neural network layer, the output end of the feed-forward neural network layer, the output end of the adapter are connected to the input end of the adapter, and the output end of the adapter, the output end of the first adder are connected to the input end of the second adder.

[0085] Figure 4 It is a schematic diagram of the principle of the above first large language model.

[0086] It should be noted that the residual connection enables the model to retain the original input features and prevent the phenomenon of gradient disappearance.

[0087] Furthermore, the formula of the adapter is as follows:

[0088] Adapter(x) = W up ·ReLU(W down ·x) + x

[0089] W down represents the dimensionality reduction layer, W up represents the dimensionality increase layer, ReLU represents the activation function, and x represents the input.

[0090] It should be noted that introducing a lightweight adapter module in the large language model enables it to quickly adapt to the characteristics of the ecological field (such as species distribution description, geographical coordinate parsing), avoids the high computational cost of full-parameter fine-tuning, and at the same time retains the general semantic understanding ability of the model.

[0091] A water ecological literature content analysis system based on multi-source data fusion and large language model includes:

[0092] A preprocessing module: obtaining the literature to be analyzed and performing preprocessing to obtain the preprocessed text;

[0093] A confidence module: inputting the preprocessed text into the first large language model to obtain a domain label and a confidence score;

[0094] A first judgment module: if the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis;

[0095] Scoring module: Calculate the literature quality score according to the citation volume, data integrity, and authoritative source of the literature to be analyzed;

[0096] Second judgment module: Judge whether the literature quality score is greater than the preset literature quality score; if the literature quality score is greater than the preset literature quality score, execute step S6; otherwise, terminate the analysis;

[0097] Name extraction module: Extract the species name and the corresponding coordinate information and GIS information according to the literature to be analyzed.

[0098] In a specific embodiment, if model quantization technology (such as 8-bit quantization of LLM) and parallel pipeline design (data chunking + multi-threaded processing) are adopted, a single node supports a daily throughput of millions of documents. Design a cache reuse mechanism to compare hash values for duplicate texts (such as news reprint content) to avoid duplicate calculations, reducing resource consumption by more than 40%. In practical applications, ecological research institutions can use the generated structured data to quickly construct species distribution heat maps, environmental protection departments can monitor ecological events in real time based on news data, and GIS suppliers can directly load standardized geographical files into the analysis platform. This system significantly reduces the data preprocessing cost, improves the scientific research and application efficiency, and has broad industrialization prospects.

[0099] Such as Figure 5 As shown, users can obtain the multi-dimensional map of the data in real time on the platform according to their own needs to measure the interpretation accuracy of the literature.

[0100] The same or similar reference numerals correspond to the same or similar components;

[0101] The terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as a limitation of this patent;

[0102] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for analyzing water ecology literature content based on multi-source data fusion and large language model, characterized in that: include: S1: Obtain the document to be analyzed and preprocess it to obtain the preprocessed text; S2: Inputting the preprocessed text into the first language model to obtain a domain label and a confidence score; S3: If the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis; S4: Calculate the quality score of the literature according to the number of citations, data completeness and authoritative source of the literature to be analyzed; S5: Determine whether the document quality score is greater than a preset document quality score; if the document quality score is greater than the preset document quality score, execute step S6; otherwise, terminate the analysis; S6: Extract species names and corresponding coordinate information and GIS information according to the document to be analyzed.

2. According to claim 1, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: The document quality scoring formula in step S4 is as follows: Score=α·Citations+β·Completeness+γ·Authority α, β, and γ represent adjustable weight coefficients, Citations represents the number of citations, Completeness represents data integrity, and Authority represents authoritative source.

3. According to claim 2, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: The calculation formula for the data integrity is as follows: S indicates whether the species name exists, G indicates whether the geographic coordinates exist, and T indicates whether the time information exists.

4. According to claim 2, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: The calculation formula of the cited amount is as follows: C represents the actual number of citations of the current document, and Cmax represents the maximum number of citations of documents in the target field.

5. According to claim 2, a method for analyzing water ecology literature content based on multi-source data fusion and large language model, characterized in that: The formula for the authoritative source is as follows: IF represents the journal impact factor, IFmax represents the highest impact factor of the journal in the field, and MediaScore represents the preset credibility rating scale.

6. According to claim 2, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: After step S6, the adjustable weight coefficients α, β, γ and preset confidence scores currently set by the user are added to the setting database; before step S1, based on the setting database, the optimal adjustable weight coefficients α, β, γ and preset confidence scores are predicted using a clustering algorithm and / or a random forest regression model.

7. According to claim 1, a method for analyzing water ecology literature content based on multi-source data fusion and large language model, characterized in that: The first language model includes: a position encoding layer and a language processing unit; the output end of the position encoding layer is connected to the input end of the language processing unit; The language processing unit includes a plurality of language processing modules connected in sequence.

8. According to claim 7, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: The language processing module includes: a multi-head self-attention layer, a first addition point, a feedforward neural network layer, an adapter, and a second addition point; The output end of the multi-head self-attention layer and the input end of the multi-head self-attention layer are connected to the input end of the first addition point, the output end of the first addition point is connected to the input end of the feedforward neural network layer, the output end of the feedforward neural network layer and the output end of the adapter are connected to the input end of the adapter, and the output end of the adapter and the output end of the first addition point are connected to the input end of the second addition point.

9. According to claim 8, a method for analyzing water ecology literature content based on multi-source data fusion and large language model is characterized in that: The formula for the adapter is as follows: Adaptor(x)=W up ·ReLU(W down x)+x W down represents the dimension reduction layer, W up ReLU represents the activation function, and x represents the input.

10. A water ecology literature content analysis system based on multi-source data fusion and large language model, applied to the analysis method according to any one of claims 1 to 9, characterized in that: include: Preprocessing module: obtain the document to be analyzed, and preprocess it to obtain the preprocessed text; Confidence module: input the preprocessed text into the first language model to obtain a domain label and a confidence score; First judgment module: if the domain label is a preset domain and the confidence score is greater than the preset confidence score, then execute step S4; otherwise, terminate the analysis; Scoring module: Calculate the quality score of the document according to the number of citations, data integrity and authoritative source of the document to be analyzed; The second judgment module: judges whether the document quality score is greater than a preset document quality score; If the document quality score is greater than the preset document quality score, step S6 is executed; otherwise, the analysis is terminated; Name extraction module: extracts species names and corresponding coordinate information and GIS information according to the document to be analyzed.

Citation Information

Patent Citations

  • Method for evaluating influence of references in scientific documents

    CN107391921A

  • Meta analysis generation method based on artificial intelligence

    CN111552776A

  • Scientific and technological achievement value evaluation method and system based on technical chain

    CN117591628A

  • Scientific field-oriented multi-modal corpus data construction method and device

    CN118170933A

  • Scientific and technical literature review automatic generation method and device

    CN118278365A