Scientific research information retrieval matching method and device based on large model

By employing a large-model-based approach for scientific research information retrieval, utilizing pre-trained language models and knowledge graphs, and combining various search strategies, this approach addresses the issues of insufficient semantic understanding, data noise interference, and coarse relationship modeling in scientific research information retrieval, achieving efficient and accurate matching of scientific research information.

CN120929557AActive Publication Date: 2025-11-11INST OF LOGISTICS SCI & TECH ACAD OF SYST ENG ACAD OF MILITARY SCI
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202511132209.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-11
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing scientific research information retrieval methods suffer from insufficient semantic understanding, data noise interference, low efficiency of multimodal retrieval fusion, and coarse relationship modeling, resulting in low retrieval accuracy and efficiency.

Method used

We employ a large-model-based approach, using a pre-trained language model for text vectorization and data cleaning. This is combined with knowledge graph construction and various search strategies, including knowledge graph semantic search, embedding vector search, and RAG search. By utilizing the relevance calculation between grouped word sets and the thesaurus, we dynamically filter entities and calculate relation values ​​and edge information to construct a high-quality retrieval graph and fuse results.

Benefits of technology

It enhances the semantic understanding capabilities of scientific research information retrieval, reduces data noise interference, improves the fusion efficiency of multimodal retrieval and the accuracy of relationship modeling, enhances the matching accuracy and recall rate in complex scientific research scenarios, and adapts to the semantic distribution characteristics of different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929557A_ABST
    Figure CN120929557A_ABST
Patent Text Reader

Abstract

The invention discloses a large model-based scientific research information retrieval matching method and device. The method comprises the following steps of: acquiring to-be-retrieved scientific research text information; the scientific research text information to be retrieved comprises a plurality of retrieval words; processing the scientific research text information to be retrieved to obtain a retrieval atlas; and carrying out retrieval processing on the retrieval atlas by utilizing a preset scientific research knowledge atlas to obtain retrieval matching result information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent text data processing and equipment system evaluation, specifically to a scientific research information retrieval and matching method and device based on a large model. Background Technology

[0002] In the current field of scientific research information retrieval, traditional retrieval methods face the following technical problems:

[0003] Insufficient semantic understanding: Existing technologies mostly rely on keyword matching or simple vector representations, which struggle to capture the complex semantic relationships in scientific research texts (such as term synonyms and domain context dependencies), leading to semantic discrepancies between search results and user needs. For example, while BERT-based vectorization can handle some semantics, it lacks modeling of the structured relationships between search terms.

[0004] Data noise interference: Scientific research texts often contain irrelevant terms, formatting errors, or words from outside the field (such as general terms mixed into professional literature). Traditional preprocessing methods (such as simple word frequency filtering) are difficult to accurately remove noise, which leads to the introduction of redundant information when constructing the retrieval map and affects the matching accuracy.

[0005] Multimodal retrieval fusion is inefficient: Existing retrieval systems often employ a single search strategy (such as relying solely on semantic search or vector matching), failing to balance retrieval speed and accuracy. For example, while embedded vector search is efficient, it lacks semantic reasoning capabilities, and although the RAG method can incorporate knowledge bases, its response speed is slow.

[0006] Coarse relationship modeling: In traditional knowledge graph construction, the relationships between entities are mostly based on predefined rules or simple similarity calculations, which makes it difficult to dynamically capture the deep connections hidden in scientific research texts (such as the relevance decay of cross-domain terms and the contextual distinction of polysemous words), resulting in insufficient structural expressiveness of the retrieval graph.

[0007] The aforementioned technical issues result in low accuracy and efficiency for existing scientific research information retrieval methods. Summary of the Invention

[0008] This invention primarily addresses the problems of insufficient semantic understanding, data noise interference, low efficiency of multimodal retrieval fusion, and coarse relationship modeling in existing scientific research information retrieval methods. This invention discloses a scientific research information retrieval matching method and device based on a large model.

[0009] In a first aspect, this invention discloses a scientific research information retrieval and matching method based on a large model, comprising:

[0010] S1, Collect the research text information to be retrieved; the research text information to be retrieved includes several search terms;

[0011] S2, process the scientific research text information to be retrieved to obtain a retrieval map;

[0012] S3. Using a preset scientific knowledge graph, perform retrieval processing on the retrieval graph to obtain retrieval matching results.

[0013] The process of processing the scientific research text information to be retrieved to obtain a retrieval map includes:

[0014] S21, preprocess the scientific research text information to be retrieved to obtain preprocessed embedded vector information;

[0015] S22, construct a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph.

[0016] The preprocessing of the scientific research text information to be retrieved to obtain preprocessed embedded vector information includes:

[0017] S211, The scientific research text information to be retrieved is represented by text vectorization to obtain a text embedding vector;

[0018] S212, perform data cleaning processing on the text embedding vector to obtain the first text information;

[0019] S213, perform category checking on the first text information to obtain preprocessed embedded vector information.

[0020] The step of constructing a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph includes:

[0021] S221, using preset search term grouping rules, all search terms in the preprocessed embedded vector information are grouped to obtain several group term sets; the group term sets include several search terms;

[0022] S222, For each group of words, calculate the relevance with the corresponding vocabulary to obtain the relevance calculation results;

[0023] S223, the set of grouped words whose relevant calculation results are greater than the preset relevant threshold value are used as entity information of the retrieval map;

[0024] S224, calculates the relationship values ​​between the information of each entity;

[0025] S225, using the relationship values ​​between the various entity information, a relationship matrix is ​​constructed; the elements in the i-th row and j-th column of the relationship matrix are the relationship values ​​between the i-th search term in the grouping word set corresponding to the first entity information and the j-th search term in the grouping word set corresponding to the second entity information.

[0026] S226, perform edge information calculation on the relation matrix to obtain the edge information between entity information in the retrieval graph;

[0027] S227, using all entity information and edge information, a retrieval graph is constructed;

[0028] The expression for calculating the edge information is:

[0029]

[0030] Among them, W N1,M1 (x) represents the Whittaker function, x is the input value of the Whittaker function, N1 and M1 are the row and column dimensions of the relation matrix, respectively, and G... ij Let σ be the element in the i-th row and j-th column of the relation matrix. ij Let σ be the j-th element of the i-th singular vector obtained by performing singular value decomposition on the relation matrix. i0 Let τ be the mean of the i-th singular vector. i Let τi be the i-th singular value obtained by performing singular value decomposition on the relation matrix, τ0 be the mean of all singular values, and b be the edge information.

[0031] The process of using a pre-defined scientific knowledge graph to perform retrieval processing on the retrieval graph to obtain retrieval matching results includes:

[0032] S31, Using the knowledge graph semantic search method, the retrieval graph is searched in the preset scientific research knowledge graph to obtain the first search result information;

[0033] S32, Using the embedded vector search method, the retrieval map is searched in the preset scientific knowledge map to obtain the second search result information;

[0034] S33, Using the RAG search method, the retrieval map is searched in the preset scientific knowledge map to obtain the third search result information;

[0035] S34. Perform fusion matching calculation on all the obtained search results information to obtain the search matching result information.

[0036] The process of fusing and matching all obtained search results to obtain search matching result information includes:

[0037] S341, perform text alignment processing on all search result information to obtain aligned search result information;

[0038] S342, calculate the importance coefficient for each aligned search result information to obtain the importance coefficient value;

[0039] S343 uses the importance coefficient value to fuse and calculate all aligned search result information to obtain the search matching result information.

[0040] The expression for calculating the importance coefficient is as follows:

[0041]

[0042] Where zy is the importance coefficient value, β i For the i-th entity information in the retrieval map, λ i M2 is the average value of all edge information of the i-th entity information in the retrieval graph, M2 is the total number of entity information in the retrieval graph, |α| represents the modulus of the aligned search result information α, and cos(α,β) i () indicates the result of the cosine similarity calculation;

[0043] The expression for the fusion calculation is:

[0044]

[0045] Where, ε j To retrieve the j-th element of the matching results, α ij This refers to the j-th element of the aligned i-th search result information.

[0046] A second aspect of this invention discloses a scientific research information retrieval and matching device based on a large model, the device comprising:

[0047] Memory containing executable program code;

[0048] A processor coupled to the memory;

[0049] The processor calls the executable program code stored in the memory to execute the scientific research information retrieval and matching method based on the large model.

[0050] In a third aspect of this invention, a computer-storable medium is disclosed, wherein the computer-storable medium stores computer instructions, which, when invoked by a computer, are used to execute the aforementioned scientific research information retrieval and matching method based on a large model.

[0051] In a fourth aspect of this invention, an information data processing terminal is disclosed, which is used to implement the aforementioned scientific research information retrieval and matching method based on a large model.

[0052] The beneficial effects of this invention are as follows:

[0053] This invention uses a pre-trained language model (such as BERT) to vectorize text, and combines data cleaning (smoothing noise and removing outliers) and category checking (filtering irrelevant words based on a scientific text library) to improve the purity and semantic representation ability of the input data, laying a high-quality foundation for subsequent graph construction.

[0054] This invention calculates the relevance between grouped word sets and the lexicon (integrating Jaccard similarity and Jaro-Winkler distance), dynamically filters highly relevant entities, and accurately characterizes the semantic relationships between entities through relation matrix and edge information calculation (introducing Whittaker function to handle singular value decomposition results), making the retrieval graph more closely reflect the knowledge structure of scientific research fields.

[0055] This invention combines knowledge graph semantic search, embedded vector search (supporting brute force / tree structure indexing) and RAG search, taking into account semantic understanding, retrieval efficiency and knowledge base reasoning ability. By weighting and fusing multi-source results with importance coefficients, it improves matching accuracy and recall in complex scientific research scenarios.

[0056] This invention introduces a decision factor and a constant factor into the relation value calculation, and dynamically adjusts the relation weights between entities based on the maximum relevance value, adapting to the semantic distribution characteristics of scientific research texts in different fields and enhancing the model's generalization ability. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention. Detailed Implementation

[0058] To better understand the content of this invention, an embodiment is provided here.

[0059] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention.

[0060] In a first aspect, this invention discloses a scientific research information retrieval and matching method based on a large model, comprising:

[0061] S1, Collect the research text information to be retrieved; the research text information to be retrieved includes several search terms;

[0062] S2, process the scientific research text information to be retrieved to obtain a retrieval map;

[0063] S3, using a preset scientific knowledge graph, perform retrieval processing on the retrieval graph to obtain retrieval matching result information;

[0064] The process of processing the scientific research text information to be retrieved to obtain a retrieval map includes:

[0065] S21, preprocess the scientific research text information to be retrieved to obtain preprocessed embedded vector information;

[0066] S22, construct a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph;

[0067] The preprocessing of the scientific research text information to be retrieved to obtain preprocessed embedded vector information includes:

[0068] S211, The scientific research text information to be retrieved is represented by text vectorization to obtain a text embedding vector;

[0069] S212, perform data cleaning processing on the text embedding vector to obtain the first text information;

[0070] S213, perform category checking on the first text information to obtain preprocessed embedded vector information.

[0071] The vectorized representation of the text can be implemented using pre-trained language models (such as BERT, RoBERTa, ALBERT, etc.).

[0072] The data cleaning process includes smoothing noisy data and smoothing or deleting outliers;

[0073] The category checking process involves using a pre-defined scientific research text database to remove texts from the first text information that are not in the database, thereby obtaining pre-processed embedded vector information.

[0074] The step of constructing a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph includes:

[0075] S221, using preset search term grouping rules, all search terms in the preprocessed embedded vector information are grouped to obtain several group term sets; the group term sets include several search terms;

[0076] S222, For each group of words, calculate the relevance with the corresponding vocabulary to obtain the relevance calculation results;

[0077] S223, the set of grouped words whose related calculation results are greater than the preset related threshold value is taken as entity information of the retrieval map; a set of grouped words whose related calculation results are greater than the preset related threshold value is taken as entity information.

[0078] S224, calculates the relationship values ​​between the information of each entity;

[0079] S225, using the relationship values ​​between the various entity information, a relationship matrix is ​​constructed; the elements in the i-th row and j-th column of the relationship matrix are the relationship values ​​between the i-th search term in the grouping word set corresponding to the first entity information and the j-th search term in the grouping word set corresponding to the second entity information.

[0080] The first and second entity information are the two entity information for calculating the relation value;

[0081] S226, perform edge information calculation on the relation matrix to obtain the edge information between entity information in the retrieval graph;

[0082] S227. Using all entity and edge information, a retrieval graph is constructed.

[0083] The expression for calculating the relevance between the grouped word set and the corresponding vocabulary is:

[0084]

[0085] Where xg is the relevance, and co1(f i ,ck j ) indicates the first relevance calculation, co2(f i ,ck j ) indicates the calculation of the second relevance, f i ck represents the i-th detection word in the grouped word set. j The term represents the j-th word in the thesaurus corresponding to the grouped term set, and N and M represent the total number of search terms contained in the grouped term set and the total number of words contained in the thesaurus corresponding to the grouped term set, respectively.

[0086] The first relevance can be calculated using Jaccard similarity, and the second relevance can be calculated using Jaro-Winkler distance.

[0087] The relevance calculation between the grouped word sets and their corresponding thesaurus combines Jaccard similarity (co1, measuring set overlap) and Jaro-Winkler distance (co2, measuring string edit distance) to simultaneously capture semantic inclusion relationships (e.g., "machine learning" and "deep learning") and spelling similarity (e.g., "algorithm" and "algorithms"), avoiding the limitations of a single indicator. A non-linear transformation of the relevance values ​​is performed using the tan function, and normalization is achieved using the maximum relevance within each group. This highlights the contribution of highly relevant terms and suppresses the interference of low-relevance terms, making entity filtering of the grouped word sets more accurate.

[0088] The expression for calculating the relationship value between the various entity information is as follows:

[0089]

[0090] Where G is the relational value, Co3 ij This represents the relevance value between the i-th search term of the first entity information and the j-th search term of the second entity information, used to calculate the relation value. max{Co3 ij} represents the maximum relevance value of all detected words between the first entity information and the second entity information used to calculate the relationship value, where δ is a preset constant factor and ψ is a preset decision factor. The relevance value can be expressed using Jaccard similarity.

[0091] When the maximum correlation between entities is max{Co3 ij When the threshold exceeds the decision factor ψ, a linear decay model is adopted. To mitigate the impact of extremely high correlation values ​​and prevent a single strong association from obscuring the overall semantic relationship, a global weighted average is used to ensure that the relationship between weakly correlated entities is still effectively modeled. A constant factor δ is used to adjust the decay rate of relation values ​​to adapt to the distribution characteristics of term relevance in different domains (e.g., term relevance is more concentrated in the biomedical field and more dispersed in the information technology field), thereby improving the model's adaptability to cross-domain data.

[0092] The expression for calculating the edge information is:

[0093]

[0094] Among them, W N1,M1 (x) represents the Whittaker function, x is the input value of the Whittaker function, N1 and M1 are the row and column dimensions of the relation matrix, respectively, and G... ij Let σ be the element in the i-th row and j-th column of the relation matrix. ij Let σ be the j-th element of the i-th singular vector obtained by performing singular value decomposition on the relation matrix. i0 Let τ be the mean of the i-th singular vector. i Let τi be the i-th singular value obtained by performing singular value decomposition on the relation matrix, τ0 be the mean of all singular values, and b be the edge information.

[0095] The edge information calculation combines the singular values ​​(reflecting the overall matrix energy) and singular vectors (reflecting the matrix characteristic distribution) of the relation matrix, and calculates the element-level distance σ. i -G ij and σ i0 -σ ij This method captures the global structure and local details of entity relationships, enabling edge information to reflect both relationship strength and graph topological features. By using the Whittaker function to perform a nonlinear mapping on weighted distances, numerical features are transformed into edge weights suitable for graph construction, enhancing the expressive power of complex semantic relationships (such as indirect associations between cross-domain terms and the transitivity of multi-hop relationships).

[0096] The preset search term grouping rules include several thesauruses; each thesaurus includes several entity words; each group term set has a corresponding thesaurus.

[0097] The process of using a pre-defined scientific knowledge graph to perform retrieval processing on the retrieval graph to obtain retrieval matching results includes:

[0098] S31, Using the knowledge graph semantic search method, the retrieval graph is searched in the preset scientific research knowledge graph to obtain the first search result information;

[0099] S32, Using the embedded vector search method, the retrieval map is searched in the preset scientific knowledge map to obtain the second search result information;

[0100] S33, Using the RAG search method, the retrieval map is searched in the preset scientific knowledge map to obtain the third search result information;

[0101] S34. Perform fusion matching calculation on all the obtained search results information to obtain the search matching result information.

[0102] The embedded vector search method can employ either brute-force search or tree-structured index search.

[0103] The process of fusing and matching all obtained search results to obtain search matching result information includes:

[0104] S341, perform text alignment processing on all search result information to obtain aligned search result information;

[0105] S342, calculate the importance coefficient for each aligned search result information to obtain the importance coefficient value;

[0106] S343 uses the importance coefficient value to fuse and calculate all aligned search result information to obtain the search matching result information.

[0107] The expression for calculating the importance coefficient is as follows:

[0108]

[0109] Where zy is the importance coefficient value, β i For the i-th entity information in the retrieval map, λ i M2 is the average value of all edge information of the i-th entity information in the retrieval graph, M2 is the total number of entity information in the retrieval graph, |α| represents the modulus of the aligned search result information α, and cos(α,β) i() indicates the result of the cosine similarity calculation;

[0110] The expression for the fusion calculation is:

[0111]

[0112] Where, ε j To retrieve the j-th element of the matching results, α ij This refers to the j-th element of the aligned i-th search result information.

[0113] The RAG search method combines retrieval and generation, providing accurate answers while enhancing contextual understanding. After a user submits a query, semantic vectorization technology is first used to transform the question and text in the knowledge base into high-dimensional vector representations. Based on similarity calculations, the most relevant text paragraphs are retrieved from the knowledge base. The retrieval results are then used as contextual input to the generative model, which uses generative models (such as T5, BART, or GPT) to generate natural language answers. This RAG architecture ensures seamless integration between retrieval results and the generation process, making the answers more accurate and context-relevant.

[0114] The knowledge graph semantic search method can be a representation learning-based knowledge graph semantic search method or a faceted knowledge graph semantic search method.

[0115] The textual vectorization of the research text information to be retrieved yields embedded text vectors. These vectors can be converted into high-dimensional vectors by semantic representation of the text data using pre-trained language models (such as BERT, RoBERTa, ALBERT, etc.). These vectors capture the semantic features of the text, facilitating subsequent retrieval and computation. In specific scenarios, the pre-trained model can be fine-tuned to further improve the semantic accuracy of the embeddings, ensuring a deeper understanding of domain-specific vocabulary or context. Furthermore, single-tower or dual-tower models can be selected based on specific task requirements to optimize the distribution of semantic vectors in the high-dimensional space, thus providing a more accurate foundation for retrieval. High-performance vector retrieval tools are used to index and store the vectors generated by the embeddings. These tools employ algorithms such as HNSW and IVF to quickly find the set of vectors most similar to the query vector. The retrieval process can also be combined with techniques such as inverted indexes and approximate nearest neighbor search to improve efficiency. To improve the accuracy of the retrieval results, a re-ranking model can be used to semantically re-rank the initial retrieval results, ensuring that the returned information is highly relevant to the user's query.

[0116] Retrieved relevant vector information is used as conditional input to generative models (such as GPT, T5, and BLOOM), and contextual information is injected into the generation process through prompting learning or the insertion of explicit knowledge. This not only improves the accuracy of the generated content but also enhances the model's performance in specific domains.

[0117] A second aspect of this invention discloses a scientific research information retrieval and matching device based on a large model, the device comprising:

[0118] Memory containing executable program code;

[0119] A processor coupled to the memory;

[0120] The processor calls the executable program code stored in the memory to execute the scientific research information retrieval and matching method based on the large model.

[0121] In a third aspect of this invention, a computer-storable medium is disclosed, wherein the computer-storable medium stores computer instructions, which, when invoked by a computer, are used to execute the aforementioned scientific research information retrieval and matching method based on a large model.

[0122] In a fourth aspect of this invention, an information data processing terminal is disclosed, which is used to implement the aforementioned scientific research information retrieval and matching method based on a large model.

[0123] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A scientific research information retrieval and matching method based on a large model, characterized in that, include: S1, Collect the research text information to be retrieved; the research text information to be retrieved includes several search terms; S2, process the scientific research text information to be retrieved to obtain a retrieval map; S3. Using a preset scientific knowledge graph, perform retrieval processing on the retrieval graph to obtain retrieval matching results.

2. The scientific research information retrieval and matching method based on a large model as described in claim 1, characterized in that, The process of processing the scientific research text information to be retrieved to obtain a retrieval map includes: S21, preprocess the scientific research text information to be retrieved to obtain preprocessed embedded vector information; S22, construct a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph.

3. The scientific research information retrieval and matching method based on a large model as described in claim 2, characterized in that, The preprocessing of the scientific research text information to be retrieved to obtain preprocessed embedded vector information includes: S211, The scientific research text information to be retrieved is represented by text vectorization to obtain a text embedding vector; S212, perform data cleaning processing on the text embedding vector to obtain the first text information; S213, perform category checking on the first text information to obtain preprocessed embedded vector information.

4. The scientific research information retrieval and matching method based on a large model as described in claim 2, characterized in that, The step of constructing a knowledge graph from the preprocessed embedded vector information to obtain a retrieval graph includes: S221, using preset search term grouping rules, all search terms in the preprocessed embedded vector information are grouped to obtain several group term sets; the group term sets include several search terms; S222, For each group of words, calculate the relevance with the corresponding vocabulary to obtain the relevance calculation results; S223, the set of grouped words whose relevant calculation results are greater than the preset relevant threshold value are used as entity information of the retrieval map; S224, calculates the relationship values ​​between the information of each entity; S225, using the relationship values ​​between the various entity information, a relationship matrix is ​​constructed; the elements in the i-th row and j-th column of the relationship matrix are the relationship values ​​between the i-th search term in the grouping word set corresponding to the first entity information and the j-th search term in the grouping word set corresponding to the second entity information. S226, perform edge information calculation on the relation matrix to obtain the edge information between entity information in the retrieval graph; S227, using all entity information and edge information, a retrieval graph is constructed; The expression for calculating the edge information is: Among them, W N1,M1 (x) represents the Whittaker function, x is the input value of the Whittaker function, N1 and M1 are the row and column dimensions of the relation matrix, respectively, and G... ij Let σ be the element in the i-th row and j-th column of the relation matrix. ij Let σ be the j-th element of the i-th singular vector obtained by performing singular value decomposition on the relation matrix. i0 Let τ be the mean of the i-th singular vector. i Let τi be the i-th singular value obtained by performing singular value decomposition on the relation matrix, τ0 be the mean of all singular values, and b be the edge information.

5. The scientific research information retrieval and matching method based on a large model as described in claim 1, characterized in that, The process of using a pre-defined scientific knowledge graph to perform retrieval processing on the retrieval graph to obtain retrieval matching results includes: S31, Using the knowledge graph semantic search method, the retrieval graph is searched in the preset scientific research knowledge graph to obtain the first search result information; S32, Using the embedded vector search method, the retrieval map is searched in the preset scientific knowledge map to obtain the second search result information; S33, Using the RAG search method, the retrieval map is searched in the preset scientific research knowledge map to obtain the third search result information; S34. Perform fusion and matching calculations on all the obtained search results to obtain the search matching results.

6. The scientific research information retrieval and matching method based on a large model as described in claim 5, characterized in that, The process of fusing and matching all obtained search results to obtain search matching result information includes: S341, perform text alignment processing on all search result information to obtain aligned search result information; S342, calculate the importance coefficient for each aligned search result information to obtain the importance coefficient value; S343 uses the importance coefficient value to fuse and calculate all aligned search result information to obtain the search matching result information.

7. The scientific research information retrieval and matching method based on a large model as described in claim 6, characterized in that, The expression for calculating the importance coefficient is as follows: Where zy is the importance coefficient value, β i For the i-th entity information in the retrieval map, λ i M2 is the average value of all edge information of the i-th entity information in the retrieval graph, M2 is the total number of entity information in the retrieval graph, |α| represents the modulus of the aligned search result information α, and cos(α,β) i () indicates the result of the cosine similarity calculation; The expression for the fusion calculation is: Where, ε j To retrieve the j-th element of the matching results, α ij This refers to the j-th element of the aligned i-th search result information.

8. A scientific research information retrieval and matching device based on a large model, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the scientific research information retrieval and matching method based on a large model as described in any one of claims 1 to 7.

9. A computer-storable medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked by the computer, are used to execute the scientific research information retrieval and matching method based on a large model as described in any one of claims 1 to 7.

10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the scientific research information retrieval and matching method based on a large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Question and answer method and device based on knowledge graph, computer equipment and storage medium

    CN110457431A

  • Operation and maintenance fault diagnosis and analysis method based on subgraph matching and distributed query

    CN114186073A

  • Intelligent retrieval method and system based on exploration and development knowledge graph and storage medium

    CN116401350A

  • Knowledge graph question-answering method for sub-graph retrieval optimization

    CN117149974A

  • College scientific research management question and answer system combining knowledge graph and large language model

    CN117609436A