Literature search method based on large language model

By constructing a domain knowledge base using a large language model, the contradiction between comprehensiveness and accuracy in literature retrieval methods is resolved, enabling large-scale and highly accurate literature retrieval and generating detailed literature analysis reports.

CN120561269BActive Publication Date: 2026-01-02WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510645351.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2026-01-02
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing literature retrieval methods present a contradiction between ensuring comprehensiveness and accuracy, making it impossible to simultaneously achieve a wide range of literature retrievals with high accuracy.

Method used

A domain knowledge base for the target field is constructed using a large language model, including core definitions, a multi-dimensional thesaurus, a topic search query, an ambiguous thesaurus, and a constraint rule base. Literature analysis reports are generated through topic search and literature filtering.

Benefits of technology

It achieves a wide range of highly accurate literature retrieval, comprehensively covering relevant literature in the target field and accurately retaining literature belonging to the target field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561269B_ABST
    Figure CN120561269B_ABST
Patent Text Reader

Abstract

The application discloses a literature retrieval method based on a large language model, and belongs to the technical field of the large language model, and the method comprises the following steps: a large language model acquires natural language text from a user; according to the natural language text, a domain knowledge base of a target domain is constructed, and the domain knowledge base comprises core definitions corresponding to the target domain, a multi-dimensional word library, a subject retrieval formula, an ambiguous word library and a constraint rule library; according to the subject retrieval formula in the domain knowledge base, literature retrieval is carried out, and a first literature set of the target domain is obtained; based on the core definitions, the ambiguous word library and the constraint rules in the domain knowledge base, the literature in the first literature set is screened, and a second literature set is obtained; and a literature analysis report is generated based on the second literature set. The method can realize large-range and high-accuracy literature retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of large language model, in particular to a literature retrieval method based on large language model. BACKGROUND

[0002] When performing literature retrieval, a Boolean logic retrieval formula is usually used for literature retrieval.

[0003] However, there is an inherent contradiction between comprehensiveness and accuracy in the retrieval formula. If the comprehensiveness of the retrieval result is to be ensured, the retrieval formula needs to be expanded, which will inevitably introduce noise literature, resulting in a decrease in the accuracy of the literature retrieval result; if high accuracy is required, the range of search terms needs to be strictly limited and the number of keywords needs to be reduced, resulting in a decrease in the number of retrieved literature and affecting the range of literature coverage. SUMMARY

[0004] The present disclosure provides a literature retrieval method based on a large language model, which can achieve large-scale and high-accuracy literature retrieval. The technical solution at least includes the following solutions:

[0005] In a first aspect, a literature retrieval method based on a large language model is provided, including: a large language model obtains a natural language text from a user, the natural language text being used for retrieving literature in a target field; according to the natural language text, a domain knowledge base of the target field is constructed, the domain knowledge base including core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguous word library and a constraint rule library corresponding to the target field; literature retrieval is performed based on the topic retrieval formula in the domain knowledge base to obtain a first literature set of the target field; the literature in the first literature set is screened based on the core definitions, the ambiguous word library and the constraint rule in the domain knowledge base to obtain a second literature set, the screening of the literature in the first literature set being used to eliminate literature in the first literature set that does not belong to the target field and to retain literature in the first literature set that belongs to the target field; and a literature analysis report is generated based on the second literature set.

[0006] Optionally, the constructing the domain knowledge base of the target domain according to the natural language text comprises: extracting a subject phrase of the natural language text; splitting the subject phrase according to a time dimension, a space dimension, a person dimension and a subject dimension; generating a core definition of the domain knowledge base according to the subject phrase; performing synonym expansion on words in different dimensions of the subject phrase to generate a multi-dimensional word library of the subject phrase, the multi-dimensional word library comprising a time word library, a space word library, a person word library and a subject word library; performing semantic analysis on words in the multi-dimensional word library to obtain an ambiguous word library in the multi-dimensional word library, the ambiguous word library comprising a plurality of ambiguous words, the ambiguous word being a word having at least one meaning inconsistent with the core definition; generating a subject search formula according to the multi-dimensional word library; and receiving constraint rules defined by a user to obtain a constraint rule set.

[0007] Optionally, the performing semantic analysis on the words in the multi-dimensional word library to obtain the ambiguous word library in the multi-dimensional word library comprises: obtaining a first semantic vector corresponding to each definition of a first ambiguous word, the first ambiguous word being any ambiguous word in the multi-dimensional word library; obtaining a second semantic vector corresponding to the core definition; comparing the similarity between each first semantic vector and the second semantic vector one by one; and in a case where there is at least one first semantic vector having a similarity smaller than a similarity threshold value with the second semantic vector, storing the first ambiguous word into the ambiguous word library.

[0008] Optionally, the core definition comprises an academic definition, a subject category, an application scenario, a key technology and a key method, and supplementary notes, wherein the academic definition is generated by using a genus-plus-differentia definition method, the subject category comprises a primary subject and a secondary subject, the application scenario comprises a core application scenario and a frontier application scenario, and the supplementary notes comprise controversial points in the core technology, key scholars and key institutions.

[0009] Optionally, the screening the documents in the first document set one by one based on the core definition, the ambiguous word library and the constraint rules in the domain knowledge base to obtain a second document set comprises: performing semantic analysis on a document title of a first document to determine whether the first document belongs to the target domain based on the core definition, the ambiguous word library and the constraint rules in the domain knowledge base, and in a case where the semantic analysis on the document title of the first document cannot determine whether the first document belongs to the target domain, performing semantic analysis on an abstract of the first document to determine whether the first document belongs to the target domain, the first document being any document in the first document set; and in a case where the first document belongs to the target domain, storing the first document into the second document set.

[0010] Optionally, the filtering of the literatures in the first literature set based on the core definition, the ambiguity word library and the constraint rules in the domain knowledge base obtains a second literature set, comprising: based on the ambiguity word library in the domain knowledge base, eliminating literatures corresponding to the ambiguous definition of a first ambiguity word from the first literature set; wherein the first ambiguity word is any word in the ambiguity word library, and the similarity between the semantic vector corresponding to the ambiguous definition and the second semantic vector is less than the similarity threshold.

[0011] The second aspect also provides a literature retrieval device based on a large language model, comprising: an acquisition module configured to acquire, by a large language model, a natural language text from a user, the natural language text being used for retrieving literatures in a target domain; a domain knowledge base construction module configured to construct a domain knowledge base of the target domain according to the natural language text, the domain knowledge base comprising core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguity word library and a constraint rule library corresponding to the target domain; a retrieval module configured to perform literature retrieval based on the topic retrieval formula in the domain knowledge base, and obtain a first literature set of the target domain; a filtering module configured to filter literatures in the first literature set based on the core definitions, the ambiguity word library and the constraint rules in the domain knowledge base, and obtain a second literature set, the filtering of the literatures in the first literature set being used for eliminating literatures in the first literature set that do not belong to the target domain, and retaining literatures in the first literature set that belong to the target domain; and a generation module configured to generate a literature analysis report based on the second literature set.

[0012] Optionally, the domain knowledge base construction module is further configured to extract a topic phrase of the natural language text; split the topic phrase according to a time dimension, a space dimension, a person dimension and a topic dimension; generate the core definitions of the domain knowledge base according to the topic phrase; perform synonym expansion on words in different dimensions in the topic phrase, and generate a multi-dimensional word library of the topic phrase, the multi-dimensional word library comprising a time word library, a space word library, a person word library and a topic word library; perform semantic analysis on words in the multi-dimensional word library, and acquire an ambiguity word library in the multi-dimensional word library, the ambiguity word library comprising a plurality of ambiguity words, the ambiguity word being a word having at least one meaning inconsistent with the core definitions; generate a topic retrieval formula according to the multi-dimensional word library; and receive constraint rules defined by the user, and obtain a constraint rule set.

[0013] Optionally, the domain knowledge base construction module is further configured to obtain a first semantic vector corresponding to each definition of a first polysemous word, the first polysemous word being any polysemous word in the multi-dimensional word library; obtain a second semantic vector corresponding to the core definition; compare the similarity between each first semantic vector and the second semantic vector one by one; and in a case where there is at least one first semantic vector and the second semantic vector between the similarity is less than the similarity threshold, store the first polysemous word in the ambiguous word library.

[0014] Optionally, in the domain knowledge base construction module, the core definition includes: academic definition, subject category, application scenario, key technology and key method, and supplementary explanation; wherein the academic definition is generated by using the genus plus species difference definition method, the subject category includes primary discipline and secondary discipline, the application scenario includes core application scenario and frontier application scenario, and the supplementary explanation includes controversial points, key scholars and key institutions in the core technology.

[0015] Optionally, the screening module is further configured to perform semantic analysis on a literature title of a first literature based on the core definition, the ambiguous word library and the constraint rule in the domain knowledge base, to determine whether the first literature belongs to the target domain, perform semantic analysis on an abstract of the first literature in a case where the semantic analysis on the literature title of the first literature cannot determine whether the first literature belongs to the target domain, to determine whether the first literature belongs to the target domain, the first literature being any literature in the first literature set; and in a case where the first literature belongs to the target domain, store the first literature in the second literature set.

[0016] Optionally, the screening module is further configured to exclude a literature corresponding to an ambiguous definition of a first ambiguous word from the first literature set based on the ambiguous word library in the domain knowledge base; wherein the first ambiguous word is any word in the ambiguous word library, and the similarity between the semantic vector corresponding to the ambiguous definition and the second semantic vector is less than the similarity threshold.

[0017] The third aspect also provides a computer device, comprising a memory and a processor, at least one computer program is stored in the memory, the at least one computer program is loaded and executed by the processor, so as to execute the literature retrieval method based on large language model described in the above embodiments.

[0018] The fourth aspect also provides a computer readable storage medium, at least one computer program is stored in the computer readable storage medium, the at least one computer program is loaded and executed by the processor, so as to execute the literature retrieval method based on large language model described in the above embodiments.

[0019] In a fifth aspect, a computer program product is provided, comprising computer programs / instructions which, when executed by a processor, implement the method of the first aspect.

[0020] The technical solutions provided by the embodiments of the present disclosure have at least the following beneficial effects:

[0021] In the embodiments of the present disclosure, the natural language text from the user is obtained by the large language model; the domain knowledge base of the target domain is constructed according to the natural language text, and the domain knowledge base includes the core definition, the multi-dimensional word library, the topic retrieval formula, the ambiguity word library and the constraint rule library corresponding to the target domain; the literature retrieval is performed according to the topic retrieval formula in the domain knowledge base to obtain a first literature set of the target domain; the first literature set can comprehensively cover the related literature of the target domain, and the literature in the first literature set is filtered based on the core definition, the ambiguity word library and the constraint rule in the domain knowledge base to obtain a second literature set, and the literature in the first literature set is filtered to remove the literature in the first literature set that does not belong to the target domain, and the literature in the first literature set that belongs to the target domain is retained; and the literature analysis report is generated based on the second literature set. Since the core definition, the ambiguity word library and the constraint rule in the domain knowledge base can accurately limit the target domain, the second literature set filtered based on the core definition, the ambiguity word library and the constraint rule in the domain knowledge base can accurately retain the literature in the target domain, and therefore the literature retrieval method in the embodiments of the present disclosure can realize large-scale and high-accuracy literature retrieval. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0023] Figure 1 A flowchart of the literature retrieval method based on the large language model provided by an exemplary embodiment of the present disclosure is shown;

[0024] Figure 2 A flowchart of the literature retrieval method based on the large language model provided by another exemplary embodiment of the present disclosure is shown;

[0025] Figure 3 is a schematic diagram of the domain knowledge base construction process;

[0026] Figure 4 A structural schematic diagram of the literature retrieval method device based on the large language model provided by an exemplary embodiment of the present disclosure is shown;

[0027] Figure 5 FIG. 1 is a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Unless otherwise defined, technical terms or scientific terms used herein should be interpreted as is normally used by one of ordinary skill in the art to which the present disclosure pertains. The use of “first”, “second”, “third”, and like terms in the present patent application specification and claims does not denote any order, quantity, or importance, but is merely used to distinguish different components. Similarly, “one” or “a” and like terms do not denote a quantity limitation, but means that at least one exists. “Include” or “contain” and like terms mean that the elements or objects appearing before “include” or “contain” cover the elements or objects listed after “include” or “contain” and their equivalents, and do not exclude other elements or objects.

[0029] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings.

[0030] Figure 1 FIG. 1 is a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. Figure 1 The method comprises the following steps.

[0031] In step 101, the large language model obtains natural language text from a user.

[0032] The natural language text is used to retrieve literature in a target field. Here, at least one retrieval field (such as a time phrase, an author phrase, an organization phrase, a topic phrase, etc.) needs to be included in the natural language text, and the meaning in the natural language text needs to clearly define at least one target field.

[0033] The solution in the embodiment of the present disclosure is described taking the case where there is only one target field in the natural language text as an example. In implementation, if there are at least two target fields in the natural language text (that is, the natural language text is used to query literature in at least two target fields), the large language model can be used to first split the natural language text input by the user according to the fields, and after splitting, the natural language text for each field can be retrieved using the method in the embodiment of the present disclosure.

[0034] The large language model can be a large language model with a parameter quantity greater than or equal to 10B, for example, a Deepseek model, etc.

[0035] In step 102, a domain knowledge base of the target domain is constructed according to the natural language text.

[0036] The domain knowledge base includes core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguity word library, and a constraint rule library corresponding to the target domain.

[0037] In step 103, literature retrieval is performed based on the topic retrieval formula in the domain knowledge base to obtain a first literature set of the target domain.

[0038] In step 104, the literature in the first literature set is screened based on the core definitions, the ambiguity word library, and the constraint rules in the domain knowledge base to obtain a second literature set.

[0039] The screening of the literature in the first literature set is used to eliminate the literature in the first literature set that does not belong to the target domain and retain the literature in the first literature set that belongs to the target domain.

[0040] In step 105, a literature analysis report is generated based on the second literature set.

[0041] In the embodiments of the present disclosure, a natural language text from a user is obtained by a large language model; a domain knowledge base of a target domain is constructed according to the natural language text, and the domain knowledge base includes core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguity word library, and a constraint rule library corresponding to the target domain; literature retrieval is performed according to the topic retrieval formula in the domain knowledge base to obtain a first literature set of the target domain; the first literature set can comprehensively cover the relevant literature of the target domain; the literature in the first literature set is screened based on the core definitions, the ambiguity word library, and the constraint rules in the domain knowledge base to obtain a second literature set; the screening of the literature in the first literature set is used to eliminate the literature in the first literature set that does not belong to the target domain and retain the literature in the first literature set that belongs to the target domain; and a literature analysis report is generated based on the second literature set. Since the core definitions, the ambiguity word library, and the constraint rules in the domain knowledge base can accurately limit the target domain, the second literature set screened based on the core definitions, the ambiguity word library, and the constraint rules in the domain knowledge base can accurately retain the literature belonging to the target domain, and therefore the literature retrieval method in the embodiments of the present disclosure can achieve literature retrieval with a large range and high accuracy.

[0042] Figure 2 A flowchart of a literature retrieval method based on a large language model provided by an example embodiment of the present disclosure is shown. The method can be executed by a computer device. Referring to Figure 2 , the method includes:

[0043] In step 201, a natural language text from a user is obtained by a large language model.

[0044] The related content of step 201 is described in the foregoing step 101, and details are omitted here.

[0045] In step 202, a domain knowledge base of the target domain is constructed according to the natural language text.

[0046] The domain knowledge base includes core definitions, multi-dimensional word libraries, topic retrieval formulas, ambiguity word libraries, and constraint rule libraries corresponding to the target domain.

[0047] Figure 3 is a schematic diagram of the domain knowledge base construction process, and steps a-g are described below in combination with Figure 3 to describe step 202. Optionally, step 202 includes steps a-g.

[0048] Step a, extracting the topic phrase of the natural language text.

[0049] In implementation, step a can first split the natural language text into multiple keywords by NLP (Natural Language Processing) technology, and then map different keywords to different retrieval fields by a large language model, so as to determine the topic phrase of the natural language text. For example, the natural language text can be split into multiple keywords by NER (Named-entity recognition) technology.

[0050] Illustratively, the user input natural language text is "help me retrieve the strategy of General Qi Jiguang against the Japanese in the southeast coastal area during the Jiajing period published by Agency A from 1990 to 2000". After keyword splitting of the natural language text by NLP technology, the large language model can obtain the retrieval fields in the natural language text, including the time phrase "from 1990 to 2000", the agency phrase "Agency A", and the topic phrase "the strategy of General Qi Jiguang against the Japanese in the southeast coastal area during the Jiajing period".

[0051] Step b, splitting the topic phrase according to the time dimension, the space dimension, the person dimension, and the theme dimension.

[0052] Since the large language model has rich semantic understanding ability, the large language model can split the topic phrase according to the time dimension, the space dimension, the person dimension, and the theme dimension by analyzing the topic phrase, and the large language model also supports cross-disciplinary term expansion and high-order logic optimization when analyzing the topic phrase.

[0053] For example, in the natural language text input by the user, the topic phrase is: the strategy of resisting the Japanese in the southeast coastal area during the Jiajing period. The large language model can identify the corresponding words of the time dimension as “Jiajing period”, the words of the spatial dimension as “southeast coastal area”, the words of the person dimension as “Qi Jiliang”, and the words of the theme dimension as “resisting the Japanese strategy”.

[0054] In addition, the large language model can correctly identify the implicit relationships (such as cause-and-effect relationship, comparison relationship, and negation relationship) in the topic phrase, and accurately determine the words corresponding to the time dimension, spatial dimension, person dimension, and theme dimension.

[0055] When splitting the four dimensions, if a certain dimension in the topic phrase is missing (for example, it does not involve the person dimension), the large language model can also retain the non-empty dimensions and skip the missing dimension.

[0056] Step c, generating the core definition of the domain knowledge base according to the topic phrase.

[0057] In the embodiments of the present disclosure, the core definition includes: academic definition, discipline category, application scenario, key technology and key method, and supplementary explanation.

[0058] Among them, the academic definition is generated by using the genus plus species difference definition method, the discipline category includes primary discipline and secondary discipline, the application scenario includes core application scenario and frontier application scenario, and the supplementary explanation includes controversial points in core technology, key scholars and key institutions.

[0059] The main sources of the core definition include but are not limited to: authoritative textbooks / papers, international standards, encyclopedias.

[0060] In the case where the source of the core definition is an authoritative textbook / paper, the article (the author / year / section needs to be marked) can be directly quoted or summarized. International standards, such as ISO / IEC definition, are standardized definitions. Encyclopedias, such as the Encyclopedia Britannica or the academic section of Wikipedia, etc.

[0061] In the implementation of the large language model generating the core definition, the input of the large language model is the topic phrase. According to the input, the large language model will generate the corresponding core definition according to the standard template of the core definition. Exemplarily, the standard template of the core definition is as follows.

[0062] 1. Academic definition

[0063] - Definition text (precise expression based on genus + species difference method, avoiding metaphor or vague description)

[0064] 2. Discipline category

[0065] - Primary discipline (such as physics, computer science)

[0066] - Secondary disciplines (e.g., quantum mechanics, natural language processing)

[0067] - Related interdisciplinary fields (e.g., bioinformatics, geochemistry, etc.)

[0068] 3. Application scenarios

[0069] - Core applications (e.g., CRISPR gene editing technology for genetic disease treatment)

[0070] - Emerging directions (e.g., application of Transformer models in protein structure prediction)

[0071] 4. Main technologies / methods

[0072] - List of key technologies (e.g., PCR amplification, mass spectrometry analysis)

[0073] - Technology classification (optional, e.g., experimental techniques / computational methods / engineering approaches)

[0074] 5. Supplementary notes (optional)

[0075] - Controversial points (e.g., the particle nature of dark matter has not been confirmed)

[0076] - Key scholars / institutions (e.g., MIT's XX Laboratory leads research in this field)

[0077] By defining the generation form of core definitions through this standard template, the structured constraints of core definitions can be achieved, avoiding randomness caused by information mixing. Each definition in the template must be associated with at least one verifiable source to improve the interpretability of the definition. The template requires clear classification of disciplines and technical terms, reducing language recombination space and improving the reliability of generated core definitions.

[0078] In addition, users can also edit and modify the core definitions generated by large language models to ensure the reliability of the core definitions.

[0079] Step d: Perform synonym expansion on the words in different dimensions of the topic phrase to generate a multi-dimensional word bank for the topic phrase.

[0080] The multi-dimensional word bank includes a time word bank, a space word bank, a character word bank, and a topic word bank.

[0081] The words corresponding to the time dimension, space dimension, character dimension, and topic dimension can be obtained through synonym expansion and standardization processing to generate the corresponding word bank (the word bank includes the words obtained after synonym expansion and standardization processing, as well as the topic phrase).

[0082] For the words corresponding to the time dimension in the topic phrase, the large language model supports the function of converting the user's description of the time dimension into a standardized time expression, and also supports the function of dynamic time range calculation. The standardized time expression is mainly used to obtain a standardized time expression; while the dynamic time range calculation is to calculate the corresponding time range in the case of dynamic time description in the user's natural language text. For example, the words related to time in the topic phrase in the user's input natural language text are "from a certain year every five years", then the time range is a dynamic time range rather than a static description value, and the large language model can dynamically calculate the corresponding time according to the user's description.

[0083] For example, in the case where the words corresponding to the time dimension in the topic phrase are "Jiajing period", the large language model can analyze the time range corresponding to "Jiajing period" (1522-1566) and convert it into a standardized time representation, which can be stored in the time word library. Similarly, the word "Jiajing period" can also be expanded, for example, expanded in multiple languages to "Jiajing period". In the case where the words corresponding to the time dimension in the topic phrase are "Jiajing period", the time word library can include: Jiajing period, 1522-1566, Jiajing period, and similar words.

[0084] For the words of the spatial dimension in the topic phrase, the large language model supports place name expansion and geographical hierarchy processing. Place name expansion refers to the expansion of geographical names, such as different ways of calling the same geographical location. Geographical hierarchy processing refers to constructing a tree structure or network structure according to the administrative or natural hierarchy of place names (such as country→city→region), supporting the completion of unstructured place names to complete hierarchical chains. After place name expansion and geographical hierarchy processing, the place name expansion results and complete hierarchical chain geographical names can be stored in the spatial word library.

[0085] For the words of the person dimension in the topic phrase, the large language model supports title and nickname expansion. For example, different titles and nicknames of the same person can be stored in the person word library.

[0086] For the words of the topic dimension in the topic phrase, the large model needs to generate core terms, interdisciplinary terms, and method and technology terms when generating a topic word library based on the words.

[0087] Core terms are the words directly related to the core definition. Core terms can be extracted from professional glossaries, authoritative literature, and various encyclopedias. Core terms include, but are not limited to, official names, abbreviations, synonyms. Wildcards can also be used to cover the inflectional forms of core terms (e.g., climat* can be used to get climate, climatic).

[0088] Disciplinary cross terms refer to words that are commonly used in multiple disciplines, but may have different meanings or application methods in different contexts. When generating the subject term library, disciplinary cross terms related to the core definition also need to be generated. Different disciplinary cross terms can be generated for different disciplines, and users can also modify or add or delete disciplinary cross terms as needed in the case of errors in the disciplinary cross terms generated by the large language model. For example, the term screening standard for atmospheric science is to screen discipline-specific disciplinary cross terms, and the disciplinary cross terms for atmospheric science are, for example, “radiative forc”; The term screening standard for ecology is to exclude general terms, and the disciplinary cross terms for ecology are, for example, “species adapt”.

[0089] Method technology terms refer to methods, technologies, tools, etc. involved in implementing the core definition. When generating method technology terms, compound words are preferred. For example, “machine learn*” is generated instead of “learn*”.

[0090] In some embodiments, the subject term library can also include scenario application terms. Scenario application terms can be divided into three categories according to application scenarios: geographic scope, industry field, and population characteristics, which represent the focus of different application scenarios. The application scenario of some fields focuses on a specific location, and the scenario application term can optionally be of the geographic scope type, such as “arcticamplificat”; The application scenario of some fields focuses on a specific industry, and the scenario application term can optionally be of the industry field type, such as “renewable energy transition”; The application scenario of some fields focuses on a certain group of people with common characteristics, and the scenario application term can optionally be of the population characteristics type, such as “indigenous community”.

[0091] In the subject term library, core terms, disciplinary cross terms, and method technology terms are mandatory, and scenario application terms are optional.

[0092] In addition, the multi-dimensional term library can also include a user-manually-added artificial term library, which is a user-manually-added term. When implemented, the terms in the artificial term library can also be added to the above-mentioned four-dimensional term library according to the type. It should be noted that the weight of the terms in the artificial term library in the multi-dimensional term library is the highest.

[0093] Step e, performing semantic analysis on the words in the multi-dimensional word library to obtain an ambiguous word library in the multi-dimensional word library.

[0094] The ambiguous word library includes a plurality of ambiguous words, and the ambiguous word is a word having at least one definition different from the core definition.

[0095] Optionally, step e includes the following four steps.

[0096] First, obtaining a first semantic vector corresponding to each definition of a first ambiguous word.

[0097] The first ambiguous word is any ambiguous word in the multi-dimensional word library. The first ambiguous word has a plurality of definitions, and each definition can obtain a first semantic vector.

[0098] Before performing the first step, it is necessary to first perform semantic analysis on the words in the multi-dimensional word library to determine the ambiguous words in the multi-dimensional word library, and then perform the first step to the fourth step in step e on each ambiguous word in the multi-dimensional word library to determine the ambiguous word library.

[0099] Second, obtaining a second semantic vector corresponding to the core definition.

[0100] Here, the first and second semantic vectors can be obtained in any manner in the related art.

[0101] Third, comparing the similarity between each first semantic vector and the second semantic vector.

[0102] The similarity here can be, for example, cosine similarity.

[0103] Fourth, in the case where there is at least one first semantic vector and the second semantic vector between which the similarity is less than a similarity threshold, storing the first ambiguous word to the ambiguous word library.

[0104] The value of the similarity threshold can be an empirical value, or can be set by the user, and the present disclosure does not limit it.

[0105] If there is at least one first semantic vector and the second semantic vector between which the similarity is less than the similarity threshold, it means that the first ambiguous word has at least one definition different from the core definition, and such definition can cause ambiguity in the retrieval process, so the first ambiguous word belongs to the ambiguous word and can be stored in the ambiguous word library.

[0106] If the similarity between each first semantic vector and the second semantic vector is greater than or equal to the similarity threshold, it means that each definition of the first ambiguous word is close to the core definition, that is, the first ambiguous word does not belong to the ambiguous word and does not need to be stored in the ambiguous word library.

[0107] The first step to the fourth step described above can be used to process each polysemous word in the multi-dimensional vocabulary, so that each ambiguous word in the multi-dimensional vocabulary can be determined, and an ambiguous vocabulary is obtained.

[0108] Step f, generating a subject retrieval formula according to the multi-dimensional vocabulary.

[0109] The subject retrieval formula is the retrieval formula corresponding to the subject phrase.

[0110] In generating the subject retrieval formula according to the multi-dimensional vocabulary, a Boolean logic retrieval formula can be used to generate a corresponding retrieval formula for each non-empty dimension corresponding vocabulary one by one. If the vocabulary corresponding to a certain dimension is empty, the dimension is skipped when generating the subject retrieval formula.

[0111] In the Boolean logic retrieval formula, the logic types include AND, OR, NOT, wildcard, etc. AND is used to connect different vocabularies across dimensions, for example, the time vocabulary, the space vocabulary, the person vocabulary and the subject vocabulary can be connected by AND. OR is used to expand the vocabulary corresponding to a single dimension, for example, the words in the time vocabulary can be connected by OR. NOT is used to add exclusion items to exclude interference words, here, the exclusion items can be user-defined. Wildcard is used for root expansion, which can achieve the effect of merging words with the same root and achieve the purpose of deduplication.

[0112] Step g, receiving constraint rules defined by the user to obtain a constraint rule set.

[0113] The constraint rule set includes multiple constraint rules defined by the user, and the constraint rule set is similar to the exclusion items in the subject retrieval formula, which can limit the target field.

[0114] Through the above steps a to g, the construction of the field knowledge base can be realized. Among them, steps a to c are used to generate the core definition of the field knowledge base; step d is used to generate the multi-dimensional vocabulary of the field knowledge base, step e is used to generate the ambiguous vocabulary of the field knowledge base; step f is used to generate the subject retrieval formula of the field knowledge base; step g is used to generate the constraint rule set.

[0115] When generating different field knowledge bases, the input of the field knowledge base can be normalized, or the random algorithm can be disabled when generating, to ensure the consistency of different field knowledge bases.

[0116] It should be noted that the user can add, delete and modify any content in the field knowledge base generated by the large language model, and the content added by the user has the highest weight in the field knowledge base.

[0117] In step 203, based on the subject retrieval formula in the field knowledge base, literature retrieval is performed to obtain a first literature set of the target field.

[0118] Generally, there are multiple retrieval fields (such as time phrases, author phrases, institution phrases, topic phrases, etc.) in the natural language text input by the user, and the topic retrieval formula is only the retrieval formula corresponding to the topic phrase. In implementation, the other retrieval fields in the natural language text input by the user also generate corresponding retrieval formulas, which are combined with the topic retrieval formula to obtain a target retrieval formula. Subsequent retrieval is performed according to the target retrieval formula.

[0119] In retrieval, the user can specify at least one database, and the large language model can perform retrieval in the database searched by the user according to the target retrieval formula.

[0120] The first literature set obtained through steps 201 to 203 can comprehensively cover the related literatures of the target field, but there may be noise literatures in the first literature set, resulting in a low accuracy of the first literature set. Therefore, the literatures in the first literature set need to be screened by step 204.

[0121] In step 204, the literatures in the first literature set are screened based on the core definitions, the ambiguity word library and the constraint rules in the domain knowledge base to obtain a second literature set.

[0122] Screening the literatures in the first literature set is to remove the literatures in the first literature set that do not belong to the target field, and retain the literatures in the first literature set that belong to the target field.

[0123] In the first possible implementation manner, step 204 includes: based on the ambiguity word library in the domain knowledge base, removing the literatures corresponding to the ambiguous definitions of the first ambiguity word from the first literature set, wherein the first ambiguity word is any word in the ambiguity word library, and the similarity between the semantic vector corresponding to each ambiguous definition and the second semantic vector is less than the similarity threshold. In this way, only the literatures corresponding to the correct definition (a definition similar to the core definition) of the first ambiguity word can be retained, and the literatures corresponding to the ambiguous definitions can be removed.

[0124] Each word in the ambiguity word library can be processed in a similar manner to the first ambiguity word, so that the literatures in the first literature set that do not conform to the core definition can be removed. Since the semantic vector corresponding to each definition of any ambiguity word in the ambiguity word library is known, and the similarity between the semantic vector corresponding to each definition and the second semantic vector is also known, the ambiguous definition is also known. Since the large language model has rich semantic understanding ability, the large language model can accurately determine the literatures corresponding to each ambiguous definition in the first literature set, thereby screening.

[0125] In a second possible implementation, the step 204 comprises: performing semantic analysis on the document title of the first document based on the core definition in the domain knowledge base, the ambiguity word library and the constraint rule to determine whether the first document belongs to the target domain, performing semantic analysis on the abstract of the first document to determine whether the first document belongs to the target domain in a case that the semantic analysis on the document title of the first document cannot determine whether the first document belongs to the target domain, the first document being any one of the first document set; and storing the first document into the second document set in a case that the first document belongs to the target domain.

[0126] Since there is a case that the document title is relatively brief, the semantic analysis on the document title of the first document cannot necessarily determine whether the first document belongs to the target domain; in this case, the semantic analysis can be performed on the abstract of the first document, since the abstract of the document is self-contained and can completely convey the key information of the article, the information quantity is large enough for the language model to make a final determination, so as to determine whether the first document belongs to the target domain.

[0127] The other documents in the first document set except the first document can also be determined by using this method.

[0128] The essence of the second implementation is to determine each document in the first document set one by one, and the semantic analysis method is used in the determination to determine whether the semantics in the title or abstract of each document is similar to the core definition, whether it meets the constraint rule, and whether it belongs to the document corresponding to the ambiguity in the ambiguity word library; if the semantics in the title or abstract of a document is similar to the core definition, meets the constraint rule and does not belong to the document corresponding to the ambiguity in the ambiguity word library, it means that the document belongs to the target domain.

[0129] The first implementation and the second implementation can also be combined, that is, the first implementation is used to remove a part of the documents from the first document set, and then the second implementation is used to determine the remaining documents in the first document set one by one. In this way, the screening efficiency of the first document set can be improved, and finally the second document set can be obtained.

[0130] The first document set can achieve comprehensive coverage of the related documents in the target domain, and the second document set can accurately retain the documents belonging to the target domain, so that the literature retrieval method in the embodiment of the present disclosure can achieve large-scale and high-accuracy literature retrieval.

[0131] In step 205, a literature analysis report is generated based on the second document set.

[0132] The essence of step 205 is the process of visualizing the second literature set and generating a report. Exemplarily, the generated literature analysis report includes, but is not limited to, trend curves, knowledge maps, etc. related to the second literature set.

[0133] In implementation, first, the second literature set is converted into multi-dimensional structured data, and the multi-dimensional structured data includes time dimension, institution dimension, author dimension, journal / conference dimension, etc. These multi-dimensional data can be stored in a format suitable for visual processing, such as structured table, matrix, graph, etc.

[0134] Then, the multi-dimensional structured data can be bound with visual elements using technologies such as ECharts, and the best visualization form is adaptively matched according to the data type. For example, time series data is mapped to a line chart coordinate system to obtain a trend curve.

[0135] Author co-occurrence, institution co-occurrence, literature citation, etc. can also be extracted to construct multi-dimensional knowledge graphs such as scholar collaboration relationship graph, institution collaboration relationship graph, and literature citation relationship graph. Using network graph algorithm, the knowledge graph constructed based on citation relationship is clustered by setting intra-class density greater than or equal to 15, inter-class density less than or equal to 3, and module resolution 1.0, thereby forming a topic clustering knowledge graph.

[0136] The above literature analysis report can be dynamically visualized in the visual interface of the computer device, and the user can interact with the visualized results.

[0137] In implementation, a discussion area can also be set up, in which users can discuss the literature analysis report with the large language model to help users understand the data, improve the charts, and reveal the rules. The discussion area text can be imported and exported or stored in a way of defining topic-based conversations. Before the final output of the literature analysis report, the user finally determines the report content restriction rules, which are used to limit the content generated by the large language model within the scope of the second literature set. The large language model summarizes the content of the discussion area and generates an analysis report output by summarizing the visual charts.

[0138] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference can be made to the above method embodiments.

[0139] Figure 4 A structure schematic diagram of a literature retrieval method and device based on a large language model provided by an exemplary embodiment of the present disclosure is shown. Referring to Figure 4 The literature retrieval method 400 based on the large language model includes an acquisition module 401, a domain knowledge base construction module 402, a retrieval module 403, a screening module 404, and a generation module 405.

[0140] The obtaining module 401 is configured to obtain natural language text from a user, the natural language text being used to retrieve literature of a target field.

[0141] The domain knowledge base construction module 402 is configured to construct a domain knowledge base of the target field according to the natural language text, the domain knowledge base including core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguity word library and a constraint rule library corresponding to the target field.

[0142] The retrieving module 403 is configured to perform literature retrieval based on the topic retrieval formula in the domain knowledge base to obtain a first literature set of the target field.

[0143] The screening module 404 is configured to screen literature in the first literature set based on the core definitions, the ambiguity word library and the constraint rules in the domain knowledge base to obtain a second literature set, the screening of the literature in the first literature set being used to eliminate literature not belonging to the target field in the first literature set and retain literature belonging to the target field in the first literature set.

[0144] The generating module 405 is configured to generate a literature analysis report based on the second literature set.

[0145] Optionally, the domain knowledge base construction module 402 is further configured to extract a topic phrase of the natural language text; split the topic phrase according to a time dimension, a space dimension, a person dimension and a topic dimension; generate core definitions of the domain knowledge base according to the topic phrase; perform synonym expansion on words in different dimensions in the topic phrase to generate a multi-dimensional word library of the topic phrase, the multi-dimensional word library including a time word library, a space word library, a person word library and a topic word library; perform semantic analysis on words in the multi-dimensional word library to obtain an ambiguity word library in the multi-dimensional word library, the ambiguity word library including a plurality of ambiguous words, the ambiguous word being a word having at least one meaning inconsistent with the core definitions; generate a topic retrieval formula according to the multi-dimensional word library; and receive constraint rules defined by the user to obtain a constraint rule set.

[0146] Optionally, the domain knowledge base construction module 402 is further configured to obtain a first semantic vector corresponding to each definition of a first ambiguous word, the first ambiguous word being any word in the multi-dimensional word library; obtain a second semantic vector corresponding to the core definitions; compare the similarity between each first semantic vector and the second semantic vector one by one; and in a case where there is at least one first semantic vector and the second semantic vector between which the similarity is less than a similarity threshold, store the first ambiguous word into the ambiguity word library.

[0147] Optionally, in the domain knowledge base construction module 402, the core definition includes: academic definition, subject category, application scenario, key technology and key method, supplementary explanation; wherein, the academic definition is generated by using the genus plus difference definition method, the subject category includes first-level discipline and second-level discipline, the application scenario includes core application scenario and frontier application scenario, and the supplementary explanation includes controversial points in the core technology, key scholars and key institutions.

[0148] Optionally, the screening module 404 is further configured to perform semantic analysis on the literature title of the first literature based on the core definition, the ambiguity word library and the constraint rule in the domain knowledge base, to determine whether the first literature belongs to the target domain, and perform semantic analysis on the abstract of the first literature when it is determined that the first literature does not belong to the target domain, to determine whether the first literature belongs to the target domain, the first literature being any one of the first literature set; and store the first literature into the second literature set when the first literature belongs to the target domain.

[0149] Optionally, the screening module 404 is further configured to exclude the literature corresponding to the ambiguous definition of the first ambiguous word from the first literature set based on the ambiguity word library in the domain knowledge base; wherein, the first ambiguous word is any word in the ambiguity word library, and the similarity between the semantic vector corresponding to the ambiguous definition and the second semantic vector is less than the similarity threshold.

[0150] It should be noted that: when the literature retrieval device based on the large language model provided in the above embodiment performs literature retrieval, only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the literature retrieval device based on the large language model provided in the above embodiment and the literature retrieval method based on the large language model embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0151] The division of the modules in the embodiments of the present disclosure is illustrative, and is only a logical functional division. In actual implementation, there can be another division manner. In addition, each functional module in each embodiment of the present disclosure can be integrated in one processor, or can be physically separated, or two or more modules can be integrated into one module. The above integrated module can be realized in the form of hardware or in the form of a software functional module.

[0152] The integrated module, if implemented in the form of a software function module and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present disclosure, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an end device (which can be a personal computer, a mobile phone, or a communication device, etc.) or a processor (processor) to perform all or part of the steps of the methods according to the embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (read-only memory, ROM), random access memory (random access memory, RAM), magnetic disk or optical disk, and various media that can store program codes.

[0153] Figure 5 is a structural schematic diagram of a computer device provided by an embodiment of the present disclosure. As shown in Figure 5 the computer device 500 includes a processor 501 and a memory 502.

[0154] The processor 501 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 can be implemented in at least one of the hardware forms of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 501 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 501 can be integrated on a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content to be displayed by the display screen. In some embodiments, the processor 501 can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0155] The memory 502 can include one or more computer-readable storage media. The memory 502 can also include high-speed random access memory and non-volatile, computer-readable storage media such as one or more magnetic disk storage devices, optical storage devices, flash memory devices, or similar storage medium. In some embodiments, the non-transitory computer-readable storage media of the memory 502 is used to store at least one instruction for being executed by the processor 501 to implement the method of retrieving literature based on a large language model provided in the embodiments of the present disclosure.

[0156] Those skilled in the art can understand that the structure shown in the above is not a limitation on the computer device 500, and the computer device 500 can include more or fewer components than those shown, or combine certain components, or adopt a different arrangement of components. Figure 5 The computer device 500 shown in the above is not a limitation on the computer device 500, and the computer device 500 can include more or fewer components than those shown, or combine certain components, or adopt a different arrangement of components.

[0157] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a computer device, the computer device is enabled to perform the method of retrieving literature based on a large language model provided in the embodiments of the present disclosure.

[0158] The embodiments of the present disclosure further provide a computer program product, including computer programs / instructions, when the computer programs / instructions are executed by a processor, the method of retrieving literature based on a large language model provided in the embodiments of the present disclosure is implemented.

[0159] The above only describes optional embodiments of the present disclosure, and does not limit the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A literature retrieval method based on a large language model, characterized in that, The method comprises: The large language model obtains natural language text from a user, the natural language text being used to retrieve literature of a target field; According to the natural language text, a domain knowledge base of the target field is constructed, the domain knowledge base comprising core definitions corresponding to the target field, a multi-dimensional word library, a topic retrieval formula, an ambiguous word library and a constraint rule library; Based on the topic retrieval formula in the domain knowledge base, literature retrieval is performed to obtain a first literature set of the target field; Based on the core definitions, the ambiguous word library and the constraint rules in the domain knowledge base, the literature in the first literature set is screened to obtain a second literature set, the screening of the literature in the first literature set being used to eliminate literature in the first literature set that does not belong to the target field and retain literature in the first literature set that belongs to the target field; A literature analysis report is generated based on the second literature set; The construction of the domain knowledge base of the target field according to the natural language text comprises: Extracting a topic phrase from the natural language text; Splitting the topic phrase according to time dimension, space dimension, person dimension and topic dimension; Generating core definitions of the domain knowledge base according to the topic phrase; Performing synonym expansion on words of different dimensions in the topic phrase to generate a multi-dimensional word library of the topic phrase, the multi-dimensional word library comprising a time word library, a space word library, a person word library and a topic word library; Performing semantic analysis on the words in the multi-dimensional word library to obtain an ambiguous word library in the multi-dimensional word library, the ambiguous word library comprising a plurality of ambiguous words, the ambiguous words being words having at least one meaning that does not conform to the core definitions; Generating a topic retrieval formula according to the multi-dimensional word library; Receiving constraint rules defined by a user to obtain a constraint rule set; The semantic analysis of the words in the multi-dimensional word library to obtain the ambiguous word library in the multi-dimensional word library comprises: Obtaining a first semantic vector corresponding to each definition of a first polysemous word, the first polysemous word being any polysemous word in the multi-dimensional word library; Obtaining a second semantic vector corresponding to the core definitions; Comparing the similarity between each first semantic vector and the second semantic vector one by one; In a case where the similarity between at least one first semantic vector and the second semantic vector is less than a similarity threshold, the first polysemous word is stored in the ambiguous word library; The core definitions comprise academic definitions, subject categories, application scenarios, key technologies and key methods, and supplementary explanations; The subject categories comprise primary subjects and secondary subjects, the application scenarios comprise core application scenarios and frontier application scenarios, and the supplementary explanations comprise controversial points in core technologies, key scholars and key institutions.

2. The method of claim 1, wherein, The screening of the literature in the first literature set based on the core definitions, the ambiguous word library and the constraint rules in the domain knowledge base to obtain the second literature set comprises: perform semantic analysis on the title of the first document based on the core definition, the ambiguity word library and the constraint rule in the field knowledge base to determine whether the first document belongs to the target field, and perform semantic analysis on the abstract of the first document to determine whether the first document belongs to the target field in a case that the semantic analysis on the title of the first document cannot determine whether the first document belongs to the target field, the first document being any one of the first document set; in a case that the first document belongs to the target field, store the first document into the second document set.

3. The method of claim 1, wherein, The filtering of the documents in the first document set based on the core definition, the ambiguity word library and the constraint rule in the field knowledge base to obtain the second document set comprises: eliminating the document corresponding to the ambiguous definition of the first ambiguous word from the first document set based on the ambiguity word library in the field knowledge base; wherein the first ambiguous word is any word in the ambiguity word library, and the similarity between the semantic vector corresponding to the ambiguous definition and the second semantic vector is less than the similarity threshold.

4. A document search device based on a large language model, characterized by, The device comprises: an acquisition module configured to acquire a natural language text from a user by a large language model, the natural language text being used to retrieve documents of a target field; a field knowledge base construction module configured to construct a field knowledge base of the target field according to the natural language text, the field knowledge base comprising core definitions, a multi-dimensional word library, a topic retrieval formula, an ambiguity word library and a constraint rule library corresponding to the target field; a retrieval module configured to perform document retrieval based on the topic retrieval formula in the field knowledge base to obtain a first document set of the target field; a filtering module configured to filter the documents in the first document set based on the core definition, the ambiguity word library and the constraint rule in the field knowledge base to obtain a second document set, the filtering of the documents in the first document set being used to eliminate the documents in the first document set that do not belong to the target field and retain the documents in the first document set that belong to the target field; a generation module configured to generate a document analysis report based on the second document set; the field knowledge base construction module is further configured to extract a topic phrase of the natural language text; split the topic phrase according to time dimension, space dimension, person dimension and topic dimension; generate the core definition of the field knowledge base according to the topic phrase; perform synonym expansion on the words in different dimensions in the topic phrase to generate a multi-dimensional word library of the topic phrase, the multi-dimensional word library comprising a time word library, a space word library, a person word library and a topic word library; perform semantic analysis on the words in the multi-dimensional word library to obtain an ambiguity word library in the multi-dimensional word library, the ambiguity word library comprising a plurality of ambiguous words, the ambiguous word being a word having at least one meaning inconsistent with the core definition; generate a topic retrieval formula according to the multi-dimensional word library; receive constraint rules defined by a user to obtain a constraint rule set; The domain knowledge base construction module is also configured to obtain a first semantic vector corresponding to each definition of the first polysemous word, the first polysemous word being any polysemous word in the multi-dimensional word library; obtain a second semantic vector corresponding to the core definition; compare the similarity between each of the first semantic vectors and the second semantic vector one by one; in the case where there is at least one first semantic vector and the second semantic vector between the similarity is less than the similarity threshold, the first polysemous word is stored in the ambiguity word library; In the domain knowledge base construction module, the core definition includes: academic definition, subject category, application scenario, key technology and key method, supplementary explanation; Wherein, the academic definition is generated by using the genus plus species difference definition method, the subject category includes primary discipline and secondary discipline, the application scenario includes core application scenario and frontier application scenario, and the supplementary explanation includes controversial points in core technology, key scholars and key institutions.

5. A computer device, comprising: The computer device comprises a memory and a processor, at least one computer program is stored in the memory, the at least one computer program is loaded and executed by the processor to realize the method of any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, the at least one computer program is loaded and executed by the processor to realize the method of any one of claims 1 to 3.

7. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to realize the method of any one of claims 1 to 3.

Citation Information

Patent Citations

  • Disambiguation method and systemin traditional Chinese medicine text word segmentation process, equipment and medium

    CN110502750A

  • Literature screening method and device, electronic equipment and storage medium

    CN119226432A