Entity identification method and device for nutritional and healthy food field

By combining a hybrid knowledge base architecture with a large language model, the accuracy and efficiency issues of named entity recognition in the field of nutritional and healthy food are solved, high-precision entity recognition and dynamic knowledge updating are achieved, and the quality and efficiency of information extraction are improved.

CN120706427AActive Publication Date: 2025-09-26SHANGHAI MENGNIU BIOTECHNOLOGY R & D CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511212017.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing named entity recognition tools in the field of nutritional and healthy food cannot accurately identify professional terms, the construction cost of static knowledge bases is high and the real-time performance is poor, and biomedical tools are seriously ambiguous when applied in the field of nutritional and health, resulting in low data extraction accuracy.

Method used

By adopting a hybrid knowledge base architecture and a large language model, combining a structured database with a vector library, and combining preliminary recognition with enhanced recognition, the large language model is used for context-aware enhanced verification to build an entity recognition system and achieve high-precision and efficient entity recognition.

Benefits of technology

It improves the accuracy and efficiency of entity recognition in the field of nutritional and healthy food, reduces the risk of misjudgment by large language models, and achieves high-quality information extraction and dynamic updating of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706427A_ABST
    Figure CN120706427A_ABST
Patent Text Reader

Abstract

The invention provides an entity recognition method and device for the field of nutritional and healthy food, and relates to the technical field of food science, and the method comprises the steps: receiving a to-be-processed text containing a named entity in the field of nutritional and healthy food; extracting a preliminary recognition entity in the to-be-processed text; inputting the preliminary recognition entity into the large language model to obtain an enhanced recognition entity output by the large language model; the big language model performs enhancement verification on the preliminary recognition entity based on the candidate entity information and the context information of the preliminary recognition entity in the to-be-processed text to generate an enhanced recognition entity; and determining a named entity in the to-be-processed text based on the enhanced recognition entity. According to the method and the device provided by the invention, preliminary recognition and enhanced recognition are combined, the preliminary recognition entity is rapidly extracted, and a large language model retrieval enhanced generation technology is introduced to carry out enhanced verification of context perception, so that the accuracy and the efficiency of entity recognition in the field of nutritional and healthy foods are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food science and technology, and in particular to an entity recognition method and device for the field of nutritious and healthy food. Background Art

[0002] With the rapid development of the nutritional and health food industry, accurately and efficiently extracting key information from massive amounts of text (such as scientific research literature, product labels, regulatory standards, patents, and market information) has become a key component in driving R&D innovation, ensuring product compliance and safety, and gaining insight into market trends. Named Entity Recognition (NER) technology, particularly for identifying specialized terms such as nutrients, bioactive ingredients, and functional ingredients, is fundamental and crucial for enabling downstream applications such as knowledge graph construction, intelligent question-answering, and public opinion monitoring.

[0003] However, current named entity recognition technology applied to the field of nutrition and health foods still faces multiple challenges. First, the named entity recognition tools used in related technologies are trained based on general corpora and cannot accurately identify specialized terms in the field of nutrition and health foods. Second, while recognition methods based on static knowledge bases or knowledge graphs can achieve high accuracy on specific datasets, the construction and maintenance of these knowledge bases relies on extensive manual effort, which is costly and time-consuming, resulting in insufficient knowledge coverage and poor real-time performance. Third, some tools that have performed well in the biomedical field (such as sciSpaCy) are trained on data that focuses on genes, proteins, diseases, and drugs. When applied to the field of nutrition and health, ambiguity often occurs and it is difficult to effectively distinguish them based on context, which seriously affects the accuracy of data extraction.

[0004] Therefore, how to improve the accuracy and efficiency of entity recognition in the field of nutritional and healthy food has become a technical problem that needs to be urgently solved in the industry. Summary of the Invention

[0005] The present invention provides an entity recognition method and device for the field of nutritious and healthy food, which are used to solve the technical problem of how to improve the accuracy and efficiency of entity recognition in the field of nutritious and healthy food.

[0006] This application provides an entity recognition method for the field of nutritious and healthy food, including: receiving a text to be processed containing named entities in the field of nutrition, health and food; Extracting preliminary recognized entities from the text to be processed; Inputting the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on candidate entity information and context information of the preliminary recognized entity in the text to be processed to generate the enhanced recognized entity; the candidate entity information is obtained after searching the knowledge base system based on the preliminary recognized entity; Determine named entities in the to-be-processed text based on the enhanced recognition entity.

[0007] In some embodiments, the knowledge base system includes a structured database and an entity vector library; The structured database is used to perform structured storage on entities, synonyms of entities, and relationships between entities in the ontology library and the synonym library; The entity vector library is used to store the representation vectors corresponding to the entities in the ontology library and the representation vectors corresponding to the synonyms in the synonym library; The ontology library is constructed based on the food science ontology and the unified medical language system; The thesaurus is constructed based on supplement terms, ingredient terms, and product terms.

[0008] In some embodiments, the candidate entity information is determined based on the following steps: Based on the entity identifier, entity name, entity type, and synonym list of the initially identified entity, querying the structured database to obtain a first query result; Based on the contextual semantic features of the preliminary identified entity in the text to be processed, querying the entity vector library to obtain a second query result; The candidate entity information is generated based on the first query result and the second query result.

[0009] In some embodiments, the knowledge base system is updated based on the following steps: Obtaining the enhanced recognition entity output by the large language model; If the enhanced recognition entity does not exist in the knowledge base system and the semantic similarity between the enhanced recognition entity and any entity in the knowledge base system is greater than a preset threshold, the enhanced recognition entity is added to the knowledge base system as a synonym.

[0010] In some embodiments, the knowledge base system is updated based on the following steps: monitoring the ontology library and / or the synonym library; In the case where a new entity is added to the ontology library and / or the synonym library, the new entity is added to the knowledge base system.

[0011] In some embodiments, inputting the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model includes: Inputting a preset prompt word, the candidate entity information, and context information of the initially recognized entity in the text to be processed into the large language model, so that the large language model performs enhanced verification on the initially recognized entity and outputs an enhanced recognized entity; The preset prompt word is used to guide the large language model to disambiguate any of the preliminary recognized entities based on the context information when any of the preliminary recognized entities corresponds to multiple candidate entity information.

[0012] In some embodiments, the preset prompt word is also used to guide the large language model to perform a normalized entity representation on the enhanced recognition entity.

[0013] This application provides an entity recognition device for the field of nutritious and healthy food, comprising: A text receiving module, used for receiving a text to be processed containing named entities in the field of nutritional and healthy food; A preliminary recognition module, used for extracting preliminary recognition entities in the text to be processed; An enhanced verification module is configured to input the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on candidate entity information and contextual information of the preliminary recognized entity in the text to be processed to generate the enhanced recognized entity; the candidate entity information is obtained after searching the knowledge base system for the preliminary recognized entity; An entity output module is used to determine the named entities in the text to be processed based on the enhanced recognition entity.

[0014] The present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the entity recognition method for the field of nutritional and healthy food is implemented.

[0015] The present application provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the entity recognition method for the field of nutritional and healthy food is implemented.

[0016] The entity recognition method and device for the field of nutritional and healthy food provided by the present invention receive a text to be processed containing named entities in the field of nutritional and healthy food; extract preliminary recognition entities from the text to be processed; input the preliminary recognition entities into a large language model to obtain enhanced recognition entities output by the large language model; determine the named entities in the text to be processed based on the enhanced recognition entities; due to the use of a combination of preliminary recognition and enhanced recognition, the language processing tools in the relevant technology can be used to quickly extract preliminary recognition entities, and the large language model retrieval enhancement generation technology is introduced to perform context-aware enhanced verification, which can effectively solve the problem of complex and ambiguous entity terms in the professional field of nutritional and healthy food. Compared with traditional entity recognition models, the accuracy and efficiency of entity recognition in the field of nutritional and healthy food are greatly improved. In addition, through preliminary recognition and retrieval in the knowledge base system, a clear candidate range and judgment basis are provided for the large language model, reducing the risk of it generating "hallucinations" and improving processing efficiency, ultimately achieving high-quality and structured information extraction of text in the field of nutritional and healthy food. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is one of the flow charts of the entity recognition method for the field of nutritional and healthy food provided by the present invention.

[0020] Figure 2 This is the second flow chart of the entity recognition method for the field of nutritional and healthy food provided by the present invention.

[0021] Figure 3 It is a structural schematic diagram of the entity recognition device for the field of nutritional and healthy food provided by the present invention.

[0022] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first," "second," and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps, units, or modules is not necessarily limited to those steps, units, or modules that are explicitly listed, but may include other steps, units, or modules that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0025] In response to the deficiencies in the related art, the present invention provides an entity recognition method and system for the field of nutritional and healthy foods. The entity recognition system is built based on a hybrid knowledge base architecture and a large language model (LLM). The hybrid knowledge representation of a structured database and a vector library is utilized, combined with a domain-optimized named entity recognition (NER) model and the retrieval-augmented generation (RAG) technology of the large language model to achieve a high-precision, efficient, and dynamically updated entity recognition and normalization system, which is suitable for professional terminology recognition, normalization, and dynamic knowledge base update in the field of nutritional and healthy foods.

[0026] Figure 1 This is one of the flow charts of the entity recognition method for the field of nutritional and healthy food provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .

[0027] Step 110: Receive the text to be processed containing named entities in the field of nutritional and healthy food.

[0028] Specifically, the entity recognition method provided by the embodiments of the present invention is applied to the field of nutritional and health foods, and is implemented by an entity recognition device. This device can be implemented through software, such as an entity recognition program running on a computer, or through hardware, such as a computer, server, or cloud platform that executes the entity recognition method.

[0029] Named entities in the field of nutritional and health foods are words or phrases with clear, independent meanings within the field. These entities form the basic units of the knowledge system in this field. These named entities can include foods, additives, supplements, nutrients, raw materials, bioactive molecules, nutritional ingredients, efficacy, health benefits, proteins, genes, and microorganisms.

[0030] The text to be processed refers to text data containing named entities in the field of nutritional and health foods. This text can come from a variety of sources, such as academic papers, patent documents, clinical trial reports, food safety standards documents, product instructions, online health forum discussions, or news reports.

[0031] Before executing the method provided by the embodiment of the present invention, the system or device may pre-process the texts in different formats, for example, by parsing and extracting the plain text contents therein to serve as input for subsequent steps.

[0032] Step 120: extract preliminary recognized entities from the text to be processed.

[0033] Specifically, after receiving the text to be processed, the text can be quickly scanned and analyzed to identify any named entities that may exist therein. These entities are referred to as preliminary recognized entities. It should be understood that "preliminary" means that the output of this step is a candidate result, which may have problems such as incomplete recognition, incorrect boundary demarcation, or inaccurate classification. Its main purpose is to provide candidate targets for subsequent precise verification, thereby improving the processing efficiency of the overall method.

[0034] The specific extraction method may adopt various natural language processing (NLP) technologies well known to those skilled in the art.

[0035] In an optional implementation, a dictionary and rule-based approach may be used to construct a dictionary containing some domain terms to perform string matching on the text to be processed.

[0036] In another optional embodiment, a statistical learning model may be used, such as a conditional random field (CRF), a hidden Markov model (HMM), or a support vector machine (SVM).

[0037] In a more preferred embodiment, a pre-trained deep learning model can be used, such as a bidirectional long short-term memory network-conditional random field (BiLSTM-CRF) model, or a NER model that is pre-trained on a general corpus and fine-tuned on a small-scale annotated corpus in the field of nutrition and health, to achieve rapid labeling and extraction of potential entities in the text.

[0038] For example, for the text to be processed, "Study the effects of GDCA on the gut microbiota of diabetic patients," this step can extract phrases such as "GDCA," "diabetes," and "gut microbiota" from the text as preliminary entities. At this point, the system may initially label "GDCA" as an entity based on the knowledge of the pre-trained model, but its specific category may be unclear or incorrectly labeled.

[0039] Step 130: Input the preliminary recognized entity into the large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on the candidate entity information and the context information of the preliminary recognized entity in the text to be processed to generate an enhanced recognized entity; the candidate entity information is obtained after searching the knowledge base system based on the preliminary recognized entity.

[0040] Specifically, the deep semantic understanding and reasoning capabilities of the large language model can be leveraged to perform enhanced verification on the initially identified entities obtained in the previous step. This step does not simply hand off the initially identified entities directly to the large language model for processing. Instead, it leverages the large language model's retrieval-enhanced generation technology to implement a closed-loop "retrieval-verification-generation" logic. Enhanced verification can include verification, disambiguation, completion, and normalization.

[0041] During the search phase, a knowledge base system can be pre-built. This knowledge base system can be a database or knowledge graph that stores extensive expertise in the field of nutritional and health foods. This knowledge base system can include, but is not limited to, standard entity names, synonyms, abbreviations, classification information (e.g., whether it is a bioactive molecule or a gene), and relationships between entities (e.g., whether a certain ingredient can treat a certain disease). For example, the knowledge base system can integrate multiple public, authoritative databases, such as the Food Ontology (FooON), the Unified Medical Language System (UMLS), the Food Database (FooDB), and the Dietary Supplement Label Database (DSLD). When the input is the preliminarily recognized entity "GDCA," the system searches the knowledge base system. Search methods can include exact string matching, fuzzy matching, or index-based searching. The search results are candidate entity information. For example, if the knowledge base system discovers that "GDCA" is a common abbreviation for two different entities, it will return two candidate entity information: Candidate 1: Name "Glycodeoxycholic acid" and Category "Biologically active molecule / bile acid." Candidate 2: Name "GNAT3" and Category "Gene / Protein."

[0042] During the validation phase, a large language model refers to a deep learning model with large parameters and powerful text understanding and generation capabilities, such as the DeepSeek series of models, the LLaMA series of models, or other similar models known to those skilled in the art. Contextual information refers to the words, phrases, or sentences before and after the initial identification entity in the text being processed. For "GDCA," the contextual information is "Study the effects of... on the intestinal flora of diabetic patients."

[0043] The large language model receives at least three types of information as input: (1) the preliminary recognition entity itself ("GDCA"); (2) candidate entity information retrieved from the knowledge base system ("glycodeoxycholic acid" and "GNAT3" and their respective categories); and (3) contextual information of the entity in the original text ("...diabetes...intestinal flora..."). The large language model uses its powerful semantic reasoning ability to analyze the semantic relevance between the contextual information and the information of each candidate entity. For example, the large language model can understand that "diabetes" and "intestinal flora" are highly related to metabolism and digestion, while "glycodeoxycholic acid", as a bile acid, has a function closely related to this; in contrast, "GNAT3", as a gene related to taste perception, has a significantly lower semantic relevance to the context. Based on this reasoning analysis, the large language model will make a judgment and determine that the most likely true meaning of "GDCA" in this context is "glycodeoxycholic acid". Finally, the large language model outputs the enhanced recognition entity. The enhanced recognition entity is the result of verification and disambiguation, and its information is more accurate and rich. For example, the output result may be structured data indicating that "GDCA" in the original text is identified as the entity "Glycodeoxycholic Acid", whose type is "Biologically Active Molecule", and may be associated with a unique identifier in the knowledge base.

[0044] During the generation phase, the large language model can normalize entity representations for enhanced recognition entities, and can also perform synonym expansion or near-sense association.

[0045] Step 140: Determine named entities in the text to be processed based on the enhanced entity recognition.

[0046] Specifically, the enhanced recognition entities output by the large language model are used as named entities in the text to be processed and are sorted and output.

[0047] The final output can take many forms. For example, the identified named entities can be highlighted in the original text, along with their standard names and categories. Alternatively, a structured list can be generated detailing all the named entities contained in the text.

[0048] The entity recognition method for the field of nutritional and healthy food provided by the embodiment of the present invention receives a text to be processed containing named entities in the field of nutritional and healthy food; extracts preliminary recognition entities from the text to be processed; inputs the preliminary recognition entities into a large language model to obtain enhanced recognition entities output by the large language model; determines the named entities in the text to be processed based on the enhanced recognition entities; due to the combination of preliminary recognition and enhanced recognition, the language processing tools in the relevant technology can be used to quickly extract preliminary recognition entities, and the large language model retrieval enhancement generation technology is introduced to perform context-aware enhanced verification, which can effectively solve the problem of complex and ambiguous entity terms in the professional field of nutritional and healthy food. Compared with the traditional entity recognition model, it greatly improves the accuracy and efficiency of entity recognition in the field of nutritional and healthy food. In addition, through preliminary recognition and retrieval in the knowledge base system, a clear candidate range and judgment basis are provided for the large language model, reducing the risk of it generating "hallucinations" and improving processing efficiency, ultimately achieving high-quality and structured information extraction of text in the field of nutritional and healthy food.

[0049] It should be noted that each embodiment of the present invention can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0050] In some embodiments, the knowledge base system includes a structured database and an entity vector library; The structured database is used to store entities, synonyms of entities, and relationships between entities in the ontology library and synonym library. The entity vector library is used to store the representation vectors corresponding to entities in the ontology library and the representation vectors corresponding to synonyms in the synonym library; The ontology library is built based on the Food Science Ontology and the Unified Medical Language System; The thesaurus is built based on supplement terms, ingredient terms, and product terms.

[0051] Specifically, the ontology library is built based on the Food Science Ontology (FoodON) and the Unified Medical Language System (UMLS). The ontology library provides a systematic, hierarchical knowledge framework that defines the core concepts within the field and their inheritance, composition, and other relationships. FoodON was chosen because it focuses on the food sector and provides a detailed classification system ranging from agricultural production to chemical composition. UMLS, on the other hand, is a large-scale meta-knowledge base that integrates hundreds of biomedical vocabularies and ontologies. It introduces a rich set of medical and health-related concepts such as diseases, symptoms, genes, and drugs to the system and provides standard mapping relationships between concepts. By integrating the two, the ontology library can provide a comprehensive and rigorous knowledge framework for the field of nutritional and healthy foods.

[0052] The synonym database is constructed based on supplement terminology, ingredient terminology, and product terminology. Unlike the standardization of the ontology database, the synonym database focuses more on collecting non-standard and diverse terminology expressions that are widely used in the real world. Its data sources may include but are not limited to: the Dietary Supplement Label Database (DSLD), which contains a large number of ingredient names labeled in commercial products; the Food Database (FooDB), which contains extremely detailed food chemical components and their synonyms; and product common names, ingredient abbreviations, etc. mined from relevant industry websites and literature. The establishment of a synonym database greatly enhances the system's ability to understand colloquial and informal text, and is an important guarantee for improving the recall rate of entity recognition.

[0053] Entities, synonyms of entities, and relationships between entities in the ontology library and thesaurus can be extracted, and entity tables, synonym tables, and relationship tables can be constructed respectively, and they are structured and stored to obtain a structured database. The entity table is used to store core entities from the ontology library. The synonym table is used to store synonyms, aliases, abbreviations, etc. from the thesaurus. The relationship table is used to store the relationships between entities, forming the basic structure of the knowledge graph. Structured storage refers to organizing and storing information according to a predefined, strictly constrained data model (such as a relational model) to facilitate efficient and accurate query, update, and management. The structured database can be implemented using a relational database, such as PostgreSQL, MySQL, or Oracle. Using PostgreSQL is a better choice in this embodiment because it has good stability and scalability.

[0054] At the same time, you can also build an entity vector library to store representation vectors corresponding to entities in the ontology library and representation vectors corresponding to synonyms in the synonym library. Representation vectors (also known as word embeddings) are a technique that converts text (in this case, entity names or synonyms) into low-dimensional, dense real-valued vectors. These vectors capture the semantic information of the text, ensuring that semantically similar words are close in distance in the vector space. By mapping entities and their synonyms into representation vectors, efficient semantic similarity retrieval is supported. The entity vector library can be implemented using pgVector.

[0055] In a preferred embodiment, the entity vector library can be integrated with a structured database. For example, when using PostgreSQL as a structured database, installing the pgVector extension enables the database itself to store and efficiently query vectors. This allows entity representation vectors to be stored directly as a field in an entity table or synonym table, enabling unified management of structured data and unstructured semantic information.

[0056] The entity recognition method for the field of nutritional and healthy food provided by the embodiment of the present invention can provide two retrieval query methods by constructing a structured database and an entity vector library, providing high-quality, high-coverage and diverse forms of candidate entity information for the subsequent enhanced verification step of the large language model, thereby ensuring that even when faced with rare, ambiguous or non-standard entities, the system can still have a high probability of finding the correct candidate entity, laying a solid foundation for the accuracy of the final recognition results.

[0057] In some embodiments, candidate entity information is determined based on the following steps: Based on the entity identifier, entity name, entity type, and synonym list of the initially identified entity, a query is performed in the structured database to obtain a first query result; Based on the contextual semantic features of the preliminarily identified entity in the text to be processed, a query is performed in the entity vector library to obtain a second query result; Based on the first query result and the second query result, candidate entity information is generated.

[0058] Specifically, an embodiment of the present invention provides a dual-channel hybrid retrieval method that combines structured precise query with vectorized semantic retrieval.

[0059] On the one hand, a query can be performed in a structured database based on the entity identifier, entity name, entity type, and synonym list of the preliminarily identified entity to obtain a first query result. For example, the text string of the preliminarily identified entity (e.g., "GDCA") is used as a query keyword to search multiple tables in the structured database. The query target tables include at least the entity table and the synonym table. The query is performed in the "Entity Name" field of the entity. The query is performed in the "Synonym" or "Abbreviation" field of the synonym table, which is key to finding abbreviations and aliases.

[0060] This query is an exact match or index-based fast search process, which aims to find all entities that are literally related to the preliminarily identified entity and have been clearly recorded in the knowledge base. The first query result returned is one or more structured entity records. For example, for the preliminarily identified entity "GDCA", a query in a structured database may result in two results because "GDCA" is the abbreviation of both "glycodeoxycholic acid" and "GNAT3", which together constitute the first query result. This result has high recall because it ensures that all candidates with literal possibilities are taken into consideration.

[0061] On the other hand, based on the contextual semantic features of the preliminarily identified entity in the text to be processed, a query is performed in the entity vector library to obtain a second query result.

[0062] This query is a vector retrieval based on semantic similarity. First, contextual semantic features need to be extracted. Specifically, the preliminarily identified entity and its adjacent words (i.e., context) in the text to be processed are concatenated to form a text fragment containing the context (e.g., "Study the effect of GDCA on the intestinal flora of diabetic patients"). Then, the text fragment is encoded into a query vector using the same representation vector model as used when building the entity vector library. Subsequently, a similarity search is performed in the entity vector library based on the query vector. The retrieval process usually calculates the cosine similarity or Euclidean distance between the query vector and each entity representation vector stored in the library. The system returns a list of entities sorted by similarity score from high to low, which is the second query result. This result reflects which entities in the library are most semantically consistent with the current context.

[0063] Continuing with the example of "GDCA," the query vector formed by its context "diabetes" and "gut flora" is closer in semantic space to concept vectors for concepts like "bile acid," "metabolism," and "digestion," but farther from vectors for concepts like "taste receptors" and "genes." Therefore, the second query result might be ranked with "glycodeoxycholic acid" (similarity 0.91) and "bile acid" (similarity 0.88) at the top, while "GNAT3" (similarity 0.12) is very far behind.

[0064] Finally, the first query result and the second query result are fused to generate high-quality candidate entity information that is ultimately provided to the large language model. In a preferred embodiment, the first query result and the second query result can be cross-validated or screened. For example, the semantic similarity scores in the second query result can be used to re-rank the candidates in the first query result. In the example of "GDCA", both "glycodeoxycholic acid" and "GNAT3" are in the first query result, but because the former has a much higher similarity score than the latter in the second query result, "glycodeoxycholic acid" will be ranked first in the fused list.

[0065] The entity recognition method for the nutritional and health food sector, provided by the embodiments of the present invention, combines the precision of structured databases with the semantic flexibility of entity vector libraries. Structured queries ensure that no candidate entities that literally match are missed, guaranteeing both recall and basic accuracy. Vector semantic queries, on the other hand, leverage contextual information to significantly improve the relevance of search results. This hybrid search mechanism provides a pre-screened and intelligently ranked list of high-quality candidates for subsequent large language models, significantly reducing the decision-making difficulty of the large language model and enabling it to reason within a smaller, more relevant scope, thereby improving the ultimate accuracy and robustness of the entire entity recognition method.

[0066] In some embodiments, the knowledge base system is updated based on the following steps: Obtain enhanced recognition entities output by large language models; When the enhanced recognition entity does not exist in the knowledge base system and the semantic similarity between the enhanced recognition entity and any entity in the knowledge base system is greater than a preset threshold, the enhanced recognition entity is added to the knowledge base system as a synonym.

[0067] Specifically, an embodiment of the present invention provides an incremental update method to enable the knowledge base system to have self-learning and dynamic evolution capabilities.

[0068] Enhanced entity recognition is a high-confidence entity recognition result that has been deeply verified and disambiguated using a large language model. If an enhanced entity does not exist in the knowledge base system, it indicates that the enhanced entity is a new entity.

[0069] The enhanced entity is compared for semantic similarity with various entities in the knowledge base. If the semantic similarity between the enhanced entity and any entity exceeds a preset threshold (which can be set as needed), it indicates that the enhanced entity is highly semantically consistent with the entity. The enhanced entity can be added to the knowledge base as a synonym, while historical records are retained to ensure data consistency.

[0070] The entity recognition method for the nutritional and health food sector, provided by the embodiments of the present invention, can capture new entities emerging within the sector from processed data in real time and automatically integrate them into the knowledge base system while meeting strict quality control requirements. This significantly addresses the information lag problem of traditional static knowledge bases, ensures the timeliness of the knowledge base system, and reduces the cost and cycle of manual knowledge base maintenance. This allows the performance of the entire entity recognition method to continuously improve over time and as more data is processed, thus possessing extremely high practical value.

[0071] In some embodiments, the knowledge base system is updated based on the following steps: Monitor the ontology and / or thesaurus; When a new entity is added to the ontology library and / or thesaurus, the new entity is added to the knowledge base system.

[0072] Specifically, an embodiment of the present invention provides a periodic alignment method to achieve external data-driven knowledge base system updates.

[0073] You can monitor the external data sources corresponding to the ontology library and / or thesaurus. This can be done in two ways: one is to set up a scheduled task to automatically trigger monitoring at intervals; the other is to set up a listening service to automatically trigger monitoring when an update event is detected in the ontology library or thesaurus.

[0074] Monitoring mainly includes querying the latest version information or update log of the external data source through interface calls, or accessing the server of the external data source to check whether there is an updated data file.

[0075] When new entities are identified in the ontology library and / or synonym library, the new entities are automatically captured and added to the knowledge base system.

[0076] To facilitate data comparison, you can also add an update timestamp field (last_updated) to the knowledge base system, use database triggers to automatically update it, and use change logs to record entity addition, deletion, and modification operations and related information.

[0077] The entity recognition method for the nutritional and health food sector provided by this embodiment of the present invention updates the knowledge base system by monitoring the ontology library and / or synonym library, ensuring rapid response to new entities and non-standard expressions. The periodic alignment method provided by this embodiment, combined with the incremental update method described in the above embodiment, constructs a dynamic knowledge system that is both robust and flexible, capable of continuous evolution. This enables the entire entity recognition method to maintain high performance and reliability over the long term.

[0078] In some embodiments, inputting the preliminary recognized entity into the large language model to obtain the enhanced recognized entity output by the large language model includes: Inputting the preset prompt word, candidate entity information, and context information of the initially recognized entity in the text to be processed into the large language model, so that the large language model performs enhanced verification on the initially recognized entity and outputs the enhanced recognized entity; The preset prompt words are used to guide the large language model to disambiguate any preliminary recognized entity based on context information when any preliminary recognized entity corresponds to multiple candidate entity information.

[0079] Specifically, the embodiment of the present invention guides and constrains the behavior of a large language model by presetting prompt words, thereby achieving efficient and reliable enhanced verification.

[0080] Preset prompts are pre-designed, structured instruction templates. They clearly explain to the large language model the role it needs to play, the specific tasks it must complete, the input structure, and the expected output format. By using preset prompts, a general-purpose large language model can be transformed into a behaviorally controllable expert system for performing domain-specific tasks, significantly improving the stability and accuracy of its output and effectively suppressing its tendency to produce irrelevant content or "hallucinations."

[0081] In this embodiment of the present invention, the preset prompt word is designed specifically for entity disambiguation. Its content clearly stipulates that when any pre-identified entity corresponds to multiple candidate entity information, the large language model is guided to disambiguate any pre-identified entity based on contextual information.

[0082] In the above embodiment, candidate entity information is generated through dual-channel retrieval using structured and vector queries, generating a ranked, high-quality list of candidate entities. For example, for "GDCA," the candidate entity information might be a list consisting of [{"Name":"Glycodeoxycholic Acid","Type":"Biologically Active Molecule"},{"Name":"GNAT3","Type":"Gene"}]. This provides the large language model with clear, limited options, narrowing its decision-making scope from an infinitely open space to a specific, highly relevant set.

[0083] In the above example, contextual information is the context of the initially recognized entity in the original text. For example, for the initially recognized entity "GDCA," its contextual information is "Study the impact of GDCA on the intestinal flora of diabetic patients..." This is the basis for the large language model to perform semantic judgment and reasoning.

[0084] Guided by preset prompts, the large language model analyzes the semantic association between contextual information ("diabetes," "gut flora") and candidate entities ("bile acid" and "gene"), ultimately making a judgment and outputting the enhanced recognized entity in the required format, for example: {"standard_name": "Glycodeoxycholic Acid", "type": "BileAcid"}. This output is then parsed by the system to determine the final named entity.

[0085] The entity recognition method for the field of nutritional and healthy food provided by the embodiment of the present invention organizes and guides the input and output of the large language model by introducing structured preset prompt words, thereby greatly improving the controllability, stability and accuracy of the large language model in performing enhanced verification tasks. It transforms an open generation task into a closed selection and judgment task with clear constraints, effectively avoiding problems such as inconsistent model output formats, divergent content or inconsistency with facts. This method enables the powerful reasoning ability of the large language model to be accurately applied to solve the problem of entity disambiguation, while ensuring that its output results can be seamlessly and automatically processed by subsequent programs.

[0086] In some embodiments, the preset prompt words are also used to guide the large language model to perform normalized entity representation on the enhanced recognition entities.

[0087] Specifically, the preset prompt words can also be used to guide the large language model to perform a new task of normalizing entity representation for enhanced recognition entities, thereby significantly improving the structure, standardization and usability of the final output results.

[0088] Normalized entity representation maps an entity, which may appear in various forms in text (such as abbreviations, colloquial names, and typos), to a unique, standardized, and information-rich structured object. The formats that can be defined for normalized entity representation include entity name, entity identifier, entity type, and synonym list.

[0089] The entity recognition method for the nutritional and health food sector, provided by embodiments of the present invention, not only addresses entity recognition accuracy issues but also significantly improves the standardization of recognition results by expanding the functionality of pre-set prompt words. The output, standardized entity representation, can be a structured object. This format can be seamlessly integrated and utilized by downstream applications (such as knowledge graph construction, database population, and data analysis systems), eliminating additional data cleaning and format conversion steps and significantly improving the automation level and end-to-end efficiency of the entire information extraction process.

[0090] Figure 2 This is the second flow chart of the entity recognition method for the field of nutritional and healthy food provided by the present invention. Figure 2 As shown, the method includes: Step 210: Build a hybrid knowledge representation and knowledge base architecture.

[0091] By combining the ontology library and the synonym library, a hybrid knowledge representation is obtained, which is stored in the structure database and entity vector library.

[0092] Step 220: Construct a domain-adaptive dual entity recognition mechanism.

[0093] The first stage of preliminary identification: A customized entity recognition tool for the field of nutrition and health food was built based on the en_core_sci_scibert pre-trained model of sciSpaCy. The tool includes a text normalization component (to handle special symbols and unify terminology), an entity recognition component (to identify various entity types such as food ingredients and nutritional components), an abbreviation detection and expansion component, and an EntityLinker component (to link to standard terminology libraries such as UMLS and MeSH to achieve entity normalization), which can quickly extract entities and standard entity identifiers (IDs).

[0094] The second stage of enhanced recognition: The "retrieval-verification-generation" three-layer enhanced recognition architecture is implemented through the large language model RAG technology to perform contextual disambiguation and missed detection completion on the preliminary recognition results.

[0095] The retrieval layer uses pgVector to implement hybrid retrieval of keyword matching and vector similarity; the verification layer verifies the sciSpaCy recognition results to solve the problem of polysemy; the generation layer uses large language models (such as DeepSeek) to generate standardized entity representations, perform synonym expansion and near-sense association.

[0096] By designing structured prompt words, the large model is guided to generate conclusions based on the retrieval results of the hybrid knowledge base, reducing the "hallucination" problem and improving the accuracy of entity recognition.

[0097] Step 230: Dynamically update the knowledge base.

[0098] The knowledge base can be updated using both incremental updates and periodic alignment. Incremental updates involve comparing new entities identified by the large model through semantic similarity. Entities with a confidence level exceeding a threshold are automatically included in the synonym database, and historical records are retained to ensure data consistency. Periodic alignment involves regularly monitoring updates to external data sources, automatically capturing and verifying newly added entities. A "last_updated" timestamp field is added to the structured knowledge base, automatically updated using database triggers, and entity addition, deletion, and modification operations, along with related information, are recorded in a change log.

[0099] The entity recognition method for the field of nutritional and healthy food provided by the embodiment of the present invention has the following technical effects: (1) tested on a dataset or document set in the field of nutritional and healthy food, the recognition ability of domain terms is improved; (2) the knowledge base can be updated in real time, and the cycle from recognition to storage of new terms is shortened, which is more efficient than the knowledge graph solution (which relies on manual review); (3) hybrid query reduces memory usage and resource consumption compared to the pure knowledge graph solution.

[0100] The following describes an apparatus provided by an embodiment of the present invention. The apparatus described below and the method described above can refer to each other.

[0101] Figure 3 This is a schematic diagram of the structure of the entity recognition device for the field of nutritional and healthy food provided by the present invention. Figure 3 As shown, the device includes: The text receiving module 310 is used to receive a text to be processed containing named entities in the field of nutritional and healthy food; A preliminary recognition module 320 is used to extract preliminary recognized entities from the text to be processed; Enhanced verification module 330 is used to input the preliminary recognized entity into the large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on candidate entity information and contextual information of the preliminary recognized entity in the text to be processed to generate an enhanced recognized entity; the candidate entity information is obtained by searching the knowledge base system based on the preliminary recognized entity; The entity output module 340 is used to determine named entities in the text to be processed based on the enhanced entity recognition.

[0102] The entity recognition device for the field of nutritional and healthy food provided by the embodiment of the present invention receives a text to be processed containing named entities in the field of nutritional and healthy food; extracts preliminary recognized entities from the text to be processed; inputs the preliminary recognized entities into a large language model to obtain enhanced recognized entities output by the large language model; determines the named entities in the text to be processed based on the enhanced recognized entities; due to the combination of preliminary recognition and enhanced recognition, the language processing tools in the relevant technology can be used to quickly extract the preliminary recognized entities, and the large language model retrieval enhancement generation technology is introduced to perform context-aware enhanced verification, which can effectively solve the problem of complex and highly ambiguous entity terms in the professional field of nutritional and healthy food. Compared with traditional entity recognition models, it greatly improves the accuracy and efficiency of entity recognition in the field of nutritional and healthy food. In addition, through preliminary recognition and retrieval in the knowledge base system, a clear candidate range and judgment basis are provided for the large language model, reducing the risk of it generating "hallucinations" and improving processing efficiency, ultimately achieving high-quality and structured information extraction from texts in the field of nutritional and healthy food.

[0103] Figure 4 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 4 As shown, the electronic device may include: a processor (Processor) 410, a communications interface (Communications Interface) 420, a memory (Memory) 430, and a communications bus (Communications Bus) 440. The processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call logic commands in the memory to execute the methods described in the above embodiments, for example: Receive a text to be processed containing named entities in the field of nutritional and healthy food; extract preliminary recognized entities from the text to be processed; input the preliminary recognized entities into a large language model to obtain enhanced recognized entities output by the large language model; the large language model performs enhanced verification on the preliminary recognized entities based on candidate entity information and contextual information of the preliminary recognized entities in the text to be processed to generate enhanced recognized entities; the candidate entity information is obtained after searching the knowledge base system based on the preliminary recognized entities; and the named entities in the text to be processed are determined based on the enhanced recognized entities.

[0104] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0105] The processor in the electronic device provided by the embodiment of the present invention can call the logic instructions in the memory to implement the above method. Its specific implementation method is consistent with the implementation method of the above method and can achieve the same beneficial effects, which will not be repeated here.

[0106] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.

[0107] Its specific implementation is consistent with the aforementioned method implementation and can achieve the same beneficial effects, so it will not be repeated here.

[0108] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0110] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for entity recognition in the field of nutritional and healthy food, characterized in that: include: receiving a text to be processed containing named entities in the field of nutrition, health and food; Extracting preliminary recognized entities from the text to be processed; Inputting the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on candidate entity information and context information of the preliminary recognized entity in the text to be processed to generate the enhanced recognized entity; the candidate entity information is obtained after searching the knowledge base system based on the preliminary recognized entity; Determine named entities in the to-be-processed text based on the enhanced recognition entity.

2. The entity recognition method for the field of nutritional and healthy food according to claim 1, characterized in that: The knowledge base system includes a structured database and an entity vector library; The structured database is used to perform structured storage on entities, synonyms of entities, and relationships between entities in the ontology library and the synonym library; The entity vector library is used to store the representation vectors corresponding to the entities in the ontology library and the representation vectors corresponding to the synonyms in the synonym library; The ontology library is constructed based on the food science ontology and the unified medical language system; The thesaurus is constructed based on supplement terms, ingredient terms, and product terms.

3. The entity recognition method for the field of nutritional and healthy food according to claim 2, characterized in that: The candidate entity information is determined based on the following steps: Based on the entity identifier, entity name, entity type, and synonym list of the initially identified entity, querying the structured database to obtain a first query result; Based on the contextual semantic features of the preliminary identified entity in the text to be processed, querying the entity vector library to obtain a second query result; The candidate entity information is generated based on the first query result and the second query result.

4. The entity recognition method for the field of nutritional and healthy food according to claim 2, characterized in that: The knowledge base system is updated based on the following steps: Obtaining the enhanced recognition entity output by the large language model; If the enhanced recognition entity does not exist in the knowledge base system and the semantic similarity between the enhanced recognition entity and any entity in the knowledge base system is greater than a preset threshold, the enhanced recognition entity is added to the knowledge base system as a synonym.

5. The entity recognition method for the field of nutritional and healthy food according to claim 2, characterized in that: The knowledge base system is updated based on the following steps: monitoring the ontology library and / or the synonym library; In the case where a new entity is added to the ontology library and / or the synonym library, the new entity is added to the knowledge base system.

6. The entity recognition method for the field of nutritional and healthy food according to claim 1, characterized in that: Inputting the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model includes: Inputting a preset prompt word, the candidate entity information, and context information of the initially recognized entity in the text to be processed into the large language model, so that the large language model performs enhanced verification on the initially recognized entity and outputs an enhanced recognized entity; The preset prompt word is used to guide the large language model to disambiguate any of the preliminary recognized entities based on the context information when any of the preliminary recognized entities corresponds to multiple candidate entity information.

7. The entity recognition method for the field of nutritional and healthy food according to claim 6, characterized in that: The preset prompt words are also used to guide the large language model to perform normalized entity representation on the enhanced recognition entity.

8. An entity recognition device for the field of nutritional and healthy food, characterized by: include: A text receiving module, used for receiving a text to be processed containing named entities in the field of nutritional and healthy food; A preliminary recognition module, used for extracting preliminary recognition entities in the text to be processed; An enhanced verification module is configured to input the preliminary recognized entity into a large language model to obtain an enhanced recognized entity output by the large language model; the large language model performs enhanced verification on the preliminary recognized entity based on candidate entity information and contextual information of the preliminary recognized entity in the text to be processed to generate the enhanced recognized entity; the candidate entity information is obtained after searching the knowledge base system for the preliminary recognized entity; An entity output module is used to determine the named entities in the text to be processed based on the enhanced recognition entity.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the entity recognition method for the field of nutritional and healthy food according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the entity recognition method for the field of nutritional and healthy food according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Public opinion analysis method, system, apparatus and storage medium

    CN109408804A

  • Medical text processing method and device, equipment and storage medium

    CN110442869A

  • Chinese named entity recognition and retrieval enhancement framework based on uncertain components

    CN116663559A

  • Unstructured text data knowledge extraction method based on large language model

    CN118036734A

  • Two-segment named entity recognition method, device and equipment and medium

    CN119416785A