Human anatomy intelligent question answering system and method based on local large model and knowledge base enhancement

By building an intelligent question-answering system for human anatomy based on local large models and enhanced knowledge base, the problems of insufficient data and data security of large models in the field of human anatomy are solved, efficient and accurate knowledge retrieval and answering are achieved, and data security and response speed are ensured.

CN120705370AInactive Publication Date: 2025-09-26CHONGQING SHAPINGBA DISTRICT ZHONGZHI MEDICAL VALLEY RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526642.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing large models in the field of human anatomy have problems with outdated data, insufficient corpus, and lack of expertise, which may lead to factual errors in the generated content. Medical data is difficult to deploy in the cloud and cannot meet the compliance requirements of local data storage and processing.

Method used

An intelligent question-answering system for human anatomy based on a local large model and knowledge base enhancement is adopted, including a data preprocessing module, a graph construction module, a knowledge base construction module and an intelligent question-answering module. The human anatomy knowledge graph is constructed through anomaly detection indicators, and the dialogue model is trained using a boundary-enhanced loss function to generate real-time answers.

Benefits of technology

It achieves the structured presentation of human anatomical knowledge, improves the accuracy and efficiency of knowledge retrieval, ensures the quality and accuracy of the knowledge graph, provides professional and accurate answers, protects the security and privacy of user data, and has a fast system response speed and supports real-time interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705370A_ABST
    Figure CN120705370A_ABST
Patent Text Reader

Abstract

The invention provides an anatomy intelligent question-answering system and method based on a local large model and knowledge base enhancement, and relates to the field of data processing, and the system comprises a data preprocessing module which is used for obtaining anatomy data, carrying out the preprocessing of the anatomy data, and generating local document data; the graph construction module is used for constructing a human anatomy knowledge graph based on the anomaly detection indexes and the local document data; the knowledge base construction module is used for constructing a knowledge base; the intelligent question and answer module is used for establishing a dialogue model and training the dialogue model based on a boundary enhanced loss function; and the intelligent question answering module is also used for acquiring real-time question information of the user, analyzing the real-time question information of the user through a dialogue model based on the human anatomy knowledge graph and the knowledge base, and generating a real-time answer, and has the advantages of improving the accuracy and speed of human anatomy intelligent question answering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a human anatomy intelligent question-answering system and method based on a local large model and knowledge base enhancement. Background Art

[0002] Large models have large-scale parameters and complex computational structures, trained using massive amounts of data and enormous computing power. Because they are pre-trained models, their raw data covers all aspects of various fields such as technology, life, and finance.

[0003] Although it has been widely used in scenarios such as open-domain question answering and text summarization, demonstrating strong semantic understanding, logical reasoning, and text generation capabilities, existing large models in specialized vertical fields such as human anatomy often suffer from a certain degree of data outdatedness and insufficient corpus. More commonly, due to a lack of anatomical expertise such as anatomical terminology and structural relationships, the generated content may contain factual errors, affecting users' understanding, learning, and even application of relevant field knowledge. At the same time, medical data may involve sensitive information and is difficult to deploy in the cloud, making it difficult for existing large models to meet the compliance requirements for local data storage and processing.

[0004] Therefore, it is necessary to provide a human anatomy intelligent question answering system and method based on a local large model and knowledge base enhancement to improve the accuracy and speed of human anatomy intelligent question answering. Summary of the Invention

[0005] The present invention provides a human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement, including: a data preprocessing module, used to obtain human anatomy data, and preprocess the human anatomy data to generate local document data; a graph construction module, used to construct a human anatomy knowledge graph based on anomaly detection indicators and local document data; a knowledge base construction module, used to construct a knowledge base; an intelligent question-answering module, used to establish a dialogue model, and train the dialogue model based on a boundary-enhanced loss function; the intelligent question-answering module is also used to obtain real-time question information from users, and analyze the real-time question information from users based on the human anatomy knowledge graph and knowledge base through the dialogue model to generate real-time answers.

[0006] Furthermore, the graph construction module constructs a human anatomy knowledge graph based on anomaly detection indicators and local document data, including: establishing and training a triple extraction model; extracting multiple triple relationships related to human anatomy terms from local document data through the triple extraction model; constructing a human anatomy knowledge graph based on the multiple triple relationships related to human anatomy terms; based on the anomaly detection indicators, judging whether the human anatomy knowledge graph meets the anomaly detection conditions, and if not, optimizing the human anatomy knowledge graph until the human anatomy knowledge graph meets the anomaly detection conditions.

[0007] Furthermore, the graph construction module determines whether the human anatomy knowledge graph meets the anomaly detection conditions based on the anomaly detection index, including: for each triple of the human anatomy knowledge graph, calculating the score of the triple in the anomaly detection index; based on the score of each triple of the human anatomy knowledge graph in the anomaly detection index, determining the score distribution characteristics of the anomaly detection index; based on the score distribution characteristics of the anomaly detection index, determining whether the human anatomy knowledge graph meets the anomaly detection conditions.

[0008] Furthermore, the graph construction module calculates the score of the triple in the anomaly detection index, including: calculating the topological consistency score of the head entity of the triple based on the degree deviation, clustering coefficient and anatomical system penalty coefficient of the head entity of the triple; calculating the topological consistency score of the tail entity of the triple based on the degree deviation, clustering coefficient and anatomical system penalty coefficient of the tail entity of the triple; calculating the relationship consistency score of the triple based on the distance between the anatomical systems to which the head entity and tail entity of the triple belong and the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph; calculating the score of the triple in the anomaly detection index based on the topological consistency score of the head entity of the triple, the topological consistency score of the tail entity of the triple and the relationship consistency score.

[0009] Furthermore, the graph construction module calculates the score of the triple in the anomaly detection index based on the topological consistency score of the head entity of the triple, the topological consistency score of the tail entity of the triple and the relationship consistency score, including: calculating the weight of the topological consistency score of the head entity of the triple based on the degree of the head entity of the relationship of the triple; calculating the weight of the topological consistency score of the tail entity of the triple based on the degree of the tail entity of the relationship of the triple; calculating the weight of the relationship consistency score of the triple based on the global frequency of the relationship of the triple; based on the weight of the topological consistency score of the head entity of the triple, the weight of the topological consistency score of the tail entity of the triple and the weight of the relationship consistency score of the triple, performing weighted summation of the topological consistency score of the head entity of the triple, the topological consistency score of the tail entity of the triple and the relationship consistency score of the triple, and calculating the score of the triple in the anomaly detection index.

[0010] Furthermore, the topological consistency score of the head entity of the triple is calculated based on the following formula:

[0011]

[0012] Among them, S topo (v) is the topological consistency score of the head entity of the triple, d(v) is the degree of the head entity of the triple, μ d is the mean degree of all nodes in the human anatomy knowledge graph, σ d is the standard deviation of the degree of all nodes in the human anatomy knowledge graph, C(v) is the clustering coefficient of the head entity of the triple, P sys (v) is the anatomical system penalty coefficient of the head entity of the triplet, α, β and γ are the anatomical system constraints, T(v) is the number of triangles between the neighbor nodes of the head entity of the triplet, k v is the degree of the head entity of the triple, sys(u) is the anatomical system to which the u-th neighbor node of the head entity of the triple belongs, sys(v) is the anatomical system to which the head entity of the triple belongs, |N(v)| is the total number of neighbor nodes of the head entity of the triple, and Π(·) is an indicator function, which is 1 if the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is different from the anatomical system to which the head entity of the triple belongs, and 0 otherwise.

[0013] Furthermore, the relation consistency score of the triples is calculated based on the following formula:

[0014]

[0015] Among them, S rel (h, r, t) is the relational consistency score of the triple with head entity h, relation r, and tail entity t, D sys (h, t) is the distance between the anatomical systems of the head entity and the tail entity of the triple. If the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is the same as that of the head entity of the triple, then D sys (h, t) is 0, λ1 is the weight, and L(h, t) is the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph.

[0016] Furthermore, the loss function based on boundary enhancement is:

[0017] L total =L CE +λ2×L boundary

[0018]

[0019] Among them, L totalis the total loss, L CE is the cross entropy loss, L boundary is the boundary perception loss, y i,k is the probability that the true i-th text block belongs to the k-th category, P i,k is the predicted probability that the i-th text block belongs to the k-th category, N is the sequence length, K is the number of label categories, and y i is the true label of the i-th text block, y i+1 is the true label of the i+1th text block, is an indicator function, which is 1 when the true labels of adjacent text blocks are different, otherwise it is 0, p i+1,k is the predicted probability that the i+1th text block belongs to the kth category, and λ2 is the weight.

[0020] Furthermore, the intelligent question-answering module analyzes the user's real-time question information based on the human anatomy knowledge map and knowledge base through a dialogue model and generates real-time answers, including: analyzing the user's real-time question information through a dialogue model to extract nouns related to human anatomy in the user's real-time question information; based on the nouns related to human anatomy in the user's real-time question information, querying target data from the human anatomy knowledge map and knowledge base to generate real-time answers.

[0021] The present invention provides a human anatomy intelligent question-answering method based on a local large model and knowledge base enhancement, which is applied to the above-mentioned human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement, including: obtaining human anatomy data, and preprocessing the human anatomy data to generate local document data; constructing a human anatomy knowledge graph based on anomaly detection indicators and local document data; constructing a knowledge base; establishing a dialogue model, and training the dialogue model based on a boundary-enhanced loss function; obtaining the user's real-time question information, and analyzing the user's real-time question information based on the human anatomy knowledge graph and knowledge base through the dialogue model to generate real-time answers.

[0022] Compared with the existing technology, the human anatomy intelligent question-answering system and method based on a local large model and knowledge base enhancement provided by the present invention has at least the following beneficial effects:

[0023] 1. A human anatomy knowledge graph is constructed using anomaly detection metrics and local document data. This presents complex human anatomy knowledge in a structured and contextualized manner, clearly demonstrating the connections between organs and their functional dependencies. This makes knowledge retrieval more accurate and efficient, allowing users to quickly access relevant anatomical information without having to search blindly through vast amounts of data. The knowledge base and knowledge graph work together to store detailed information and case data in specific fields, further enriching the knowledge system and providing users with more comprehensive knowledge support.

[0024] 2. During the knowledge graph construction process, judgment and optimization based on anomaly detection indicators can promptly detect and correct errors or incomplete information in the knowledge graph, ensuring the quality and accuracy of the knowledge graph. This is crucial for fields such as human anatomy, which require extremely high knowledge accuracy, and avoids misleading information caused by incorrect knowledge.

[0025] 3. The intelligent question-answering module analyzes and answers user questions based on a constructed human anatomy knowledge graph and knowledge base. It fully leverages the structured knowledge in the knowledge graph and the detailed information in the knowledge base to provide more accurate and professional answers. For example, when users ask questions about the functions of human organs or the impact of diseases, the system can provide comprehensive and accurate responses based on the associations in the knowledge graph and the case data in the knowledge base. The conversational model is trained using a boundary-enhanced loss function, enabling it to better handle edge cases and complex questions, improving the accuracy and robustness of question-answering. Even for uncommon or ambiguous questions, the system can provide reasonable responses. It can capture real-time user questions and quickly generate responses, enabling real-time interaction with users. This provides users with a convenient way to learn and query, allowing them to ask questions to the system anytime, anywhere and obtain the human anatomy knowledge they need.

[0026] 4. Based on a large local model and knowledge base, all data is processed and stored locally, eliminating the risk of data leakage and ensuring the security and privacy of user data. This is particularly important in areas involving sensitive information such as human anatomy. Local deployment enables faster system response, eliminating external network reliance, and providing users with smoother and more efficient services. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:

[0028] Figure 1 This is a module diagram of a human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement according to some embodiments of this specification;

[0029] Figure 2 This is a diagram of a development framework for a human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement according to some embodiments of this specification;

[0030] Figure 3 is a schematic diagram of a data sample after extraction according to some embodiments of this specification;

[0031] Figure 4is a structural diagram of a triple extraction model according to some embodiments of this specification;

[0032] Figure 5 is a flowchart of constructing a knowledge base according to some embodiments of this specification;

[0033] Figure 6 is a schematic diagram of a functional web page presentation process according to some embodiments of this specification;

[0034] Figure 7 It is a flowchart of the human anatomy intelligent question-answering method based on a local large model and knowledge base enhancement as shown in some embodiments of this specification. DETAILED DESCRIPTION

[0035] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0036] Figure 1 This is a module diagram of a human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement according to some embodiments of this specification, such as Figure 1 As shown, the human anatomy intelligent question-answering system based on the local large model and knowledge base enhancement can include a data preprocessing module, an atlas construction module, a knowledge base construction module and an intelligent question-answering module.

[0037] The data preprocessing module is used to obtain human anatomical data, preprocess the human anatomical data, and generate local document data.

[0038] Specifically, Figure 2 This is a development framework diagram of a human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement according to some embodiments of this specification, such as Figure 2 As shown, for existing raw data, especially those that cannot be directly uploaded to RAG for automatic preprocessing, including basic data files, books that greatly exceed the upload capacity, lecture notes with messy formats or contents, etc., these materials need to be extracted separately and subjected to data preprocessing operations according to different logics.

[0039] For example, for text data in PDF format, the data preprocessing module needs to convert it into editable TXT documents in batches through programming. Because the content formats are highly similar, regular expressions are used to break sentences and remove redundant text data that does not conform to the standard format. The specific expression is as follows:

[0040] ^\d{2}\.\d{4}\s+[\u4e00-\u9fa5]+\s+[a-zA-Z]+(?:\s+[a-zA-Z]+)+\s+[\u4e00-\u9fa5].*$

[0041] ^ and $ indicate the beginning and end of the expression, \d{2}\.\d{4}\s+ is used to match decimal formats such as 02.2132, [\u4e00-\u9fa5]+\s+[a-zA-Z]+(?:\s+[a-zA-Z]+)+\s+ is used to match the Chinese and English versions of anatomical terms, and [\u4e00-\u9fa5].* is used to match the description sentence part of anatomical terms.

[0042] Figure 3 is a schematic diagram of a data sample after extraction according to some embodiments of this specification, such as Figure 3 As shown, to facilitate subsequent data cleaning, uncommon symbols are inserted at breakpoints for segmentation. This also prepares for mitigating cross-row data loss and mispartitioning when uploading to a dedicated knowledge base for embedded vectorization. This also allows for the use of uncommon symbols to control the size of individual text blocks (tokens) in the database.

[0043] As for the existing human anatomy atlas, due to its mixed characteristics of pictures and text, we first program it to re-segment it by page number, and preliminarily label it by renaming the pictures with the corresponding page numbers. Then, by using optical character recognition (OCR) technology, we separate the text information contained in each page of the picture and establish a txt document annotation that corresponds to the picture one by one. Finally, the picture location, picture information and text information are stored in Json data format and synchronously imported into the local MySQL database for storage for subsequent use.

[0044] A graph construction module is used to build a human anatomy knowledge graph based on anomaly detection indicators and local document data.

[0045] Specifically include:

[0046] Build and train a triplet extraction model;

[0047] Extract multiple triple relationships related to human anatomy terms from local document data through a triple extraction model;

[0048] Construct a human anatomy knowledge graph based on the triple relationship of multiple human anatomy terms;

[0049] Based on the anomaly detection index, it is determined whether the human anatomy knowledge graph meets the anomaly detection conditions. If not, the human anatomy knowledge graph is optimized until the human anatomy knowledge graph meets the anomaly detection conditions.

[0050] Specifically, Figure 4 is a structural diagram of a triple extraction model according to some embodiments of this specification, such as Figure 4 As shown in the figure, after initially organizing local document data, specifically structuring the text data of the human anatomy term book TXT document, a sample of data was first sampled and labeled. This was done using the doccano tool platform, which significantly reduces data annotation time with its visual interface and convenient operation. A small number of triple relationships related to human anatomy terms were isolated, in the format of head entity-relationship-tail entity. This was then used to export a sample JSONL file, providing sample guidance for data conversion. The Universal Information Extraction (UIE) framework model was used to connect to the sample, and the model was evaluated using a validation set. The anatomical entity extraction model was adjusted several times to facilitate the isolation of triple relationships and the construction of a knowledge graph during subsequent conversations. After establishing this model, relational data was extracted from the large dataset and linked to a Neo4j database. Database creation operations were executed using Cyber ​​statements, enabling knowledge graph visualization and enriching the interface for easy access.

[0051] Since this knowledge graph is related to the field of human anatomy, a large part of the anomalies appear as unreasonable connections (such as the connection between the heart and bones) or topological structure abnormalities. To solve this problem, it is necessary to judge whether the human anatomy knowledge graph meets the anomaly detection conditions based on anomaly detection indicators.

[0052] In some embodiments, the graph construction module determines whether the human anatomy knowledge graph meets the anomaly detection conditions based on the anomaly detection index, including:

[0053] For each triple in the human anatomy knowledge graph, calculate the score of the triple in the anomaly detection index;

[0054] Based on the score of each triplet in the human anatomy knowledge graph in the anomaly detection index, the score distribution characteristics of the anomaly detection index are determined;

[0055] Based on the score distribution characteristics of the anomaly detection indicators, it is determined whether the human anatomy knowledge graph meets the anomaly detection conditions.

[0056] In some embodiments, the graph construction module calculates the score of the triplet in the anomaly detection index, including:

[0057] Calculate the topological consistency score of the head entity of the triple based on the degree deviation, clustering coefficient and anatomical system penalty coefficient of the head entity of the triple;

[0058] The topological consistency score of the triple's tail entity is calculated based on the degree deviation, clustering coefficient, and anatomical system penalty coefficient of the triple's tail entity;

[0059] The relationship consistency score of the triple is calculated based on the distance between the anatomical systems to which the head entity and the tail entity of the triple belong and the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph;

[0060] Based on the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score, the score of the triple in the anomaly detection index is calculated.

[0061] For example, the score of a triplet in anomaly detection metrics can be calculated based on the following formula:

[0062] K(h,r,t)=w h ×S topo (h)+w t ×S topo (t)+w r ×S rel (h,r,t)

[0063] Among them, K(h,r,t) is the score of the triplet in the anomaly detection index, S topo (h) is the topological consistency score of the head entity of the triple, S topo (t) is the topological consistency score of the tail entity of the triple, S rel (h, r, t) is the relationship consistency score of the triple, w h 、w t and w r is the weight.

[0064] In some embodiments, the topological consistency score of the head entity of the triple is calculated based on the following formula:

[0065]

[0066] Among them, S topo (v) is the topological consistency score of the head entity of the triple, that is, S topo (h), d(v) is the degree of the head entity of the triple, μ dis the mean degree of all nodes in the human anatomy knowledge graph, σ d is the standard deviation of the degree of all nodes in the human anatomy knowledge graph, C(v) is the clustering coefficient of the head entity of the triple, P sys (v) is the anatomical system penalty coefficient of the head entity of the triplet, α, β and γ are the anatomical system constraints, T(v) is the number of triangles between the neighbor nodes of the head entity of the triplet, k v is the degree of the head entity of the triple, sys(u) is the anatomical system to which the u-th neighbor node of the head entity of the triple belongs, sys(v) is the anatomical system to which the head entity of the triple belongs, |N(v)| is the total number of neighbor nodes of the head entity of the triple, and Π(·) is an indicator function, which is 1 if the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is different from the anatomical system to which the head entity of the triple belongs, and 0 otherwise. =1-C(v) is the complete normalized deviation. If the deviation is too large (e.g., the heart is connected to too many unrelated entities), it will appear abnormal. 1-C(v) means that if the clustering coefficient is low, the possibility of abnormality increases.

[0067] The calculation method of the topological consistency score of the tail entity of the triple is the same as the calculation method of the topological consistency score of the head entity of the triple, which will not be repeated here.

[0068] In some embodiments, the relation consistency score of a triple is calculated based on the following formula:

[0069]

[0070] Among them, S rel (h, r, t) is the relational consistency score of the triple with head entity h, relation r, and tail entity t, D sys (h, t) is the distance between the anatomical systems of the head entity and the tail entity of the triple. If the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is the same as that of the head entity of the triple, then D sys (h, t) is 0, which is used to punish unreasonable connections across systems, λ1 is the weight, and L(h, t) is the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph. It is a Sigmoid function. If L(h,t) is too large, it means that abnormality increases.

[0071] In some embodiments, the graph construction module calculates the score of the triple in the anomaly detection index based on the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score, including:

[0072] Calculate the weight of the topological consistency score of the head entity of the triple based on the degree of the head entity of the triple's relationship;

[0073] Based on the degree of the tail entity of the triple's relationship, the weight of the topological consistency score of the tail entity of the triple is calculated;

[0074] Calculate the weight of the triple's relationship consistency score based on the global frequency of the triple's relationship;

[0075] Based on the weight of the topological consistency score of the triple's head entity, the weight of the topological consistency score of the triple's tail entity, and the weight of the triple's relationship consistency score, the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score are weighted and summed to calculate the score of the triple in the anomaly detection index.

[0076] For example, the weight of the topological consistency score of the head entity of a triple can be calculated based on the following formula:

[0077]

[0078] Among them, w v is the weight of the topological consistency score of the triple’s head entity, d(v) is the degree of the triple’s head entity, d(u) is the degree of the triple’s head entity or the degree of the triple’s tail entity, d r is the global frequency of relation r.

[0079] The calculation method of the weight of the topological consistency score of the tail entity of the triple is the same as the calculation method of the weight of the topological consistency score of the head entity of the triple, which will not be repeated here.

[0080] The weight of the relation consistency score of a triple can be calculated based on the following formula:

[0081]

[0082] It is understandable that in the human anatomy knowledge graph, if the “connection” relationship appears in 50 triples (such as (heart, connection, aorta), (lung, connection, trachea), etc.), then d r =50. High-degree nodes and high-frequency relationships contribute more to the anomaly score. In other words, if certain relationships (such as "connection") appear frequently, it means that they may be core relationships. Therefore, when performing anomaly detection, we should pay more attention to the rationality of these relationships. Assume that the human anatomy knowledge graph contains 100 triple relationships, d r (connection) = 60 (more relationships, such as arteries connecting to the heart), d r(contains) = 60 (fewer relations, such as lung contains pulmonary artery). Taking the triple (heart, connection, stomach) as an example, d(heart) = 10, d(stomach) = 5, d r (connection) = 60, at this time w h =0.133, w t =0.067, w r =0.8, w r Dominant, anomaly detection focuses more on whether the "connection" is reasonable (for example, checking whether it is a cross-system connection).

[0083] Set the score threshold τ = μ K +k·σ K , if K(h,r,t)>τ, it is marked as abnormal. When there is no abnormal triple, and the score of each triple in the human anatomy knowledge graph in the abnormality detection index follows an approximate normal distribution When , it is determined that the abnormality detection condition is met.

[0084] When the human anatomy knowledge graph does not meet the anomaly detection conditions, the human anatomy knowledge graph can be instructed to be optimized based on the identified anomaly triples. For example, the human anatomy knowledge graph can be optimized manually based on the identified anomaly triples.

[0085] The knowledge base construction module is used to build the knowledge base.

[0086] Figure 5 is a flowchart of constructing a knowledge base according to some embodiments of this specification, such as Figure 5 As shown, the knowledge base utilizes a combination of Retrieval-Augmented Generation (RAG), a local MySQL database, and a Neo4j database. This cloud-based and local approach ensures data privacy and rapid scheduling. Both the local MySQL database and the Neo4j database have fixed setup procedures, requiring no special differentiated processing. Only preliminary data preprocessing is required. Therefore, establishing the knowledge base primarily focuses on building an accurate and usable RAG database.

[0087] In the RAG system, the original data is first uploaded to the system. These data can be text files, PDF documents, database entries, etc. Since the amount of data may be large, storage capacity and upload efficiency need to be considered. Here, it is controlled within 500MB. After uploading, it is usually raw and unstructured, so it needs to be preprocessed so that subsequent steps can efficiently use this data. If PDF or image data is uploaded, the system will mobilize the built-in OCR technology or other similar technologies to extract text. In order to improve the recognition rate and data accuracy, the number of uploads of this type of file is often reduced and the data is preprocessed locally into structured data. For other types of data, since the subsequent embedding model is usually limited by the input length token amount, the system will first divide the long document into smaller fragments for a segmentation operation, such as segmentation by paragraphs, sentences or a fixed number of characters, and at the same time remove irrelevant characters, repeated content, and low-value information (such as headers and footers).

[0088] The processed data needs to be converted into a machine-understandable representation, that is, it needs to be understood and used by the large model. Here, the bge-large-zh-v1.5 model that has been embedded and installed in the RAG system is selected to convert the processed text into a high-dimensional vector representation, which can capture the semantic information of the text, thereby realizing efficient text retrieval and semantic similarity calculation.

[0089] Here, select the conversion mode as general, set the block token count to no more than 600, and select segmentation identifiers from the uploaded file to be processed. For ease of segmentation, consider using characters such as periods, commas, and colons. Each data fragment (chunk) is vectorized to generate a corresponding vector representation. This vector is then associated with the original text fragment and stored in a vector database. For a high-dimensional vector database, choose FAISS, which supports fast similarity search. All vectors are indexed to accelerate retrieval. Metadata (such as original text, source file, and location) is also stored to facilitate tracing back to specific content.

[0090] It can be understood that the bidirectional mapping between the Neo4j graph database and the MySQL relational database supports multi-scale visualization presentation by designing a unified data model based on anatomical ontology.

[0091] The intelligent question-answering module is used to build a dialogue model and train the dialogue model based on a boundary-enhanced loss function.

[0092] Specifically, the dialogue model is implemented based on a large model, deployed locally through Ollama and called through RAGFlow. When accessed through the Web application interface, it is switched through instructions or buttons to achieve the combination of pre-training and specialized data.

[0093] During intelligent question-and-answer conversations, the intelligent Q&A page can also display images of human anatomical entities that may be contained in the question, a knowledge topology of the entity, and data sources. This is achieved by integrating entity recognition and database matching between the user's question and the big model. When the question is transmitted to the big model, the dialogue system simultaneously passes the question into the pre-set entity recognition model, extracting any nouns related to human anatomy. This noun then queries the database for any corresponding images, maps, and records, and finally displays them on the page for the user to view.

[0094] The entity recognition task is driven by Transformer-based pre-trained language models (such as BERT and RoBERTa). To better adapt to specialized terms in the field of human anatomy (such as "cervical triangle" and "groin"), the BERT pre-trained model is used as a basis and its general representation capabilities are adjusted to specific task requirements through fine-tuning, namely, extracting human anatomy-related entities. This process is modeled as a sequence labeling task, where each word is labeled as belonging to a specific category (such as "anatomical term") or not. To improve the model's performance in this task, an innovative optimization strategy is introduced during fine-tuning, and the dialogue model is trained based on a boundary-enhanced loss function. The fine-tuning process mainly includes the following steps: 1. Convert the input sentence into BERT's input format (word segmentation and embedding representation); 2. Add a task-specific classification layer on the BERT output layer; 3. Define a loss function that incorporates boundary awareness; 4. Adjust the model parameters through gradient descent, including BERT's pre-trained parameters and the newly added layers.

[0095] After the input sentence is processed by BERT's word segmenter, it is converted into a token sequence and an embedded representation is generated. BERT outputs a context-dependent representation matrix whose dimensions are the sequence length N and the hidden layer dimension (default 768). For each token, it is represented as h i On top of the BERT output, add a fully connected layer to h i Mapped to the label space, the score is calculated as follows:

[0096] Z i =h i W+b

[0097] in, is a weight matrix with a dimension of [768, K] (K is the number of label categories), is the bias vector, is the score vector for the i-th token, representing its raw score for each label. Finally, the scores are converted into a probability distribution using the Softmax activation function. To simplify computation and adapt to the task characteristics, the BIO (Begin, Inside, Outside) annotation method is used, but the number of label categories is initially set to two ("anatomical terms" and "non-anatomical terms").

[0098] In some embodiments, the loss function based on boundary enhancement is:

[0099] L total =L CE +λ2×L boundary

[0100]

[0101] Among them, L total is the total loss, L CE is the cross entropy loss, L boundary is the boundary perception loss, y i,k is the probability that the true i-th text block belongs to the k-th category, P i,k is the predicted probability that the i-th text block belongs to the k-th category, N is the sequence length, K is the number of label categories, and y i is the true label of the i-th text block, y i+1 is the true label of the i+1th text block, is an indicator function, which is 1 when the true labels of adjacent text blocks are different, otherwise it is 0, p i+1,k is the predicted probability that the i+1th text block belongs to the kth category, and λ2 is the weight.

[0102] The calculation of the gradient is mainly the gradient calculation of the weight W of the fully connected layer:

[0103]

[0104] Transpose of the input The difference between the prediction and the actual value is the gradient of the bias b:

[0105]

[0106] Finally, the optimizer is used to update the parameters, and η is the learning rate to adapt to the acceleration task:

[0107]

[0108] During fine-tuning, a shallow fine-tuning strategy was selected to freeze the pre-trained parameters of BERT and only train the W and b of the fully connected layer. To prevent overfitting, an L2 regularization term was added with a regularization coefficient of α. The optimizer used was Adam, and the learning rate was set to 3×10-5 , to balance the convergence speed and stability. The loss function calculates the gradient through back propagation and updates the parameters in combination with the improved gradient formula.

[0109] Figure 6 is a schematic diagram of a functional web page presentation process according to some embodiments of this specification, such as Figure 6 As shown, the webpage presentation utilizes the Flask backend framework, combined with a frontend framework comprised of HTML, CSS, and JavaScript. The backend uses Flask routing to establish a comprehensive application on client 5000, listening on local ports 11434 and 80, respectively. This connects to Ollama to deploy a large local model and enhance the RAGFlow knowledge base, forwarding traffic to clients 5001 and 5002 for use. Meanwhile, the frontend consists of four pages: Home, Q&A, Graph, and Edit. The Home page primarily explains the overall application's usage and precautions. The Q&A page allows for different Q&A modes, displaying images and graphs related to the Q&A content. The Graph page visualizes the entire knowledge graph, facilitating the search and modification of specific content, helping to grasp the overall knowledge base. The Edit page provides an integrated interface for managing the entire knowledge base, allowing for modifications, deletions, or expansions of underlying knowledge, enhancing the overall system's intelligence.

[0110] The intelligent question-answering module is also used to obtain users' real-time question information, and analyze the users' real-time question information based on the human anatomy knowledge map and knowledge base through a dialogue model to generate real-time answers.

[0111] Specifically include:

[0112] Analyze users' real-time question information through the dialogue model and extract nouns related to human anatomy from the users' real-time question information;

[0113] Based on the nouns related to human anatomy in the user's real-time question information, the target data (such as pictures, maps and records) are queried from the human anatomy knowledge graph and knowledge base to generate real-time answers.

[0114] As expected, Ollama deeply couples a large, locally deployed model with RAGFlow's retrieval-enhanced generation technology to build a medical knowledge question-answering system with offline reasoning capabilities. This secure interaction between the local model and the cloud-based knowledge base ensures the privacy of anatomical data while pushing the boundaries of single-model knowledge. For question-answering, a fine-tuning of the entity noun recognition model was implemented, achieving dual-track optimization for both overall question-answering speed and the speed of displaying relevant content.

[0115] Figure 7This is a flowchart of the human anatomy intelligent question answering method based on the local large model and knowledge base enhancement shown in some embodiments of this specification, such as Figure 7 As shown, the human anatomy intelligent question answering method based on local large model and knowledge base enhancement can include the following steps.

[0116] Acquire human anatomical data, pre-process the human anatomical data, and generate local document data;

[0117] Build a human anatomy knowledge graph based on anomaly detection metrics and local document data;

[0118] Build a knowledge base;

[0119] Build a dialogue model and train it based on a boundary-enhanced loss function.

[0120] Obtain the user's real-time question information, and analyze the user's real-time question information based on the human anatomy knowledge graph and knowledge base through the dialogue model to generate real-time answers.

[0121] The human anatomy intelligent question-answering method based on local large models and knowledge base enhancement can be applied to the human anatomy intelligent question-answering system based on local large models and knowledge base enhancement, which will not be repeated here.

[0122] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.

Claims

1. An intelligent question-answering system for human anatomy based on a local large model and knowledge base enhancement, characterized by: include: A data preprocessing module is used to obtain human anatomical data, preprocess the human anatomical data, and generate local document data; A graph construction module, which is used to construct a human anatomy knowledge graph based on anomaly detection indicators and local document data; Knowledge base construction module, used to construct the knowledge base; Intelligent question-answering module, used to build a dialogue model and train the dialogue model based on a boundary-enhanced loss function; The intelligent question-answering module is also used to obtain real-time question information from users, and analyze the real-time question information from users based on the human anatomy knowledge graph and knowledge base through a dialogue model to generate real-time answers.

2. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 1 is characterized in that: The graph construction module constructs a human anatomy knowledge graph based on anomaly detection indicators and local document data, including: Build and train a triplet extraction model; Extract multiple triple relationships related to human anatomy terms from local document data through a triple extraction model; Construct a human anatomy knowledge graph based on the triple relationship of multiple human anatomy terms; Based on the anomaly detection index, it is determined whether the human anatomy knowledge graph meets the anomaly detection conditions. If not, the human anatomy knowledge graph is optimized until the human anatomy knowledge graph meets the anomaly detection conditions.

3. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 2 is characterized in that: The graph construction module determines whether the human anatomy knowledge graph meets the anomaly detection conditions based on the anomaly detection index, including: For each triple in the human anatomy knowledge graph, calculate the score of the triple in the anomaly detection index; Based on the score of each triplet in the human anatomy knowledge graph in the anomaly detection index, the score distribution characteristics of the anomaly detection index are determined; Based on the score distribution characteristics of the anomaly detection indicators, it is determined whether the human anatomy knowledge graph meets the anomaly detection conditions.

4. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 3 is characterized in that: The graph construction module calculates the score of the triple in the anomaly detection index, including: Calculate the topological consistency score of the head entity of the triple based on the degree deviation, clustering coefficient and anatomical system penalty coefficient of the head entity of the triple; The topological consistency score of the triple's tail entity is calculated based on the degree deviation, clustering coefficient, and anatomical system penalty coefficient of the triple's tail entity; The relationship consistency score of the triple is calculated based on the distance between the anatomical systems to which the head entity and the tail entity of the triple belong and the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph; Based on the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score, the score of the triple in the anomaly detection index is calculated.

5. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 4 is characterized in that: The graph construction module calculates the score of the triple in the anomaly detection index based on the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score, including: Calculate the weight of the topological consistency score of the head entity of the triple based on the degree of the head entity of the triple's relationship; Based on the degree of the tail entity of the triple's relationship, the weight of the topological consistency score of the tail entity of the triple is calculated; Calculate the weight of the triple's relationship consistency score based on the global frequency of the triple's relationship; Based on the weight of the topological consistency score of the triple's head entity, the weight of the topological consistency score of the triple's tail entity, and the weight of the triple's relationship consistency score, the topological consistency score of the triple's head entity, the topological consistency score of the triple's tail entity, and the relationship consistency score are weighted and summed to calculate the score of the triple in the anomaly detection index.

6. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 4 is characterized in that: The topological consistency score of the head entity of the triple is calculated based on the following formula: Among them, S topo (v) is the topological consistency score of the head entity of the triple, d(v) is the degree of the head entity of the triple, μ d is the mean degree of all nodes in the human anatomy knowledge graph, σ d is the standard deviation of the degree of all nodes in the human anatomy knowledge graph, C(v) is the clustering coefficient of the head entity of the triple, P sys (v) is the anatomical system penalty coefficient of the head entity of the triplet, α, β and γ are the anatomical system constraints, T(v) is the number of triangles between the neighbor nodes of the head entity of the triplet, k v is the degree of the head entity of the triple, sys(u) is the anatomical system to which the u-th neighbor node of the head entity of the triple belongs, sys(v) is the anatomical system to which the head entity of the triple belongs, |N(v)| is the total number of neighbor nodes of the head entity of the triple, and Π(·) is an indicator function, which is 1 if the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is different from the anatomical system to which the head entity of the triple belongs, and 0 otherwise.

7. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 6 is characterized in that: The relation consistency score of a triple is calculated based on the following formula: Among them, S rel (h, r, t) is the relational consistency score of the triple with head entity h, relation r, and tail entity t, D sys (h, t) is the distance between the anatomical systems of the head entity and the tail entity of the triple. If the anatomical system to which the u-th neighbor node of the head entity of the triple belongs is the same as that of the head entity of the triple, then D sys (h, t) is 0, λ1 is the weight, and L(h, t) is the shortest path length from the head entity to the tail entity of the triple in the human anatomy knowledge graph.

8. The human anatomy intelligent question-answering system based on local large model and knowledge base enhancement according to claim 7 is characterized in that: The loss function based on boundary enhancement is: L total =L CE +λ2×L boundary Among them, L total is the total loss, L CE is the cross entropy loss, L boundary is the boundary perception loss, y i,k is the probability that the true i-th text block belongs to the k-th category, P i,k is the predicted probability that the i-th text block belongs to the k-th category, N is the sequence length, K is the number of label categories, and y i is the true label of the i-th text block, y i+1 is the true label of the i+1th text block, is an indicator function, which is 1 when the true labels of adjacent text blocks are different, otherwise it is 0, p i+1,k is the predicted probability that the i+1th text block belongs to the kth category, and λ2 is the weight.

9. The human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement according to any one of claims 1 to 8, characterized in that: The intelligent question-answering module analyzes the user's real-time question information based on the human anatomy knowledge graph and knowledge base through a dialogue model and generates real-time answers, including: Analyze users' real-time question information through the dialogue model and extract nouns related to human anatomy from the users' real-time question information; Based on the nouns related to human anatomy in the user's real-time question information, the target data is queried from the human anatomy knowledge graph and knowledge base to generate real-time answers.

10. An intelligent question-answering method for human anatomy based on a local large model and knowledge base enhancement, characterized by: The human anatomy intelligent question-answering system based on a local large model and knowledge base enhancement as described in any one of claims 1 to 9 comprises: Acquire human anatomical data, pre-process the human anatomical data, and generate local document data; Build a human anatomy knowledge graph based on anomaly detection metrics and local document data; Build a knowledge base; Build a dialogue model and train it based on a boundary-enhanced loss function. Obtain the user's real-time question information, and analyze the user's real-time question information based on the human anatomy knowledge graph and knowledge base through the dialogue model to generate real-time answers.