Method for establishing acupuncture point knowledge base by using traditional Chinese medicine ancient books
By analyzing the text sequences of ancient Chinese books in traditional Chinese medicine, using Jieba and word2vec technology to extract acupuncture points vocabulary, an accurate and easy-to-retrieve acupuncture points knowledge map was constructed, which solved the problem of difficulty in retrieving information of ancient Chinese books in traditional Chinese medicine, and improved the accuracy and readability of the knowledge map.
Patent Information
- Application Number
- CN202510469657.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-15
AI Technical Summary
The records of acupuncture acupoints in ancient Chinese books appear in the form of classical Chinese, with complex content and lack of systematic description, which leads to difficulty in retrieving information. The knowledge graph constructed by existing technology is relatively accurate and difficult to retrieve.
By obtaining the sequence of ancient Chinese medicine texts, using Jieba word segmentation tool to extract acupoint vocabulary, combining the word2vec algorithm to analyze the relative position and similar relationship between split vocabulary and acupoint vocabulary, calculate the complexity of vocabulary distribution and the degree of variable description, and construct acupoint knowledge graph based on search attention.
It improves the accuracy and readability of the knowledge graph, simplifies the information retrieval of acupoint vocabulary, and enhances the reliability of knowledge inheritance.
Smart Images

Figure CN120496751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical knowledge graphs, and in particular to a method for establishing an acupuncture point knowledge base using ancient Chinese medical books. Background Art
[0002] Traditional Chinese Medicine (TCM), a treasure of traditional Chinese medicine, boasts thousands of years of history and profound cultural heritage. Acupuncture, a crucial component of TCM, regulates bodily functions by stimulating specific acupuncture points on the human body, achieving the goals of treating illness and maintaining health. However, records of acupuncture points in ancient TCM texts are often written in classical Chinese, with complex content and diverse expressions, posing significant challenges to modern TCM practitioners and learners. While these ancient texts contain a wealth of knowledge about acupuncture points, the lack of systematic organization and standardized descriptions makes information retrieval difficult and limits knowledge transfer. Consequently, analysis of these ancient TCM texts is often conducted to construct knowledge graphs.
[0003] When constructing acupuncture point knowledge graphs, existing technologies usually segment the text sequences in ancient books and then extract the entity information to construct the knowledge graph. However, since this method is based on the global properties of vocabulary for analysis, and in ancient books, there may be polysemy or changes in the sentence structure of polysemous words, the importance of different words will vary. Therefore, directly extracting entity information will result in the final constructed knowledge graph with low accuracy and great difficulty in retrieval. Summary of the Invention
[0004] In order to solve the problem that in ancient books, there may be polysemy of a word or changes in the sentence structure of the polysemy word, the importance of different words may vary. Therefore, directly extracting entity information will lead to low accuracy of the final constructed knowledge graph and great retrieval difficulty. The purpose of this invention is to provide a method for establishing an acupuncture point knowledge base using ancient Chinese medicine books. The technical solution adopted is as follows:
[0005] Obtain text sequences from ancient Chinese medical books;
[0006] Performing vocabulary segmentation on the text sequence to obtain all split words, and extracting acupoint words from all the split words, wherein the same acupoint words belong to the same type of acupoint words; determining an operation statement corresponding to each acupoint word based on the position of the acupoint word in the text sequence;
[0007] In all the operation statements corresponding to each acupoint vocabulary, the similarity between the relative positions of the split words and the acupoint vocabulary, as well as the similarity between the split words, are analyzed to determine the lexical distribution complexity corresponding to the operation statements of each acupoint vocabulary; the quantitative distribution characteristics of the split words in all the operation statements of each acupoint vocabulary and the lexical distribution complexity corresponding to the acupoint vocabulary are combined to obtain the descriptive variability value of each acupoint vocabulary;
[0008] The differences between the description variability values corresponding to all acupoint words are compared to determine the retrieval attention of each acupoint word; the acupoint words are marked in the text sequence based on the retrieval attention of all acupoint words to construct an acupoint knowledge graph.
[0009] Furthermore, the method for obtaining the character sequence includes:
[0010] Obtain grayscale text images from ancient Chinese medical books, and perform threshold segmentation and connected domain detection on the grayscale text images, thereby obtaining the minimum circumscribed rectangle of each connected domain as a character selection box for each character;
[0011] The character selection box is recognized based on the OCR tool to obtain the corresponding characters. After the characters are unified and special characters are removed, the text sequence is obtained based on the preset reading order.
[0012] Furthermore, the word segmentation of the text sequence to obtain all split words and extracting acupoint words from all the split words includes:
[0013] Performing word segmentation on the text sequence based on the Jieba word segmentation tool to obtain all the split words;
[0014] All split words are matched based on a preset acupoint vocabulary library, so as to extract acupoint words from all split words.
[0015] Furthermore, the method for obtaining the operation statement includes:
[0016] In the text sequence, the split words between each two adjacent acupoint words are combined into a sentence sequence;
[0017] In terms of word order, the sentence sequence adjacent to each acupoint vocabulary is combined with the acupoint vocabulary to form the operation sentence corresponding to each acupoint vocabulary.
[0018] Furthermore, the method for obtaining the vocabulary distribution complexity includes:
[0019] Obtain the vocabulary vector corresponding to each split word and acupoint word in the text sequence based on the word2vec algorithm;
[0020] Preset an attention window, and in the attention window corresponding to each acupoint vocabulary, calculate the cosine similarity between the vocabulary vector of each split vocabulary and the vocabulary vector of each acupoint vocabulary as the vocabulary similarity value;
[0021] In each acupoint vocabulary, one acupoint vocabulary is selected as the target vocabulary, and the remaining vocabulary is selected as the contrast vocabulary. The target vocabulary is combined with each contrast vocabulary in pairs, so as to obtain all vocabulary combinations corresponding to the target vocabulary.
[0022] In each vocabulary combination, the difference between the lexical similarity values of the split words in the attention window of the target word and the comparison word, as well as the similarity relationship between the relative positions of the split words and the acupoint words were analyzed to determine the lexical distribution complexity factor of the target word in each vocabulary combination;
[0023] The normalized value of the mean of the vocabulary distribution complexity factors of the target vocabulary in all vocabulary combinations is used as the vocabulary distribution complexity corresponding to the operation statement of the target vocabulary.
[0024] Furthermore, the method for obtaining the vocabulary distribution complexity factor includes:
[0025] In the attention window of the target vocabulary, any split vocabulary is selected as the vocabulary to be analyzed. In each vocabulary combination of the target vocabulary, the absolute value of the difference between the vocabulary similarity value of each split vocabulary in the attention window of the target vocabulary and the comparison vocabulary is calculated as the deviation factor.
[0026] The split word with the smallest deviation factor is used as the adaptation word of the word to be analyzed, and the character distance between the word to be analyzed and the target word, as well as the character distance between the adaptation word and the comparison word, is obtained. The absolute value of the difference between the two character distances is used as the distance factor corresponding to the word to be analyzed;
[0027] Analyze the relative position between the word to be analyzed and the target word, and the relative position between the adaptation word and the comparison word, and determine the position similarity factor corresponding to the word to be analyzed;
[0028] Determining a vocabulary distribution difference factor of the vocabulary to be analyzed based on a distance factor and a position similarity factor corresponding to the vocabulary to be analyzed, wherein the distance factor is positively correlated with the vocabulary distribution difference factor, and the position similarity factor is negatively correlated with the vocabulary distribution difference factor;
[0029] The mean of all lexical distribution difference factors of the target vocabulary in each lexical combination is taken as the lexical distribution complexity factor of the target vocabulary in each lexical combination.
[0030] Furthermore, the method for obtaining the position similarity factor includes:
[0031] The word order of the character sequence is the positive direction, and the step length is one character;
[0032] In the attention window corresponding to the target word, the target word is taken as the starting point and the word to be analyzed is taken as the end point, thereby obtaining the position vector corresponding to the word to be analyzed;
[0033] In the attention window corresponding to the contrasting vocabulary, the contrasting vocabulary is used as the starting point and the matching vocabulary of the vocabulary to be analyzed is used as the end point, thereby obtaining the position vector corresponding to the matching vocabulary;
[0034] The cosine similarity between the two position vectors is calculated and the numerical adjustment is performed to obtain the position similarity factor corresponding to the word to be analyzed.
[0035] Furthermore, the method for obtaining the description variability value includes:
[0036] In each acupoint vocabulary, the absolute value of the difference in the number of split words in the two sentence sequences before and after each acupoint vocabulary is counted as the vocabulary number difference value;
[0037] The difference between the maximum and minimum values of the vocabulary quantity difference value is taken as the vocabulary quantity difference range;
[0038] The ratio of the vocabulary quantity difference value of each acupoint vocabulary to the extreme difference of the vocabulary quantity difference is used as the first variable factor;
[0039] The ratio of the vocabulary distribution complexity corresponding to each acupoint vocabulary to the maximum value of the distribution complexity of all acupoint vocabulary is used as the second variable factor of each acupoint vocabulary;
[0040] The value obtained by normalizing the product of the first variable factor and the second variable factor of each acupoint vocabulary is used as the description variable degree value of each acupoint vocabulary.
[0041] Furthermore, the method for obtaining the retrieval attention includes:
[0042] In each acupoint vocabulary, the mean of the description variability values of all acupoint vocabulary is taken as the mean feature value;
[0043] The mean of the mean characteristic values of all acupoint words was taken as the baseline value;
[0044] The difference between the mean characteristic value of each acupoint word and the reference value is normalized and used as the retrieval attention corresponding to each acupoint word.
[0045] Furthermore, the retrieval attention based on all acupoint words is used to mark the acupoint words in the text sequence to construct an acupoint knowledge graph, including:
[0046] In the text sequence, all acupoint words are marked based on their retrieval attention, to obtain a marked text sequence;
[0047] The marked text sequence is used as the input of the SDP tool to obtain the SDP node graph;
[0048] Inputting the SDP node graph into a pre-trained neural network to obtain the weights of each statement path in the SDP node graph;
[0049] All sentence paths containing acupoint vocabulary in the SDP node graph are extracted to obtain the acupoint knowledge graph.
[0050] The present invention has the following beneficial effects:
[0051] Obtaining a sequence of characters in ancient Chinese medical books is a step used to obtain ancient book data and provide a data source for subsequent processing. By word segmentation, split words are obtained, and complex character sequences can be converted into easy-to-process word units, thereby facilitating the extraction of acupoint words from all the split words. Since in ancient Chinese medical books, the description sentences or operation sentences for different acupoint words are different, and even the description sentences or operation sentences for the same acupoint words in different semantics or situations may also be different, the present invention determines the operation sentence corresponding to each acupoint word based on the position of the acupoint word in the character sequence, providing key information for the subsequent construction of an accurate knowledge graph. Furthermore, in all the operation sentences corresponding to each acupoint word, the similarity relationship between the relative positions of the split words and the acupoint words, as well as the similarity between the split words, is analyzed to obtain the vocabulary distribution complexity corresponding to each acupoint word; this can provide an in-depth understanding of the context of the acupoint words in ancient books, and explain the distribution patterns of the words in ancient books, thereby providing important clues for accurately interpreting acupoint knowledge and enhancing the accuracy and readability of the subsequent knowledge graph construction. Furthermore, by comprehensively considering the quantitative distribution characteristics of the split words in the operational statements of each acupoint vocabulary and the complexity of the lexical distribution corresponding to the acupoint vocabulary, we can obtain the descriptive variability value of the acupoint vocabulary. The descriptive variability value can be used to evaluate the tolerance of the information representation of acupoint vocabulary in ancient books, that is, the diversity of acupoint vocabulary descriptions. Finally, by comparing the differences between the descriptive variability values of acupoint vocabulary, we can determine the retrieval attention of each acupoint vocabulary. The retrieval attention can be used to describe the importance and priority of each acupoint vocabulary in the construction of the knowledge graph. Therefore, based on the retrieval attention of all acupoint vocabulary, acupoint vocabulary can be marked in the text sequence, thereby constructing a more accurate and easy-to-search acupoint knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 A flowchart of a method for establishing an acupuncture point knowledge base using ancient Chinese medical books provided by one embodiment of the present invention;
[0054] Figure 2 A flow chart of a method for obtaining vocabulary distribution complexity provided by one embodiment of the present invention;
[0055] Figure 3 A schematic diagram of an SDP node graph provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail a method for establishing an acupuncture point knowledge base using ancient Chinese medical texts, including its specific implementation, structure, features, and effectiveness. In the following description, references to different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0057] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0058] The following describes in detail a method for establishing an acupuncture point knowledge base using ancient Chinese medical books provided by the present invention with reference to the accompanying drawings.
[0059] See also Figure 1 , which shows a flow chart of a method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to an embodiment of the present invention, the method comprising the following steps:
[0060] Step S1: Obtain a text sequence from ancient Chinese medical books.
[0061] Ancient Chinese medical books record a great deal of medical skills and their operating procedures. Therefore, analyzing acupuncture technical knowledge from ancient books can extract knowledge information from the ancient books, thereby migrating the ancient book information from book text to computer data, which is conducive to the analysis of digital knowledge in ancient books. Therefore, we can first obtain the text sequence in ancient Chinese medical books.
[0062] Preferably, in one embodiment of the present invention, the method for obtaining a character sequence includes:
[0063] When using a high-definition camera to capture text images in ancient Chinese medical books, it is necessary to ensure that the images are clear and have moderate contrast to reduce the difficulty of subsequent processing. The text images are grayscaled to obtain grayscale text images.
[0064] Since the text in ancient books is usually in black font and the background is white or yellow, threshold segmentation and connected domain detection are performed on the grayscale text image to identify all connected black pixel areas. At this time, each connected domain can represent a character, so the minimum bounding rectangle of each connected domain is obtained as the character selection box for each character. It should be noted that when obtaining the character selection box, some minimum bounding rectangles can also be fine-tuned. For example, when the area of a minimum bounding rectangle is less than a preset area threshold, the corresponding connected domain is merged with the connected domain corresponding to the minimum bounding rectangle whose center is closest to it, thereby obtaining a character selection box; the area threshold is set to the average area of the minimum bounding rectangles of all connected domains.
[0065] Finally, the character selection box is recognized based on the OCR tool to obtain the corresponding characters. Since there may be some variants or simplified forms of words in ancient books, all characters are unified through the preset font comparison dictionary. The variants or simplified forms in the classical Chinese text are modified to reduce the interference caused by different font forms. After unification, the punctuation marks in the sentence are removed through the preset symbol dictionary, and the text sequence is obtained based on the preset reading order.
[0066] It should be noted that the grayscale image acquisition method can adopt the average grayscale method, the maximum value method or the weighted grayscale method, and all of them are technical means well known to those skilled in the art, and are not limited or elaborated here; the threshold segmentation method can adopt the Otsu threshold segmentation method, which and the connected domain detection, the minimum circumscribed rectangle acquisition process and the use of OCR tools are all well-known technologies and are not elaborated here; the preset reading order is the writing direction of the characters in the ancient books, such as from top to bottom, from right to left, or from top to bottom, from left to right, etc.; the preset font reference dictionary and symbol dictionary can be obtained by using programming languages such as Python in combination with relevant font processing libraries (Pillow, fontforge) to read the font file and extract character information, and then generate a dictionary containing character encoding and symbol encoding. It can also be obtained according to other methods, and the specific acquisition process is not limited or elaborated here.
[0067] Step S2: Segment the text sequence into words to obtain all split words, and extract acupoint words from all the split words, wherein the same acupoint words belong to the same type of acupoint words; determine the operation statement corresponding to each acupoint word based on the position of the acupoint word in the text sequence.
[0068] The purpose of the embodiment of the present invention is to establish an acupuncture point knowledge base. Therefore, in the process of constructing the knowledge graph, more attention should be paid to the acupuncture point vocabulary. Therefore, the text sequence can be segmented to obtain all the split words, thereby extracting the acupuncture point vocabulary. Furthermore, since classical Chinese is usually used in ancient books, a single word is often used to express an action or a process. At the same time, there are situations such as polysemy in classical Chinese. Therefore, some words in classical Chinese need to be combined to accurately express their meaning in a specific scenario. Therefore, after obtaining the acupuncture point vocabulary, it is necessary to determine the description or operation that needs to be performed on it in the ancient book. Therefore, based on the position of the acupuncture point vocabulary in the text sequence, the operation statement corresponding to each acupuncture point vocabulary is determined. At this time, each acupuncture point vocabulary is combined with its corresponding operation statement to form a structured expression, which is convenient for subsequent more in-depth analysis.
[0069] Preferably, in one embodiment of the present invention, word segmentation is performed on the text sequence to obtain all split words, and acupoint words are extracted from all the split words, including:
[0070] The Jieba word segmentation tool has a good effect on the word segmentation of Chinese texts and can accurately identify the boundaries of Chinese words. This is especially important for texts such as ancient Chinese medical books with special language expressions and complex vocabulary. Therefore, in an embodiment of the present invention, the Jieba word segmentation tool is used to perform word segmentation on the text sequence to obtain all the split words.
[0071] Then, all the split words are matched based on the preset acupoint vocabulary library, so as to extract the acupoint vocabulary from all the split words.
[0072] For example, if the text sequence is "Acupuncture treatment for headaches often involves pricking the temples to dredge the meridians and harmonize qi and blood," Jieba's word segmentation tool can be used to segment this text sequence into the following: " / acupuncture / treatment / headaches / often / acupuncture / at the temples / to / dredge / the meridians / and / harmony / qi and blood / ." "Temples" is an acupuncture point, so we can extract them based on the pre-set acupuncture point vocabulary.
[0073] It should be noted that the preset acupoint vocabulary library can use data mining technology to extract a large amount of acupoint vocabulary from traditional Chinese medicine literature, textbooks and online resources to construct the acupoint vocabulary library. It can also be constructed by other methods, which will not be elaborated or limited here; the use process of the Jieba word segmentation tool is a well-known technology and will not be elaborated here.
[0074] At this point, all acupoint words in the text sequence can be obtained. Since the same acupoint word may appear multiple times at different positions in the text sequence, the same acupoint word is regarded as the same acupoint word, that is, the temple is regarded as an acupoint word, which may appear at different positions in the text sequence.
[0075] Since in terms of word order, acupoint words are usually preceded and followed by specific operation methods or the therapeutic effects produced, and the operation methods required for different therapeutic effects may be different, different types of acupoint words contain different amounts of information, that is, there may be multiple situations for their corresponding acupuncture techniques or effects. Therefore, in order to improve the accuracy of the subsequent construction of the knowledge graph and better analyze the amount of information contained in each acupoint word, in an embodiment of the present invention, the operation statement corresponding to each acupoint word is determined based on the position of the acupoint word in the text sequence, thereby providing support for subsequent processes.
[0076] Preferably, in one embodiment of the present invention, the method for obtaining the operation statement includes:
[0077] In the text sequence, the split words between every two adjacent acupoint words are combined into a sentence sequence.
[0078] Then, in terms of word order, the sentence sequence adjacent to each acupoint vocabulary is combined with the acupoint vocabulary to form the operation sentence corresponding to each acupoint vocabulary.
[0079] For example, “ / Acupuncture / Treatment / Headache / Often / Use / Taiyang / Temple / For / Acupuncture / To / Achieve / Unblocking / Meridians / Harmonizing / Qi / and / Blood / ”, where “Taiyang / Temple” is an acupoint vocabulary, “Acupuncture / Treatment / Headache / Often / Use / Taiyang / Temple / For / Acupuncture / To / Achieve / Unblocking / Meridians / And / Harmonizing / Qi / and / Blood / ”, where “Taiyang / Temple” is an acupoint vocabulary, and “Acupuncture / Treatment / Headache / Often / Use / Taiyang / Temple / For / Acupuncture / To ...
[0080] Step S3: In all the operation statements corresponding to each acupoint vocabulary, the similarity relationship between the relative positions of the split vocabulary and the acupoint vocabulary, as well as the similarity between the split vocabulary are analyzed, so as to determine the vocabulary distribution complexity corresponding to the operation statement of each acupoint vocabulary; the description variability value of each acupoint vocabulary is obtained by comprehensively considering the quantity distribution characteristics of the split vocabulary in all the operation statements of each acupoint vocabulary and the vocabulary distribution complexity corresponding to the acupoint vocabulary.
[0081] Because the operation statement of each acupoint word contains the acupuncture operation of moving the needle on the acupoint, as well as the record of the corresponding symptoms and the effects produced, the description variability value of each acupoint word is quantified by analyzing the similarities and other features between the operation statements of the acupoint word, reflecting the diversity and complexity of the ways in which each acupoint word is described in ancient books, and is used to evaluate the amount of knowledge information of each acupoint word in ancient book chapters.
[0082] By analyzing the similarity between the relative positions of the split words and the acupoint words in all the operational statements corresponding to each acupoint word, we can reveal the collocation between the split words and the acupoint words in the operational statements when describing acupoints. The similarity between the split words also reflects the repetition and diversity of word usage when describing or operating on the same acupoint word. By combining these factors, we can quantify the lexical distribution complexity corresponding to the operational statements for each acupoint word. This metric helps to assess the complexity of acupoint word descriptions in ancient Chinese medical texts, thereby revealing the importance and uniqueness of acupoint word descriptions.
[0083] Preferably, in one embodiment of the present invention, the method for obtaining vocabulary distribution complexity includes:
[0084] See also Figure 2 , which shows a flow chart of a method for obtaining vocabulary distribution complexity in one embodiment of the present invention, the method comprising the following steps:
[0085] Step S301: In each operation statement, obtain the vocabulary similarity value between each split vocabulary and the corresponding acupoint vocabulary.
[0086] Based on the word2vec algorithm, the vocabulary vector corresponding to each split word and acupoint word in the text sequence is obtained.
[0087] In view of the different lengths of the operation statements of acupoint vocabulary, an attention window is preset to facilitate the subsequent analysis of vocabulary similarity values, so as to ensure that the number of split words involved in the calculation in different operation statements remains consistent.
[0088] Cosine similarity is a common method to measure the similarity between two vectors. Therefore, in the attention window corresponding to each acupoint word, the cosine similarity between the vocabulary vector of each split word and the vocabulary vector of each acupoint word is calculated as the vocabulary similarity value.
[0089] It should be noted that the process of obtaining word vectors using the word2vec algorithm is a well-known technology and will not be described in detail here. The length of the preset attention window is set to be centered on the acupoint vocabulary, with 5 split words taken before and after the word order. The specific length can be adjusted according to the actual scenario and is not limited here.
[0090] Step S302: In each acupoint vocabulary, target vocabulary and comparison vocabulary are determined, thereby obtaining all vocabulary combinations corresponding to each target vocabulary.
[0091] In each acupoint vocabulary, one acupoint vocabulary is selected as the target vocabulary, and the remaining vocabulary is selected as the comparison vocabulary. Then, the target vocabulary is combined with each comparison vocabulary in pairs, so as to obtain all vocabulary combinations corresponding to the target vocabulary.
[0092] For example, the acupoint word x appears three times in the text sequence, which are recorded as x1, x2 and x3 respectively. Then when x1 is used as the target word, the contrasting words are x2 and x3, and all the word combinations corresponding to the target word x1 are (x1, x2) and (x1, x3).
[0093] Step S303: In each vocabulary combination, the difference between the lexical similarity values of the split vocabulary in the attention window of the target vocabulary and the comparison vocabulary, as well as the similarity relationship between the relative positions of the split vocabulary and the acupoint vocabulary are analyzed to determine the lexical distribution complexity factor of the target vocabulary in each vocabulary combination.
[0094] In the target word's attention window, any split word is selected as the word to be analyzed. For each word combination in the target word, the absolute difference between the lexical similarity values of each split word in the attention window of the target word and the comparison word is calculated as the deviation factor. Because the target word and the comparison word belong to the same acupoint vocabulary, the smaller the deviation factor between the split words in the attention windows of the target word and the comparison word, the higher the similarity between the two words.
[0095] Therefore, the split word with the smallest deviation factor is used as the adaptive word for the word to be analyzed, and the character distance between the word to be analyzed and the target word, as well as the character distance between the adaptive word and the comparison word are obtained. The absolute value of the difference between these two character distances is used as the distance factor corresponding to the word to be analyzed. The larger the distance factor, the greater the deviation between the position distance between the word to be analyzed and the target word and the position distance between the adaptive word and the comparison word. In this case, when the word to be analyzed and the adaptive word are relatively similar, the position distribution of the two is quite different.
[0096] Then, the relative position between the word to be analyzed and the target word, and the relative position between the adaptation word and the comparison word, are analyzed, and the deviation relationship between them is determined to determine the position similarity factor corresponding to the word to be analyzed: the word order of the character sequence is taken as the positive direction, and the step length is one character.
[0097] In the attention window corresponding to the target word, the target word is taken as the starting point and the word to be analyzed is taken as the end point, thereby obtaining the position vector corresponding to the word to be analyzed.
[0098] In the attention window corresponding to the contrasting vocabulary, the contrasting vocabulary is used as the starting point and the adaptation vocabulary of the vocabulary to be analyzed is used as the end point, thereby obtaining the position vector corresponding to the adaptation vocabulary.
[0099] Calculate the cosine similarity between the two position vectors as the similarity coefficient. The range of the similarity coefficient is [-1, 1]. The closer it is to 1, the more similar it is. This means that the position distribution of the word to be analyzed and the corresponding adaptation word in the operation statement is more consistent. In order to facilitate subsequent calculations, the similarity coefficient is numerically adjusted so that its value range is [0, 1]. The closer it is to 1, the more consistent the position distribution is. The adjusted value is used as the position similarity factor corresponding to the word to be analyzed. The numerical adjustment process can be performed using the formula x represents the independent variable.
[0100] Then, the vocabulary distribution difference factor of the vocabulary to be analyzed is determined based on the distance factor and position similarity factor corresponding to the vocabulary to be analyzed. The formula model of the vocabulary distribution difference factor includes:
[0101]
[0102] Among them, CC represents the vocabulary distribution difference factor of the vocabulary to be analyzed; D represents the distance factor corresponding to the vocabulary to be analyzed; WX represents the position similarity factor corresponding to the vocabulary to be analyzed; γ1 represents the preset first parameter.
[0103] In the formula model of the vocabulary distribution difference factor, based on the above analysis, it can be seen that the larger the distance factor is, the greater the deviation between the position distance between the vocabulary to be analyzed and the target vocabulary and the position distance between the adaptive vocabulary and the comparison vocabulary. In this case, when the vocabulary to be analyzed and the adaptive vocabulary are relatively similar, the position distribution of the two is quite different; the larger the position similarity factor is, the more consistent the position distribution of the vocabulary to be analyzed and the adaptive vocabulary in their respective operation statements is, so the distance factor is positively correlated with the vocabulary distribution difference factor, and the position similarity factor is negatively correlated with the vocabulary distribution difference factor. Therefore, the above formula model is constructed to implement the above logic. At this time, the larger the vocabulary distribution difference factor is, the more dissimilar the position distribution of the vocabulary to be analyzed and its corresponding adaptive vocabulary in their respective operation statements is.
[0104] It should be noted that the purpose of presetting the first parameter γ1 is to prevent the denominator from being 0. Here, the value can be 0.001. The specific value can be adjusted according to the implementation scenario and is not limited here.
[0105] At this point, in each vocabulary combination, each split vocabulary in the attention window of the target vocabulary will correspond to a vocabulary distribution difference factor. Finally, the mean of all vocabulary distribution difference factors of the target vocabulary in each vocabulary combination will be used as the vocabulary distribution complexity factor of the target vocabulary in each vocabulary combination. The larger the vocabulary distribution complexity factor is, the less similar the distribution of vocabulary in the operation statements of the acupoint vocabulary in this vocabulary combination is, which means that in this vocabulary combination, different operations may have been performed on the same acupoint vocabulary, and different treatment effects have been obtained.
[0106] Step S304: The vocabulary distribution complexity factors of the target vocabulary in all vocabulary combinations are integrated to determine the vocabulary distribution complexity corresponding to the operation sentence of the target vocabulary.
[0107] The normalized value of the mean of the vocabulary distribution complexity factor of the target vocabulary in all vocabulary combinations is used as the vocabulary distribution complexity corresponding to the operation statement of the target vocabulary. The greater the vocabulary distribution complexity, the higher the semantic independence of the acupoint vocabulary, and the higher the complexity of the description of the acupoint vocabulary in ancient Chinese medical books, thereby increasing its importance and uniqueness accordingly. Normalization is a technical means well known to those skilled in the art, and the normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.
[0108] Since analyzing the difference in the number of split words in the two sentence sequences before and after the acupoint vocabulary can also characterize the distribution of the split words in the operation sentences of the acupoint vocabulary to a certain extent, thereby reflecting the complexity of the description of the acupoint vocabulary, in an embodiment of the present invention, the description variability value of each acupoint vocabulary is obtained by comprehensively considering the number distribution characteristics of the split words in all operation sentences of each acupoint vocabulary and the vocabulary distribution complexity corresponding to the acupoint vocabulary.
[0109] Preferably, in one embodiment of the present invention, the method for obtaining the description variability value includes:
[0110] For each acupoint vocabulary, the absolute value of the difference in the number of split words in the two sentence sequences before and after each acupoint vocabulary is counted as the vocabulary quantity difference value. The vocabulary quantity difference value can reflect the degree of change in the vocabulary quantity of each acupoint vocabulary when used in different contexts. The larger the value, the higher the degree of change in the vocabulary quantity.
[0111] At this time, each acupoint word in each acupoint word corresponds to a word quantity difference value. The range is an important indicator reflecting the range of data fluctuation. Therefore, here, the difference between the maximum and minimum values of the word quantity difference value is taken as the word quantity difference range. The word quantity difference range can reveal the maximum extent of the change in the word quantity in each acupoint word.
[0112] The ratio of the difference value of the number of words in each acupoint vocabulary to the extreme difference of the number of words is used as the first variable factor; the larger the first variable factor is, the more different the word distribution methods can be used to express information in the current acupoint vocabulary, which means that the current acupoint vocabulary is less targeted.
[0113] Then the ratio of the vocabulary distribution complexity corresponding to each acupoint vocabulary to the maximum value of the distribution complexity of all acupoint vocabulary is used as the second variable factor of each acupoint vocabulary. The larger the second variable factor is, the more complex the description of the acupoint vocabulary in ancient Chinese medical books is, and thus the importance and uniqueness should also be improved.
[0114] Finally, the product of the first variable factor and the second variable factor of each acupoint vocabulary is normalized to obtain the value of the description variability value of each acupoint vocabulary. At this time, the larger the description variability value, the more diverse the description of the acupoint vocabulary in ancient Chinese medical books. Therefore, it should receive higher attention in the subsequent process to ensure that the extracted information is more complete when the corresponding information is extracted. Normalization is a technical means well known to those skilled in the art. The normalization function can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.
[0115] Step S4: Compare the differences between the description variability values corresponding to all acupoint words to determine the retrieval attention of each acupoint word; mark the acupoint words in the text sequence based on the retrieval attention of all acupoint words to construct the acupoint knowledge graph.
[0116] Based on the above steps, the description variability value of each acupoint vocabulary can be obtained. In each acupoint vocabulary, when the description variability values of different acupoint vocabulary are relatively stable, it means that the description changes of this acupoint vocabulary in ancient books are relatively simple. Therefore, the logical structure expressed by the subsequent construction of the knowledge graph is less likely to have errors. Therefore, the differences between the description variability values of all acupoint vocabulary can be compared among all acupoint vocabulary to determine the retrieval attention of each acupoint vocabulary. The retrieval attention measures the importance of each acupoint vocabulary in the subsequent construction of the knowledge graph and the degree of attention required. Finally, based on the retrieval attention of all acupoint vocabulary, the acupoint vocabulary can be marked in the text sequence summary to construct the acupoint knowledge graph.
[0117] Preferably, in one embodiment of the present invention, the method for obtaining retrieval attention includes:
[0118] In each acupoint vocabulary, the mean of the description variability values of all acupoint vocabulary is taken as the mean eigenvalue, which represents the general level of the description variability of all acupoint vocabulary in this type of acupoint vocabulary.
[0119] Then, the mean of the mean characteristic values of all types of acupoint vocabulary is taken as the benchmark value. When the mean characteristic value of a certain acupoint vocabulary is greater than the benchmark value, it means that the acupoint vocabulary may contain more information, so a higher level of attention is required to capture more complete information; conversely, when the mean characteristic value of a certain acupoint vocabulary is less than the benchmark value, it means that the acupoint vocabulary contains less information, and a lower level of attention can capture all the information.
[0120] Therefore, in the embodiment of the present invention, the difference between the mean characteristic value of each acupoint word and the reference value is normalized and used as the search attention corresponding to each acupoint word. The normalization process here uses the Sigmoid() function.
[0121] Based on the above steps, the retrieval attention of each acupoint word in the text sequence can be obtained. Then each acupoint word in the text sequence also has a retrieval attention. Therefore, the acupoint words in the text sequence are marked based on the retrieval attention of all acupoint words, which can be used to construct an acupoint knowledge graph.
[0122] Preferably, in one embodiment of the present invention, the process of constructing the acupoint knowledge graph includes:
[0123] In the text sequence, all acupoint words are marked based on their retrieval attention to obtain a marked text sequence, which helps the SDP tool and neural network to recognize and process acupoint words in the subsequent process.
[0124] Semantic Dependency Parsing (SDP) is a task that analyzes the semantic relationships between words in a sentence and represents them as a graph structure. Therefore, the labeled text sequence is used as the input of the SDP tool to obtain the SDP node graph. Figure 3 , which shows a schematic diagram of an SDP node graph provided in one embodiment of the present invention.
[0125] The SDP node graph is then input into a pre-trained neural network. Through the trained neural network, the weight of each statement path in the SDP node graph can be evaluated, thereby obtaining the weight of each statement path in the SDP node graph.
[0126] Finally, all the statement paths containing acupoint vocabulary in the SDP node graph are extracted to obtain the acupoint knowledge graph.
[0127] It should be noted that the use of the SDP tool and the training process of the neural network are well-known technologies and will not be described in detail here; the neural network here can use GNN; since some acupoint words may appear adjacent to each other in the text sequence, when constructing the acupoint knowledge graph, the sentence paths of adjacent acupoint words can also be merged to obtain the final acupoint knowledge graph.
[0128] In summary, obtaining a sequence of characters in ancient Chinese medical books is a step used to obtain ancient book data and provide a data source for subsequent processing. By word segmentation, split words are obtained, and complex character sequences can be converted into easy-to-process word units, thereby facilitating the extraction of acupoint words from all the split words. Since in ancient Chinese medical books, the description sentences or operation sentences for different acupoint words are different, and even the same acupoint words have different description sentences or operation sentences under different semantics or situations, the present invention determines the operation sentence corresponding to each acupoint word based on the position of the acupoint word in the character sequence, providing key information for the subsequent construction of an accurate knowledge graph. Furthermore, in all the operation sentences corresponding to each acupoint word, the similarity relationship between the relative position of the split words and the acupoint words, as well as the similarity between the split words, is analyzed to obtain the vocabulary distribution complexity corresponding to each acupoint word; it can provide an in-depth understanding of the context of the acupoint words in ancient books, and explain the distribution rules of the words in ancient books, thereby providing important clues for accurately interpreting acupoint knowledge and enhancing the accuracy and readability of the subsequent knowledge graph construction. Furthermore, by comprehensively considering the quantitative distribution characteristics of the split words in the operational statements of each acupoint vocabulary and the complexity of the lexical distribution corresponding to the acupoint vocabulary, we can obtain the descriptive variability value of the acupoint vocabulary. The descriptive variability value can be used to evaluate the tolerance of the information representation of acupoint vocabulary in ancient books, that is, the diversity of acupoint vocabulary descriptions. Finally, by comparing the differences between the descriptive variability values of acupoint vocabulary, we can determine the retrieval attention of each acupoint vocabulary. The retrieval attention can be used to describe the importance and priority of each acupoint vocabulary in the construction of the knowledge graph. Therefore, based on the retrieval attention of all acupoint vocabulary, acupoint vocabulary can be marked in the text sequence, thereby constructing a more accurate and easy-to-search acupoint knowledge graph.
[0129] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0130] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A method for establishing an acupuncture point knowledge base using ancient Chinese medical books, characterized in that: The method comprises: Obtain text sequences from ancient Chinese medical books; Performing vocabulary segmentation on the text sequence to obtain all split words, and extracting acupoint words from all the split words, wherein the same acupoint words belong to the same type of acupoint words; determining an operation statement corresponding to each acupoint word based on the position of the acupoint word in the text sequence; In all the operation statements corresponding to each acupoint vocabulary, the similarity between the relative positions of the split words and the acupoint vocabulary, as well as the similarity between the split words, are analyzed to determine the lexical distribution complexity corresponding to the operation statements of each acupoint vocabulary; the quantitative distribution characteristics of the split words in all the operation statements of each acupoint vocabulary and the lexical distribution complexity corresponding to the acupoint vocabulary are combined to obtain the descriptive variability value of each acupoint vocabulary; The differences between the description variability values corresponding to all acupoint words are compared to determine the retrieval attention of each acupoint word; the acupoint words are marked in the text sequence based on the retrieval attention of all acupoint words to construct an acupoint knowledge graph.
2. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The method for obtaining the character sequence includes: Obtain grayscale text images from ancient Chinese medical books, and perform threshold segmentation and connected domain detection on the grayscale text images, thereby obtaining the minimum circumscribed rectangle of each connected domain as a character selection box for each character; The character selection box is recognized based on the OCR tool to obtain the corresponding characters. After the characters are unified and special characters are removed, the text sequence is obtained based on the preset reading order.
3. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The word segmentation of the text sequence to obtain all split words and extracting acupoint words from all the split words includes: Performing word segmentation on the text sequence based on the Jieba word segmentation tool to obtain all the split words; All split words are matched based on a preset acupoint vocabulary library, so as to extract acupoint words from all split words.
4. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The method for obtaining the operation statement includes: In the text sequence, the split words between each two adjacent acupoint words are combined into a sentence sequence; In terms of word order, the sentence sequence adjacent to each acupoint vocabulary is combined with the acupoint vocabulary to form the operation sentence corresponding to each acupoint vocabulary.
5. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The method for obtaining the vocabulary distribution complexity includes: Obtain the vocabulary vector corresponding to each split word and acupoint word in the text sequence based on the word2vec algorithm; Preset an attention window, and in the attention window corresponding to each acupoint vocabulary, calculate the cosine similarity between the vocabulary vector of each split vocabulary and the vocabulary vector of each acupoint vocabulary as the vocabulary similarity value; In each acupoint vocabulary, one acupoint vocabulary is selected as the target vocabulary, and the remaining vocabulary is selected as the contrast vocabulary. The target vocabulary is combined with each contrast vocabulary in pairs, so as to obtain all vocabulary combinations corresponding to the target vocabulary. In each vocabulary combination, the difference between the lexical similarity values of the split words in the attention window of the target word and the comparison word, as well as the similarity relationship between the relative positions of the split words and the acupoint words were analyzed to determine the lexical distribution complexity factor of the target word in each vocabulary combination; The normalized value of the mean of the vocabulary distribution complexity factors of the target vocabulary in all vocabulary combinations is used as the vocabulary distribution complexity corresponding to the operation statement of the target vocabulary.
6. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 5, characterized in that: The method for obtaining the vocabulary distribution complexity factor includes: In the attention window of the target vocabulary, any split vocabulary is selected as the vocabulary to be analyzed. In each vocabulary combination of the target vocabulary, the absolute value of the difference between the vocabulary similarity value of each split vocabulary in the attention window of the target vocabulary and the comparison vocabulary is calculated as the deviation factor. The split word with the smallest deviation factor is used as the adaptation word of the word to be analyzed, and the character distance between the word to be analyzed and the target word, as well as the character distance between the adaptation word and the comparison word, is obtained. The absolute value of the difference between the two character distances is used as the distance factor corresponding to the word to be analyzed; Analyze the relative position between the word to be analyzed and the target word, and the relative position between the adaptation word and the comparison word, and determine the position similarity factor corresponding to the word to be analyzed; Determining a vocabulary distribution difference factor of the vocabulary to be analyzed based on a distance factor and a position similarity factor corresponding to the vocabulary to be analyzed, wherein the distance factor is positively correlated with the vocabulary distribution difference factor, and the position similarity factor is negatively correlated with the vocabulary distribution difference factor; The mean of all lexical distribution difference factors of the target vocabulary in each lexical combination is taken as the lexical distribution complexity factor of the target vocabulary in each lexical combination.
7. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 6, characterized in that: The method for obtaining the position similarity factor includes: The word order of the character sequence is the positive direction, and the step length is one character; In the attention window corresponding to the target word, the target word is taken as the starting point and the word to be analyzed is taken as the end point, thereby obtaining the position vector corresponding to the word to be analyzed; In the attention window corresponding to the contrasting vocabulary, the contrasting vocabulary is used as the starting point and the matching vocabulary of the vocabulary to be analyzed is used as the end point, thereby obtaining the position vector corresponding to the matching vocabulary; The cosine similarity between the two position vectors is calculated and the numerical adjustment is performed to obtain the position similarity factor corresponding to the word to be analyzed.
8. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 4, characterized in that: The method for obtaining the description variability value includes: In each acupoint vocabulary, the absolute value of the difference in the number of split words in the two sentence sequences before and after each acupoint vocabulary is counted as the vocabulary number difference value; The difference between the maximum and minimum values of the vocabulary quantity difference value is taken as the vocabulary quantity difference range; The ratio of the vocabulary quantity difference value of each acupoint vocabulary to the extreme difference of the vocabulary quantity difference is used as the first variable factor; The ratio of the vocabulary distribution complexity corresponding to each acupoint vocabulary to the maximum value of the distribution complexity of all acupoint vocabulary is used as the second variable factor of each acupoint vocabulary; The value obtained by normalizing the product of the first variable factor and the second variable factor of each acupoint vocabulary is used as the description variable degree value of each acupoint vocabulary.
9. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The method for obtaining the retrieval attention includes: In each acupoint vocabulary, the mean of the description variability values of all acupoint vocabulary is taken as the mean feature value; The mean of the mean characteristic values of all acupoint words was taken as the baseline value; The difference between the mean characteristic value of each acupoint word and the reference value is normalized and used as the retrieval attention corresponding to each acupoint word.
10. The method for establishing an acupuncture point knowledge base using ancient Chinese medical books according to claim 1, characterized in that: The acupoint vocabulary is marked in the text sequence based on the retrieval attention of all acupoint vocabulary to construct an acupoint knowledge graph, including: In the text sequence, all acupoint words are marked based on their retrieval attention, to obtain a marked text sequence; The marked text sequence is used as the input of the SDP tool to obtain the SDP node graph; Inputting the SDP node graph into a pre-trained neural network to obtain the weights of each statement path in the SDP node graph; All sentence paths containing acupoint vocabulary in the SDP node graph are extracted to obtain the acupoint knowledge graph.