Cognitive heuristic-based efficient instruction fine-tuning method and device for traditional chinese medicine large model

By conducting multi-dimensional evaluations of traditional Chinese medicine datasets, high-quality fine-tuning datasets were selected. The large-scale traditional Chinese medicine question-and-answer model was then fine-tuned, solving the evaluation difficulties in existing technologies, improving the model's accuracy and generalization ability, and reducing data requirements.

CN122472216APending Publication Date: 2026-07-28北京中科闻歌科技股份有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京中科闻歌科技股份有限公司
Filing Date
2026-06-26
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

When using large language models for fine-tuning in the field of traditional Chinese medicine, existing technologies lack intuitive and quantifiable evaluation indicators, making it impossible to accurately assess the quality of the fine-tuned dataset and the model's performance.

Method used

By conducting data quality analysis on the original TCM dataset, the first TCM dataset was selected. Based on the internal and external difficulty, the second TCM dataset was further selected. Finally, the TCM fine-tuning dataset was used to fine-tune the TCM question-and-answer model, and a transparent and quantifiable method for determining fine-tuning data was constructed.

Benefits of technology

It enabled accurate evaluation of the large-scale TCM question-and-answer model, ensured high dataset quality, improved the model's fine-tuning effect and generalization ability, reduced data requirements, and achieved cost reduction and efficiency improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472216A_ABST
    Figure CN122472216A_ABST
Patent Text Reader

Abstract

This disclosure relates to a cognitively inspired, efficient method and apparatus for fine-tuning a large-scale TCM (Traditional Chinese Medicine) model, applicable to the field of machine learning technology. The method involves performing data quality analysis on the original TCM dataset, obtaining a first TCM dataset from it, and then obtaining a second TCM dataset based on the inherent difficulty of data representation in the first dataset. Based on the external difficulty of data representation in the second TCM dataset, a fine-tuning dataset is obtained from the second dataset. This fine-tuning dataset is then used to fine-tune a pre-designed large-scale TCM question-and-answer model, resulting in a well-tuned model. Therefore, by constructing a transparent and quantifiable method for determining fine-tuning data from multiple evaluation perspectives, it is beneficial to trace which data is of high quality, thereby achieving the goal of quantifying and explaining the quality of the fine-tuning dataset, and ultimately facilitating the accurate evaluation of the fine-tuning effect of the large-scale TCM question-and-answer model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of machine learning technology, and in particular to a method and apparatus for efficient instruction fine-tuning of a large-scale TCM model based on cognitive inspiration. Background Technology

[0002] Large Language Models (LMMs) have demonstrated powerful capabilities in general tasks, but adapting them to the field of traditional Chinese medicine (TCM) relies heavily on fine-tuned datasets specific to TCM.

[0003] To select fine-tuning datasets from massive datasets in the field of Traditional Chinese Medicine (TCM), related techniques utilize large models to score and filter these datasets, thereby identifying the fine-tuning datasets. These fine-tuning datasets are then used to fine-tune the large-scale TCM question-answering model. However, this approach essentially exploits the black-box logic of large models, making it difficult to trace which data points are of high quality. This results in an inability to quantify and interpret the quality of the fine-tuning datasets, thus hindering the evaluation of the fine-tuning effect on the large-scale TCM question-answering model. Summary of the Invention

[0004] To address the aforementioned technical issues, this disclosure provides a method and apparatus for efficient instruction fine-tuning of a large-scale TCM model based on cognitive inspiration.

[0005] Firstly, this disclosure provides a method for efficient instruction fine-tuning of a large-scale TCM model based on cognitive inspiration, including: Perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset, and obtain the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset; Based on the inherent difficulty of representing traditional Chinese medicine data in the first traditional Chinese medicine dataset, a second traditional Chinese medicine dataset is obtained from the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. Based on the external difficulty of the TCM data in the second TCM dataset, a TCM fine-tuning dataset is obtained from the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The pre-set TCM question-and-answer model was fine-tuned using the TCM fine-tuning dataset to obtain a fine-tuned model.

[0006] Secondly, this disclosure provides a highly efficient instruction fine-tuning device for a large-scale TCM model based on cognitive inspiration, comprising: The first acquisition module is used to perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset and acquire the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset; The second acquisition module is used to acquire a second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset based on the inherent difficulty of the traditional Chinese medicine data representation in the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. The third acquisition module is used to acquire a TCM fine-tuning dataset from the second TCM dataset based on the external difficulty of the TCM data in the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The fine-tuning module is used to fine-tune the preset TCM question-and-answer model using the TCM fine-tuning dataset to obtain a fine-tuned model.

[0007] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the methods provided in the first aspect.

[0008] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method provided in the first aspect.

[0009] The technical solution provided in this disclosure has the following advantages compared with the prior art: This disclosure discloses a method, apparatus, device, and medium for efficient fine-tuning of a large-scale TCM (Traditional Chinese Medicine) model based on cognitive heuristics. For a massive original TCM dataset, the method first performs data quality analysis on the TCM-related data within the original dataset, obtaining a first TCM dataset. Then, based on the inherent difficulty of representing the TCM-related data in the first dataset, a second TCM dataset is obtained. Next, based on the external difficulty of representing the TCM-related data in the second dataset, a fine-tuning dataset is obtained. Finally, the fine-tuning dataset is used to fine-tune a pre-defined large-scale TCM question-and-answer model, resulting in a well-tuned model. Therefore, by constructing a transparent and quantifiable method for determining fine-tuning data through multi-dimensional evaluation, it is beneficial to trace which TCM-related data is of high quality, thereby achieving the goal of quantifying and explaining the quality of the fine-tuning dataset and facilitating accurate evaluation of the fine-tuning effect of the large-scale TCM question-and-answer model. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0011] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating an efficient instruction fine-tuning method for a large-scale TCM model based on cognitive inspiration, provided in this embodiment of the disclosure; Figure 2 A logical diagram illustrating an efficient instruction fine-tuning method for a large-scale TCM model based on cognitive inspiration, provided in this embodiment of the disclosure; Figure 3 A schematic diagram of the structure of a high-efficiency instruction fine-tuning device for a large-scale model of traditional Chinese medicine based on cognitive inspiration, provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0013] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0014] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0015] In related technologies, the method of selecting fine-tuned datasets from massive datasets in the field of traditional Chinese medicine by utilizing the black-box logic of large models lacks intuitive and quantifiable evaluation indicators, and the generalization ability of the basic model is limited.

[0016] One of the shortcomings is the lack of intuitive and quantifiable evaluation indicators. Specifically, it means that specific quantitative dimensions cannot be provided, making it impossible to directly apply these evaluation standards to the process of artificial data construction and annotation, and thus failing to provide clear guidance for data construction in vertical fields.

[0017] One drawback is the limited generalization ability of the basic model. Specifically, it removes only obviously low-quality data, but retains a large number of redundant or "simple samples" with extremely low cognitive difficulty. These samples cannot effectively improve the model's boundary capabilities and may even impair the model's generalization ability on complex tasks.

[0018] To address the aforementioned issues, this embodiment provides a cognitively inspired, highly efficient instruction fine-tuning method for a large-scale TCM model. The following is a combination of... Figures 1-2 This disclosure describes a method for efficient instruction fine-tuning of a large-scale TCM model based on cognitive inspiration, provided in this embodiment. In this embodiment, the method can be executed by an electronic device or a server. The electronic device can include a desktop computer, laptop, tablet, or other smart devices. The server can include a cloud server or a server cluster. Specifically, this embodiment demonstrates that the method can be executed using an electronic device.

[0019] Figure 1 The illustration shows a flowchart of an efficient instruction fine-tuning method for a large-scale TCM model based on cognitive inspiration provided in this disclosure.

[0020] As shown in Figure 1, the efficient instruction fine-tuning method for the large-scale TCM model based on cognitive inspiration may include the following steps.

[0021] S110. Perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset, and obtain the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset.

[0022] In this embodiment, the electronic device can convert the original TCM dataset into structured data and then perform initial quality screening to remove low-quality data that contains obvious factual errors, logical inconsistencies, or does not conform to user preferences, thereby obtaining the initial screened dataset.

[0023] The original TCM dataset can be a dataset within the field of TCM. Optionally, the original TCM dataset includes TCM prescription data, modern pharmacology data, etc.

[0024] Among them, the first traditional Chinese medicine dataset can be a dataset with a high quality score.

[0025] In some embodiments, the specific implementation method of S110 includes: inputting the original traditional Chinese medicine dataset into a preset quality analysis model, performing data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset, and obtaining the quality score of each traditional Chinese medicine data in the original traditional Chinese medicine dataset; and obtaining the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset based on the quality score of each traditional Chinese medicine data.

[0026] Specifically, the electronic device can use a high-performance reward model as a preset quality analysis model to calculate the reward score for each instruction-response pair in the original TCM dataset, and use the reward score as the quality score for each TCM category of data. Then, the quality scores of each TCM category of data in the original TCM dataset are sorted in descending order, and a preset number of TCM category data with the lowest quality scores are filtered out from the original TCM dataset, or TCM category data with quality scores less than a preset score threshold are filtered out from the original TCM dataset. The first TCM dataset is constructed based on the remaining TCM category data.

[0027] Optionally, the reward score for each instruction-response pair in the original TCM dataset can be calculated using the following method:

[0028] in, It consists of each instruction-response pair in the original Traditional Chinese Medicine dataset; It is the quality score of each instruction-response pair, which serves as the quality score for each category of traditional Chinese medicine data.

[0029] Optionally, the preset quantity can be the last 80% of the quantity, or it can be any other quantity.

[0030] Optionally, the preset score threshold can be 0.8 or other data.

[0031] In this way, by conducting quality analysis on the TCM-related data in the original TCM dataset, it can be ensured that the retained first TCM dataset has basic answer accuracy and semantic coherence.

[0032] S120. Based on the inherent difficulty of data representation in the first traditional Chinese medicine dataset, obtain the second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset.

[0033] To further filter out high-value datasets, electronic devices assess the inherent knowledge depth and breadth involved in each piece of TCM data in the first TCM dataset to determine the inherent difficulty of each data representation. Based on the inherent difficulty, the first TCM dataset is further filtered to determine the second TCM dataset.

[0034] The inherent difficulty represents the depth of content corresponding to each piece of TCM-related data in the field of Chinese medicine.

[0035] The second TCM dataset is a collection of TCM-related data that is inherently more difficult to analyze.

[0036] In some embodiments, the specific implementation method of S120 includes: S1201, performing cognitive reasoning analysis on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the reasoning ability score of each traditional Chinese medicine data in the first traditional Chinese medicine dataset; S1202, performing cross-domain analysis on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the information complexity of each traditional Chinese medicine data in the first traditional Chinese medicine dataset; S1203, determining the intrinsic difficulty score of the representation of each traditional Chinese medicine data based on the reasoning ability score and the information complexity of each traditional Chinese medicine data; S1204, obtaining the second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset according to the intrinsic difficulty score of the representation of each traditional Chinese medicine data.

[0037] The specific implementation method of S1201 includes: inputting the first traditional Chinese medicine dataset into a preset classification model, classifying and parsing the data in the first traditional Chinese medicine dataset to obtain the cognitive level to which each traditional Chinese medicine data belongs; and performing Bloom cognitive calculation on the data in the first traditional Chinese medicine dataset based on the progressive difficulty between the cognitive levels to which different traditional Chinese medicine data belongs to, to obtain the reasoning ability score of each traditional Chinese medicine data.

[0038] Optionally, the preset classification model can be a base large language model.

[0039] Optionally, the cognitive hierarchy may include levels such as memory, understanding, application, analysis, evaluation, and creation, and these levels are ordered from low to high.

[0040] Optionally, Bloom's cognitive computation can be performed on the data in the first traditional Chinese medicine dataset, which can be achieved through the following method:

[0041] in, This indicates whether the data related to traditional Chinese medicine belongs to the corresponding cognitive level. (Lower-level weight coefficients are small, higher-level weight coefficients are large). and These represent the minimum and maximum scores for reasoning ability, respectively. It is the Bloom score for TCM-related data, which is the reasoning ability score. The higher the score, the stronger the cognitive reasoning ability required for that TCM-related data.

[0042] The specific implementation method of S1202 includes: using a preset vector extraction model to vectorize the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain domain description features; and calculating the complexity of the domain description features to obtain the information complexity of each type of traditional Chinese medicine data.

[0043] Among them, the domain description feature is used to describe the subject characteristics associated with each piece of traditional Chinese medicine data.

[0044] Information complexity is used to characterize the breadth of disciplines that each piece of TCM data needs to traverse. Optionally, higher information complexity means more disciplines involved and greater semantic differences between disciplines; conversely, lower information complexity means fewer disciplines involved and smaller semantic differences between disciplines.

[0045] Specifically, the electronic device first extracts the core subject set associated with each piece of TCM data based on the large model in the preset vector extraction model, generates a detailed description text, and then transforms the detailed description text into a high-dimensional dense feature vector based on the text embedding network in the preset vector extraction model, thereby obtaining the domain description features; then, it calculates the number of subjects contained in each piece of TCM data and the sum of the cosine distances between the vectors of each subject to complete the complexity calculation and obtain the information complexity of each piece of TCM data.

[0046] Optionally, the information complexity of each TCM category data can be obtained by calculating the complexity of the domain description features. This can be achieved through the following method:

[0047] in, This refers to current data related to traditional Chinese medicine. The related disciplines involved; and These represent the size of the set containing the minimum and maximum number of disciplines in the entire dataset, respectively. This refers to the number of combinations of any two disciplines selected from the first dataset of traditional Chinese medicine. This refers to the characteristic vector of subject description. and The cosine distance between them.

[0048] Optionally, feature vector and The cosine distance between them can be determined by the following method:

[0049] Optionally, and The greater the cosine distance between them, the more difficult it is to integrate the table name information, which further leads to a high score for interdisciplinary complexity.

[0050] The specific implementation methods of S1203 and S1204 are as follows: the reasoning ability score and information complexity of each TCM data are weighted and averaged to obtain the intrinsic difficulty score of each TCM data representation. Then, the intrinsic difficulty scores of each TCM data representation are sorted in descending order. Data with an intrinsic difficulty score lower than a predetermined number are filtered out from the first TCM dataset, or data with an intrinsic difficulty score lower than a predetermined score threshold are filtered out from the first TCM dataset. The second TCM dataset is constructed based on the remaining TCM data.

[0051] Optionally, the preset quantity can be the last 50% of the quantity, or it can be any other quantity.

[0052] Optionally, the preset score threshold can be 0.5 or other data.

[0053] In this way, by penetrating the knowledge reasoning hierarchy involved in the semantic evaluation instructions, the inherent difficulty of representing TCM-related data can be assessed, and data with low cognitive difficulty can be eliminated by utilizing the inherent difficulty of TCM-related data representation.

[0054] S130. Based on the external difficulty of representing traditional Chinese medicine data in the second traditional Chinese medicine dataset, obtain the traditional Chinese medicine fine-tuning dataset from the second traditional Chinese medicine dataset.

[0055] To further filter out high-value datasets, electronic devices assess the learning challenge of each TCM-related data point in the second TCM dataset in terms of text structure and distribution characteristics from the perspective of text morphological scalability and internal distribution characteristics of the dataset. This determines the external difficulty of representing each TCM-related data point, and then further filters the second TCM dataset based on the external difficulty to determine the TCM fine-tuning dataset.

[0056] Among them, external difficulty represents the breadth of content corresponding to each piece of TCM-related data in the field of Chinese medicine.

[0057] Among them, the TCM fine-tuning dataset is a collection of data that is externally challenging.

[0058] In some embodiments, the specific implementation method of S130 includes: S1301, performing a burden analysis on the traditional Chinese medicine data in the second traditional Chinese medicine dataset to obtain a burden score for each traditional Chinese medicine data in the second traditional Chinese medicine dataset; S1302, performing an isolation analysis on the traditional Chinese medicine data in the second traditional Chinese medicine dataset to obtain an isolation score for each traditional Chinese medicine data in the second traditional Chinese medicine dataset; S1303, determining the external difficulty score for representing each traditional Chinese medicine data based on the burden score and the isolation score of each traditional Chinese medicine data; S1304, obtaining a traditional Chinese medicine fine-tuning dataset from the second traditional Chinese medicine dataset based on the external difficulty score for representing each traditional Chinese medicine data.

[0059] The specific implementation method of S1301 includes: obtaining instruction data with instruction semantic information and response data with response semantic information from the TCM category data contained in the second TCM dataset; merging the instruction data and response data in the second TCM dataset to obtain instruction-response merged data; determining the first length corresponding to the instruction data, the second length corresponding to the response data, the minimum length of the instruction-response merged data, and the maximum length of the instruction-response merged data; and calculating the instruction-response expansion index of each TCM category data in the second TCM dataset based on the first length, the second length, the minimum length, and the maximum length, as the burden score of each TCM category data in the second TCM dataset.

[0060] Among them, the response expansion index is used to characterize the degree to which traditional Chinese medicine data provides contextual information.

[0061] Optionally, a larger response expansion index indicates that the TCM data provides limited contextual information, requiring the model to call more intrinsic world knowledge to generate long texts. Therefore, the burden score of TCM data is lower. Conversely, a smaller response expansion index indicates that the TCM data provides more contextual information, requiring the model to call less intrinsic world knowledge to generate long texts. Therefore, the burden score of TCM data is higher.

[0062] Optionally, the expansion index of each TCM category data instruction response in the second TCM dataset can be determined using the following method:

[0063] in, This refers to the first length corresponding to the instruction data; This refers to the second length of the response data; This refers to the minimum length of the merged data in the instruction response; This refers to the maximum length of the merged data in the instruction response.

[0064] The specific implementation method of S1302 includes: converting the TCM data in the second TCM dataset into vector representation; performing cluster analysis on the vector representation to obtain multiple clusters corresponding to the second TCM dataset; for the current TCM data in the second TCM dataset, obtaining the nearest neighbor cluster and the cluster to which the current TCM data belongs from the multiple clusters; determining the first average distance between the current TCM data and each data in the nearest neighbor cluster, and determining the second average distance between the current TCM data and each data in its own cluster; and calculating the isolation degree of the TCM data in the second TCM dataset based on the first average distance and the second average distance to obtain the isolation degree score of each TCM data.

[0065] Specifically, firstly, the electronic device maps the TCM category data in the second TCM dataset into vector representations (e.g., TF-IDF vector representations). Then, it uses a preset clustering algorithm (e.g., K-Means clustering algorithm) to cluster the second TCM dataset, resulting in multiple clusters. Finally, for each TCM category data in the second TCM dataset and the distance between each TCM category data in its corresponding cluster, the isolation degree of each TCM category data is calculated to obtain the isolation degree score of each TCM category data.

[0066] Optionally, the isolation degree of each TCM-related data point can be calculated using the following method:

[0067] in, This refers to the first average distance between the current TCM data and all TCM data in the nearest neighbor cluster; It refers to the second average distance between the current TCM data and the TCM data in its respective cluster.

[0068] The specific implementation methods of S1303 and S1304 are as follows: the burden score and the isolation score of each data are weighted and averaged to obtain the external difficulty score of each data. Then, the external difficulty scores of each data are sorted in descending order. Data with a lower external difficulty score than a predetermined number are filtered out from the second TCM dataset, or data with an external difficulty score less than a predetermined score threshold are filtered out from the second TCM dataset. The remaining data are used to construct a TCM fine-tuning dataset.

[0069] Optionally, the preset quantity can be the last 50% of the quantity, or it can be any other quantity.

[0070] Optionally, the preset score threshold can be 0.5 or other data.

[0071] In this way, by combining external difficulties such as text expansion ratio and long-tail distribution characteristics in global space, redundant and simple samples are eliminated by the elimination rate, and only long-tail difficult samples that are unique in feature space and have low isolation are retained, thus achieving a comprehensive and three-dimensional capture of the potential training value of fine-tuning data.

[0072] S140. Use the TCM fine-tuning dataset to fine-tune the preset TCM question-and-answer model to obtain the fine-tuned model.

[0073] In this embodiment, within the TCM question-and-answer vertical domain to which the TCM fine-tuning dataset belongs, considering the requirements of interdisciplinary collaboration, detailed answers, and relatively small dataset size, this embodiment uses the TCM fine-tuning dataset as the labeled fine-tuning dataset for the TCM question-and-answer domain. This labeled fine-tuning dataset is then used to fine-tune the pre-set TCM question-and-answer model, resulting in a fine-tuned model. This approach eliminates the need to blindly collect tens of thousands of domain-specific data points, enabling the acquisition of a fine-tuning dataset within the vertical domain for labeling guidance.

[0074] In this way, while providing clear guidance for data construction in vertical industries, the goal of reducing costs and increasing efficiency is also achieved.

[0075] In some embodiments, after performing S130, the method further includes: using a traditional Chinese medicine fine-tuning dataset to fine-tune a pre-trained model in the general domain to obtain a fine-tuned model in the general domain.

[0076] In the general domain of models, the TCM fine-tuning dataset is used as the training set, and techniques for efficiently fine-tuning large pre-trained models (such as low-rank adaptation techniques) are used to perform supervised fine-tuning on the pre-trained models in the general domain, resulting in the fine-tuned models in the general domain.

[0077] In this way, the fine-tuned model will outperform the model trained using the original TCM dataset in generalization tests.

[0078] This disclosure discloses a cognitively inspired, efficient instruction fine-tuning method for a large-scale TCM (Traditional Chinese Medicine) model. For a massive original TCM dataset, the method first performs data quality analysis on the TCM-related data within the original dataset, obtaining a first TCM dataset. Then, based on the inherent difficulty of representing the TCM-related data in the first dataset, a second TCM dataset is obtained. Next, based on the external difficulty of representing the TCM-related data in the second dataset, a fine-tuning dataset is obtained. Finally, the fine-tuning dataset is used to fine-tune a pre-defined large-scale TCM question-and-answer model, resulting in a well-tuned model. Therefore, by constructing a transparent and quantifiable method for determining fine-tuning data through multi-dimensional evaluation, it is beneficial to trace which TCM-related data is of high quality, thereby achieving the goal of quantifying and explaining the quality of the fine-tuning dataset and facilitating accurate evaluation of the fine-tuning effect of the large-scale TCM question-and-answer model.

[0079] In another embodiment of this application, the overall implementation logic of the efficient instruction fine-tuning method for a large-scale TCM model based on cognitive inspiration is explained.

[0080] Figure 2 This illustration shows a logical diagram of another cognitively inspired method for efficient instruction fine-tuning of a large-scale TCM model, as provided in an embodiment of this disclosure.

[0081] like Figure 2 As shown, the efficient instruction fine-tuning method for the large-scale TCM model based on cognitive inspiration can include the following steps.

[0082] S210. Input the original TCM dataset into the reward model to score the data quality, and obtain the quality score of each TCM data in the original TCM dataset.

[0083] Specifically, the electronic device will parse the original TCM dataset into a structured format, and then call a high-performance reward model to perform inference and scoring on the structured TCM dataset to obtain a quality score for each data point.

[0084] S220. Sort the quality scores of each TCM data in descending order, and remove the low-quality data in the bottom 80% to obtain the first TCM dataset.

[0085] Specifically, electronic devices are sorted in descending order of quality score, and only the top 20% of samples are selected. Invalid data with logical fallacies or low value are removed to form a high-quality first TCM dataset.

[0086] S230. Perform Bloom cognitive analysis on each piece of traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain a reasoning ability score for each piece of traditional Chinese medicine data.

[0087] Specifically, the electronic device first constructs prompt words and inputs the definitions of Bloom's six levels (memory, comprehension, application, analysis, evaluation, and creation) into the base model; then, the model performs multi-label classification on each piece of traditional Chinese medicine data in the first traditional Chinese medicine dataset and outputs the corresponding level judgment (0 or 1); next, it calculates and normalizes the Bloom score to obtain the reasoning ability score of each piece of traditional Chinese medicine data, with higher levels having greater weight scores.

[0088] S240. Perform interdisciplinary complexity calculation on each piece of TCM data in the first TCM dataset to obtain the interdisciplinary complexity score for each piece of TCM data.

[0089] Specifically, the electronic device uses a large model to infer the core subject categories involved in the instructions, then uses a text embedding model to transform the descriptions of each subject category into high-dimensional dense vectors. Finally, by calculating the cosine distance between subject vectors, the subject span is quantified, and combined with the number of subjects, the cross-disciplinary complexity of each TCM-related data in the first TCM dataset is calculated to obtain the cross-disciplinary complexity score of each TCM-related data.

[0090] S250. The reasoning ability score and the interdisciplinary complexity score of each piece of TCM data are weighted and fused to obtain the intrinsic difficulty score of each piece of TCM data.

[0091] Specifically, the electronic device takes the average of the reasoning ability score and the interdisciplinary complexity score of each piece of TCM data as the intrinsic difficulty score of each piece of TCM data.

[0092] S260. Sort the intrinsic difficulty scores of each TCM data in descending order, and remove the low cognitive difficulty data in the bottom 50% to obtain the second TCM dataset.

[0093] Specifically, electronic devices are sorted in descending order of their inherent difficulty scores. The first TCM dataset is sorted by inherent difficulty, and the top 50% of the more difficult data are retained to obtain the second TCM dataset.

[0094] S270. Determine the instruction response expansion index for each piece of traditional Chinese medicine data in the second traditional Chinese medicine dataset.

[0095] Specifically, the electronic device uses a string processing library to extract the first length of the instruction data, the second length of the response data, the minimum length of the merged instruction and response data, and the maximum length of the merged instruction and response data. Then, based on the first length, the second length, the minimum length, and the maximum length, it calculates the instruction-response expansion index score, which serves as the instruction-response expansion index for each piece of TCM-related data, thereby quantifying the difficulty of model generation divergence.

[0096] S280. Determine the isolation degree of each piece of traditional Chinese medicine data in the second traditional Chinese medicine dataset.

[0097] Specifically, the electronic device uses a preset machine learning library to extract the feature matrix of the sample set, and then uses a clustering algorithm to divide the data in the second traditional Chinese medicine dataset into multiple semantic clusters, and calculates the isolation degree of each traditional Chinese medicine data based on the semantic clusters.

[0098] S290. The instruction response expansion index and the isolation degree of each piece of TCM data are weighted and fused to obtain the external difficulty score of each piece of TCM data.

[0099] Specifically, the electronic device takes the average of the instruction response expansion index and the isolation degree of each piece of TCM data as the external difficulty score of each piece of TCM data.

[0100] S291. Sort the external difficulty scores of each TCM data in descending order, and remove the data with the simple structure in the bottom 50% of the external difficulty scores to obtain the TCM fine-tuning dataset.

[0101] Specifically, electronic devices are sorted in descending order of external difficulty score. The second TCM dataset is sorted by external difficulty, and 50% of the simpler data is filtered out again to obtain a core subset of about 5% of the original data volume.

[0102] S292. Use the TCM fine-tuning dataset as a vertical domain labeled fine-tuning dataset.

[0103] Specifically, electronic devices can transform the evaluation dimensions of the system into human annotation instruction specifications, including: Specification 1, requiring vertical domain annotation users to include interdisciplinary requirements when writing instructions (e.g., combining traditional Chinese medicine formulas with modern pharmacology characteristics for analysis); Specification 2, requiring detailed answers that involve the "evaluation and creation" level; Specification 3, under this guidance, only a very small number (e.g., 200) of high-difficulty domain instructions are built, that is, fine-tuning to generate a domain-wide model that is superior to the conventional training of tens of thousands of vertical domain question-and-answer data, thereby achieving cost reduction and efficiency improvement.

[0104] S293. Using the TCM fine-tuning dataset, supervised fine-tuning is performed on pre-trained models in general domains.

[0105] Specifically, the 5% of high-difficulty data output mentioned above is used as the training set to perform supervised fine-tuning on the large model.

[0106] In this way, the fine-tuned model will outperform the model trained using the original full dataset in generalization tests.

[0107] This disclosure also provides a cognitive-inspired efficient instruction fine-tuning device for a large-scale traditional Chinese medicine model, used to implement the above-mentioned efficient instruction fine-tuning method for a cognitive-inspired model of traditional Chinese medicine. The following is a detailed description... Figure 3 The following explanation is provided. In this embodiment, the efficient instruction fine-tuning device for the cognitively inspired large-scale model of traditional Chinese medicine can be an electronic device or a server. The electronic device can include a desktop computer, laptop, tablet, and other smart devices. The server can include a cloud server or a server cluster.

[0108] Figure 3 A schematic diagram of the structure of a cognitively inspired large-scale TCM model high-efficiency instruction fine-tuning device provided in this disclosure is shown.

[0109] like Figure 3 As shown, the cognitively inspired large-scale TCM model high-efficiency instruction fine-tuning device 300 may include: The first acquisition module 310 is used to perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset and acquire the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset. The second acquisition module 320 is used to acquire a second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset based on the inherent difficulty of the traditional Chinese medicine data representation in the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. The third acquisition module 330 is used to acquire a TCM fine-tuning dataset from the second TCM dataset based on the external difficulty of the TCM data in the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The fine-tuning module 340 is used to fine-tune the preset TCM question-and-answer model using the TCM fine-tuning dataset to obtain the fine-tuned model.

[0110] This disclosure discloses a cognitively inspired, high-efficiency instruction fine-tuning device for a large-scale TCM (Traditional Chinese Medicine) model. For a massive original TCM dataset, it first performs data quality analysis on the TCM-related data within the original dataset, obtaining a first TCM dataset. Then, based on the inherent difficulty of representing the TCM-related data in the first dataset, a second TCM dataset is obtained. Next, based on the external difficulty of representing the TCM-related data in the second dataset, a fine-tuning dataset is obtained. Finally, the fine-tuning dataset is used to fine-tune a pre-defined large-scale TCM question-and-answer model, resulting in a well-tuned model. Thus, by constructing a transparent and quantifiable method for determining fine-tuning data through multi-dimensional evaluation, it is beneficial to trace which TCM-related data is of high quality, thereby achieving the goal of quantifying and explaining the quality of the fine-tuning dataset, and ultimately facilitating accurate evaluation of the fine-tuning effect of the large-scale TCM question-and-answer model.

[0111] In some embodiments of this disclosure, the first acquisition module 310 includes: The first analysis unit is used to input the original TCM dataset into a preset quality analysis model, perform data quality analysis on the TCM data in the original TCM dataset, and obtain the quality score of each TCM data in the original TCM dataset. The first acquisition unit is used to acquire a first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset based on the quality scores of each type of traditional Chinese medicine data.

[0112] In some embodiments of this disclosure, the second acquisition module 320 includes: The second analysis unit is used to perform cognitive reasoning analysis on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the reasoning ability score of each traditional Chinese medicine data in the first traditional Chinese medicine dataset. The third analysis unit is used to perform cross-domain analysis on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the information complexity of each traditional Chinese medicine data in the first traditional Chinese medicine dataset. The first determining unit is used to determine the inherent difficulty score of the representation of each type of traditional Chinese medicine data based on the reasoning ability score of each type of traditional Chinese medicine data and the information complexity of each type of traditional Chinese medicine data. The second acquisition unit is used to acquire the second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset based on the inherent difficulty score represented by each type of traditional Chinese medicine data.

[0113] In some embodiments of this disclosure, the second analysis unit is specifically used for: Input the first TCM dataset into a preset classification model to classify and analyze the TCM data in the first TCM dataset, and obtain the cognitive level to which each TCM data belongs. Based on the progressive difficulty between the cognitive levels to which different TCM data belong, Bloom's cognitive calculation is performed on the TCM data in the first TCM dataset to obtain the reasoning ability score of each TCM data category.

[0114] In some embodiments of this disclosure, the third analysis unit is specifically used for: Using a pre-defined vector extraction model, the TCM data in the first TCM dataset is vectorized to obtain domain description features; The complexity of the domain description features is calculated to obtain the information complexity of each type of TCM data.

[0115] In some embodiments of this disclosure, the third acquisition module 330 includes: The fourth analysis unit is used to perform burden analysis on the traditional Chinese medicine data in the second traditional Chinese medicine dataset to obtain the burden score of each traditional Chinese medicine data in the second traditional Chinese medicine dataset. The fifth analysis unit is used to perform isolation degree analysis on the traditional Chinese medicine data in the second traditional Chinese medicine dataset and obtain the isolation degree score of each traditional Chinese medicine data in the second traditional Chinese medicine dataset. The second determining unit is used to determine the external difficulty score of each type of traditional Chinese medicine data based on the burden score of each type of traditional Chinese medicine data and the isolation degree of each type of traditional Chinese medicine data. The third acquisition unit is used to acquire the TCM fine-tuning dataset from the second TCM dataset based on the external difficulty scores represented by the TCM data of each category.

[0116] In some embodiments of this disclosure, the fourth analysis unit is specifically used for: Extract instruction data with instruction semantic information and response data with response semantic information from the traditional Chinese medicine data contained in the second traditional Chinese medicine dataset; The instruction data and the response data in the second traditional Chinese medicine dataset are merged to obtain the instruction-response merged data; Determine the first length corresponding to the instruction data, the second length corresponding to the response data, the minimum length of the instruction response merged data, and the maximum length of the instruction response merged data; Based on the first length, the second length, the minimum length, and the maximum length, calculate the instruction response expansion index for each category of traditional Chinese medicine data in the second traditional Chinese medicine dataset, which serves as the burden score for each category of traditional Chinese medicine data in the second traditional Chinese medicine dataset.

[0117] In some embodiments of this disclosure, the fifth analysis unit is specifically used for: Convert the TCM-related data in the second TCM dataset into vector representations; Cluster analysis is performed on the vector representation to obtain multiple clusters corresponding to the second traditional Chinese medicine dataset; For the current data in the second traditional Chinese medicine dataset, the nearest neighbor cluster of the current traditional Chinese medicine data and the cluster to which the current traditional Chinese medicine data belongs are obtained from the multiple clusters; Determine the first average distance between the current TCM data and each TCM data in the nearest neighbor cluster, and determine the second average distance between the current TCM data and each TCM data in the cluster to which it belongs; Based on the first average distance and the second average distance, the isolation degree of the TCM data in the second TCM dataset is calculated to obtain the isolation degree score of each TCM data.

[0118] In some embodiments of this disclosure, the fine-tuning module 340 is specifically used for: The TCM fine-tuning dataset is used as a labeled fine-tuning dataset in the field of TCM question-and-answer, and the preset TCM question-and-answer model is fine-tuned using the labeled fine-tuning dataset to obtain the fine-tuned model.

[0119] It should be noted that, Figure 3 The cognitively inspired large-scale TCM model high-efficiency instruction fine-tuning device 300 shown can execute... Figures 1-2 The various steps in the method embodiment shown are implemented. Figures 1-2 The processes and effects in the method embodiments shown are not described in detail here.

[0120] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0121] like Figure 4 As shown, the electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0122] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0123] Memory 402 may include a large-capacity storage device for advertising or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway device. In a particular embodiment, memory 402 is a non-volatile solid-state memory. In a particular embodiment, memory 402 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0124] The processor 401 acquires and executes computer program instructions stored in the memory 402 to perform the steps of the efficient instruction fine-tuning method for a large-scale model of traditional Chinese medicine based on cognitive inspiration provided in this embodiment of the present disclosure.

[0125] In one example, the electronic device may also include a transceiver 403 and a bus 404. Wherein, as... Figure 4 As shown, the processor 401, memory 402 and transceiver 403 are connected via bus 404 and communicate with each other.

[0126] Bus 404 includes hardware, software, or both. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 404 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0127] The following are embodiments of the computer-readable storage medium provided in this disclosure. This computer-readable storage medium belongs to the same inventive concept as the cognitive-inspired method for fine-tuning instructions of a large-scale traditional Chinese medicine model described above. For details not described in detail in the embodiments of the computer-readable storage medium, please refer to the embodiments of the cognitive-inspired method for fine-tuning instructions of a large-scale traditional Chinese medicine model described above.

[0128] This embodiment provides a storage medium containing computer-executable instructions. When executed by a computer processor, these computer-executable instructions are used to perform a cognitively inspired, high-efficiency instruction fine-tuning method for a large-scale model of traditional Chinese medicine, including: Perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset, and obtain the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset; Based on the inherent difficulty of representing traditional Chinese medicine data in the first traditional Chinese medicine dataset, a second traditional Chinese medicine dataset is obtained from the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. Based on the external difficulty of the TCM data in the second TCM dataset, a TCM fine-tuning dataset is obtained from the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The pre-set TCM question-and-answer model was fine-tuned using the TCM fine-tuning dataset to obtain a fine-tuned model.

[0129] Of course, the computer-executable instructions provided in the embodiments of this disclosure are not limited to the above-described method operations, but can also execute related operations in the cognitive-inspired large-scale TCM model efficient instruction fine-tuning method provided in any embodiment of this disclosure.

[0130] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this disclosure can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer cloud platform (which can be a personal computer, server, or network cloud platform, etc.) to execute the cognitive-inspired large-scale TCM model efficient instruction fine-tuning method provided in the various embodiments of this disclosure.

[0131] Note that the above description is merely a preferred embodiment and the technical principles employed in this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this disclosure. Therefore, although this disclosure has been described in detail through the above embodiments, it is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this disclosure, and the scope of this disclosure is determined by the scope of the appended claims.

Claims

1. A highly efficient instruction fine-tuning method for a large-scale TCM model based on cognitive inspiration, characterized in that, include: Perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset, and obtain the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset; Based on the inherent difficulty of representing traditional Chinese medicine data in the first traditional Chinese medicine dataset, a second traditional Chinese medicine dataset is obtained from the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. Based on the external difficulty of the TCM data in the second TCM dataset, a TCM fine-tuning dataset is obtained from the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The pre-set TCM question-and-answer model was fine-tuned using the TCM fine-tuning dataset to obtain a fine-tuned model.

2. The method according to claim 1, characterized in that, The step of performing data quality analysis on the traditional Chinese medicine (TCM) data in the original TCM dataset, and obtaining the first TCM dataset from the original TCM dataset, includes: The original TCM dataset is input into a preset quality analysis model to perform data quality analysis on the TCM data in the original TCM dataset and obtain the quality score of each TCM data in the original TCM dataset. Based on the quality scores of each type of TCM data, a first TCM dataset is obtained from the original TCM dataset.

3. The method according to claim 1, characterized in that, The process of obtaining a second traditional Chinese medicine (TCM) dataset from the first TCM dataset, based on the inherent difficulty of representing TCM-related data in the first TCM dataset, includes: Cognitive reasoning analysis is performed on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the reasoning ability score of each traditional Chinese medicine data in the first traditional Chinese medicine dataset; Perform cross-domain analysis on the traditional Chinese medicine data in the first traditional Chinese medicine dataset to obtain the information complexity of each traditional Chinese medicine data in the first traditional Chinese medicine dataset; Based on the reasoning ability score and information complexity of each type of TCM data, the inherent difficulty score of the representation of each type of TCM data is determined. Based on the inherent difficulty scores of each type of TCM data, the second TCM dataset is obtained from the first TCM dataset.

4. The method according to claim 3, characterized in that, The cognitive reasoning analysis of the TCM-related data in the first TCM dataset, to obtain the reasoning ability score of each TCM-related data in the first TCM dataset, includes: Input the first TCM dataset into a preset classification model to classify and analyze the TCM data in the first TCM dataset, and obtain the cognitive level to which each TCM data belongs. Based on the progressive difficulty between the cognitive levels to which different TCM data belong, Bloom's cognitive calculation is performed on the TCM data in the first TCM dataset to obtain the reasoning ability score of each TCM data category.

5. The method according to claim 3, characterized in that, The step of performing cross-domain analysis on the TCM-related data in the first TCM dataset to obtain the information complexity of each TCM-related data in the first TCM dataset includes: Using a pre-defined vector extraction model, the TCM data in the first TCM dataset is vectorized to obtain domain description features; The complexity of each data item in the traditional Chinese medicine category is obtained by performing complexity calculations on the domain description features.

6. The method according to claim 1, characterized in that, The step of obtaining a TCM fine-tuning dataset from the second TCM dataset based on the external difficulty represented by TCM data in the second TCM dataset includes: A burden analysis was performed on the traditional Chinese medicine data in the second traditional Chinese medicine dataset to obtain the burden score of each traditional Chinese medicine data category in the second traditional Chinese medicine dataset; An isolation degree analysis was performed on the TCM data in the second TCM dataset to obtain the isolation degree score of each TCM data in the second TCM dataset; Based on the burden score and isolation degree of each type of TCM data, the external difficulty score of each type of TCM data is determined. Based on the external difficulty scores represented by each type of TCM data, the TCM fine-tuning dataset is obtained from the second TCM dataset.

7. The method according to claim 6, characterized in that, The burden analysis of the TCM data in the second TCM dataset, to obtain the burden score of each TCM data category in the second TCM dataset, includes: Extract instruction data with instruction semantic information and response data with response semantic information from the traditional Chinese medicine data contained in the second traditional Chinese medicine dataset; The instruction data and the response data in the second traditional Chinese medicine dataset are merged to obtain the instruction-response merged data; Determine the first length corresponding to the instruction data, the second length corresponding to the response data, the minimum length of the instruction response merged data, and the maximum length of the instruction response merged data; Based on the first length, the second length, the minimum length, and the maximum length, calculate the instruction response expansion index for each category of traditional Chinese medicine data in the second traditional Chinese medicine dataset, which serves as the burden score for each category of traditional Chinese medicine data in the second traditional Chinese medicine dataset.

8. The method according to claim 6, characterized in that, The isolation degree analysis of the TCM data in the second TCM dataset, to obtain the isolation degree score of each TCM data category in the second TCM dataset, includes: Convert the traditional Chinese medicine data in the second traditional Chinese medicine dataset into vector representations; Cluster analysis is performed on the vector representation to obtain multiple clusters corresponding to the second traditional Chinese medicine dataset; For the current TCM category data in the second TCM dataset, the nearest neighbor cluster of the current TCM category data and the cluster to which the current TCM category data belongs are obtained from the multiple clusters; Determine the first average distance between the current TCM data and each TCM data in the nearest neighbor cluster, and determine the second average distance between the current TCM data and each TCM data in the cluster to which it belongs; Based on the first average distance and the second average distance, the isolation degree of the TCM data in the second TCM dataset is calculated to obtain the isolation degree score of each TCM data.

9. The method according to claim 1, characterized in that, The step of fine-tuning the pre-set TCM question-and-answer model using the TCM fine-tuning dataset to obtain the fine-tuned model includes: The TCM fine-tuning dataset is used as a labeled fine-tuning dataset in the field of TCM question-and-answer, and the preset TCM question-and-answer model is fine-tuned using the labeled fine-tuning dataset to obtain the fine-tuned model.

10. A highly efficient instruction fine-tuning device for a large-scale traditional Chinese medicine model based on cognitive inspiration, characterized in that, include: The first acquisition module is used to perform data quality analysis on the traditional Chinese medicine data in the original traditional Chinese medicine dataset and acquire the first traditional Chinese medicine dataset from the original traditional Chinese medicine dataset. The second acquisition module is used to acquire a second traditional Chinese medicine dataset from the first traditional Chinese medicine dataset based on the inherent difficulty of the data representation in the first traditional Chinese medicine dataset, wherein the inherent difficulty represents the content depth of each piece of traditional Chinese medicine data in the field of traditional Chinese medicine. The third acquisition module is used to acquire a TCM fine-tuning dataset from the second TCM dataset based on the external difficulty of the TCM data in the second TCM dataset, wherein the external difficulty represents the content breadth of each TCM data in the field of TCM. The fine-tuning module is used to fine-tune the preset TCM question-and-answer model using the TCM fine-tuning dataset to obtain a fine-tuned model.