Training Method of Model, Data Processing Method and Device, Equipment, Medium
By training the target matrix prediction model and generating the target parameter matrix, the problem of low accuracy of symptom entity data in intelligent decision-making of traditional Chinese medicine diseases is solved, and the effect of intelligent decision-making of traditional Chinese medicine diseases is improved.
Patent Information
- Application Number
- CN202210073114.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-01-21
AI Technical Summary
In the intelligent decision-making of traditional Chinese medicine diseases, it is difficult to effectively map user spoken symptom entity data and Western medicine diagnosis data to a standardized traditional Chinese medicine context, resulting in low accuracy of symptom entity data.
A model training method is proposed, by obtaining the traditional Chinese medicine disease and symptom type training set, traditional Chinese medicine diagnosis rule data, symptom entity normalization data, etc., data matching and extraction are carried out, target matrix prediction model is trained, and target parameter matrix is generated to improve the accuracy of symptom entity data.
By improving the accuracy of symptom entity data, ensuring the effectiveness of intelligent decision-making in traditional Chinese medicine diseases, the application ability of traditional Chinese medicine diagnostic rules is enhanced.
Smart Images

Figure CN114399001B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a method for training a model, a method and device for data processing, a device, and a medium. Background Art
[0002] With the development of artificial intelligence technology, intelligent decision-making for traditional Chinese medicine diseases has been widely applied in many fields. In related technologies, it is usually necessary to map the symptom entity data (such as in the form of free text) input by the user consultation terminal to standardized symptom entity data, such as a symptom entity list. However, the data sources of the collected symptom entity data usually include the user's colloquial symptom entity data and / or the symptom entity data corresponding to Western medicine diagnoses, which cannot ensure that they are collected in the context of traditional Chinese medicine, resulting in a large error in the standardized symptom entity data. Summary of the Invention
[0003] The main purpose of the embodiments of the present disclosure is to propose a method for training a model, a method and device for data processing, a device, and a medium, which can improve the accuracy of symptom entity data and ensure the effect of intelligent decision-making for traditional Chinese medicine diseases.
[0004] To achieve the above object, a first aspect of the embodiments of the present disclosure proposes a method for training a model for training a target matrix prediction model, and the training method includes:
[0005] Obtain a training set of traditional Chinese medicine diseases and syndromes, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data;
[0006] Perform data matching processing on the training set of traditional Chinese medicine diseases and syndromes, the traditional Chinese medicine diagnosis rule data, and the symptom entity normalization data to obtain seed symptom entity data;
[0007] According to the seed symptom entity data, perform data extraction processing on the first symptom entity data to obtain non-seed symptom entity data;
[0008] According to the seed symptom entity data, the non-seed symptom entity data, the training set of traditional Chinese medicine diseases and syndromes, the traditional Chinese medicine diagnosis rule data, the first symptom entity data, and a preset training rule, perform training processing on a preset prediction model to obtain the target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix.
[0009] A second aspect of the embodiments of the present disclosure proposes a data processing method applied to a target matrix prediction model trained by the model training method according to the first aspect embodiment. The target matrix prediction model is used to generate a target parameter matrix, and the data processing method includes:
[0010] Obtain the third symptom entity data to be processed and the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules;
[0011] According to the preset search function, screen out the first vector parameter data corresponding to the third symptom entity data from the target parameter matrix;
[0012] According to the search function, screen out the second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix;
[0013] Perform similarity processing on the first vector parameter data and the second vector parameter data according to the semantic similarity calculation method to obtain the semantic similarity data of the first vector parameter data corresponding to the second vector parameter data;
[0014] According to the semantic similarity data, perform screening processing on the first vector parameter data to obtain the fifth symptom entity data corresponding to the first vector parameter data, where the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data indicates non-conformity with the preset traditional Chinese medicine rules;
[0015] Map the fifth symptom entity data into the fourth symptom entity data.
[0016] The third aspect of the embodiments of the present disclosure proposes a training device for a model, which is used to train a target matrix prediction model. The training device includes:
[0017] The first acquisition module is used to acquire the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, the first symptom entity data, and the symptom entity normalization data;
[0018] The data matching module is used to perform data matching processing on the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, and the symptom entity normalization data to obtain the seed symptom entity data;
[0019] The data extraction module is used to perform data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data;
[0020] The model training module is used to train a preset prediction model according to the seed symptom entity data, the non-seed symptom entity data, the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, the first symptom entity data, and the preset training rules to obtain the target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix.
[0021] A fourth aspect of the embodiments of the present disclosure provides a data processing device, which is applied to a target matrix prediction model trained by the model training method described in the embodiments of the first aspect. The target matrix prediction model is used to generate a target parameter matrix. The data processing device includes:
[0022] A second acquisition module, configured to acquire third symptom entity data to be processed and fourth symptom entity data that conforms to preset traditional Chinese medicine rules;
[0023] A first search module, configured to screen and obtain first vector parameter data corresponding to the third symptom entity data from the target parameter matrix according to a preset search function;
[0024] A second search module, configured to screen and obtain second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix according to the search function;
[0025] A data calculation module, configured to perform a similarity process on the first vector parameter data and the second vector parameter data according to a semantic similarity calculation method to obtain semantic similarity data of the first vector parameter data corresponding to the second vector parameter data;
[0026] A data screening module, configured to perform a screening process on the first vector parameter data according to the semantic similarity data to obtain fifth symptom entity data corresponding to the first vector parameter data, where the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data indicates non - compliance with the preset traditional Chinese medicine rules;
[0027] A data mapping module, configured to map the fifth symptom entity data into the fourth symptom entity data.
[0028] A fifth aspect of the embodiments of the present disclosure provides a computer device, which includes a memory and a processor. Among them, a program is stored in the memory, and when the program is executed by the processor, the processor is configured to execute the method described in any one of the embodiments of the first aspect of the present application or the method described in any one of the embodiments of the second aspect of the present application.
[0029] A sixth aspect of the embodiments of the present disclosure provides a storage medium, which is a computer - readable storage medium. The storage medium stores computer - executable instructions, and the computer - executable instructions are used to cause a computer to execute the method described in any one of the embodiments of the first aspect of the present application or the method described in any one of the embodiments of the second aspect of the present application.
[0030] The training method, data processing method, device, equipment, and medium of the model proposed in the embodiments of the present disclosure obtain a traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data; perform data matching processing on the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, and symptom entity normalization data to obtain seed symptom entity data; perform data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data; train a preset prediction model according to the seed symptom entity data, non-seed symptom entity data, traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and a preset training rule to obtain a target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix. The embodiments of the present disclosure can improve the accuracy of symptom entity data and ensure the intelligent decision-making effect of traditional Chinese medicine diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flowchart of the training method of the model provided by the embodiment of the present disclosure;
[0032] Figure 2 It is a schematic diagram of the traditional Chinese medicine diagnosis rule data provided by the embodiment of the present disclosure;
[0033] Figure 3 is Figure 1 a flowchart of step S120 in
[0034] Figure 4 is Figure 1 a first flowchart of step S130 in
[0035] Figure 5 is Figure 4 a second flowchart of step S130 in
[0036] Figure 6 is Figure 1 a flowchart of step S140 in
[0037] Figure 7 It is a schematic diagram of the category symptom entity data provided by the embodiment of the present disclosure;
[0038] Figure 8 It is a flowchart of the data processing method provided by the embodiment of the present disclosure;
[0039] Figure 9 It is a schematic diagram of the data processing method provided by the embodiment of the present disclosure;
[0040] Figure 10 It is a block diagram of the module structure of the model training device provided by the embodiment of the present disclosure;
[0041] Figure 11Block diagram of the module structure of the data processing device provided by the embodiments of the present disclosure;
[0042] Figure 12 It is a schematic diagram of the hardware structure of the computer device provided by the embodiments of the present disclosure. Detailed implementation manners
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the present application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0044] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description and claims and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.
[0046] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0047] The block diagrams shown in the drawings are only functional entities, and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0048] The flowcharts shown in the drawings are only illustrative, and do not necessarily include all the content and operations / steps, nor do they have to be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0049] First, parse several terms involved in this application:
[0050] Artificial intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0051] Semantic similarity: Semantic similarity is an intermediate task and an essential intermediate level in most natural language processing tasks. It has extensive applications in natural language processing, such as word sense disambiguation, information retrieval, and machine translation, etc. Semantic similarity can be applied to query translation to disambiguate and translate query keywords using semantic similarity to improve translation quality; or it can be applied to query expansion to make the expanded content more relevant to the original query, so as to improve the recall rate and accuracy of retrieval.
[0052] Symptom standardization, such as conforming to preset traditional Chinese medicine rules, is the cornerstone of any clinical decision-making system. The symptom descriptions corresponding to traditional Chinese medicine are more diverse and more targeted, and their expression methods are very different from the symptom terms in Western medicine or the colloquial expressions of patients. For example: "Aversion to cold" and "fear of cold" are actually two symptom entity data with similar but different meanings in the concept of traditional Chinese medicine.
[0053] With the development of artificial intelligence technology, intelligent decision-making for traditional Chinese medicine (TCM) diseases has been widely applied in many fields. However, in the intelligent diagnosis system for TCM diseases, the diagnosis process implemented through the intelligent diagnosis system for TCM diseases usually involves first creating a set of diagnosis rules driven by a collection of typical TCM symptoms according to the TCM theoretical system, such as creating a set of preset TCM rules. To apply this set of TCM rules well, it is necessary to map the symptom entity data input from the user's interrogation terminal, such as symptom descriptions (in the form of free text), into a standardized symptom entity data, such as mapping it into a standardized symptom entity list. However, in the related art, the data sources of the collected standardized symptom entity list usually include the user's colloquial symptom entity data, such as the description of the disease condition, and / or the symptom entity data corresponding to Western medicine diagnosis, such as the description of Western medicine diagnosis. It cannot be guaranteed that the data is collected in the context of TCM, which easily leads to a large error in the standardized symptom entity list. For example, a large amount of symptom entity data relied on by a whole set of intelligent decision-making for TCM diseases cannot be mapped to the collected and sorted standardized symptom entity list, ultimately resulting in very limited efficacy that a whole set of intelligent decision-making for TCM diseases can exert.
[0054] Based on this, the embodiments of the present disclosure propose a training method for a model, a data processing method, a device, a device, and a medium, which can improve the accuracy of symptom entity data to ensure the effect of intelligent decision-making for TCM diseases.
[0055] The embodiments of the present disclosure provide a training method for a model, a data processing method, a device, a device, and a medium, which will be specifically described through the following embodiments. First, the training method for the model in the embodiments of the present disclosure will be described.
[0056] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0057] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0058] The training method of the model provided by the embodiments of the present disclosure relates to the field of artificial intelligence and also to the field of virtual reality. The training method of the model provided by the embodiments of the present disclosure can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the training method of the model, etc., but is not limited to the above forms.
[0059] The embodiments of the present disclosure can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0060] Referring to Figure 1 , according to the training method of the model in the first aspect embodiment of the present disclosure, for training a target matrix prediction model, the training method includes but is not limited to the steps of:
[0061] Step S110, obtaining a traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data;
[0062] Step S120, performing data matching processing on the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, and symptom entity normalization data to obtain seed symptom entity data;
[0063] Step S130, performing data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data;
[0064] Step S140: Train a preset prediction model according to the seed symptom entity data, non-seed symptom entity data, traditional Chinese medicine (TCM) disease and syndrome training set, TCM diagnosis rule data, first symptom entity data, and preset training rules to obtain a target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix.
[0065] It can be understood that the following definitions are made: T (Train Data, abbreviated as T) represents the TCM disease and syndrome training set, R (Rule, abbreviated as R) represents the TCM diagnosis rule data, E (Entities, abbreviated as E) represents the first symptom entity data, NE (Normalised Entities, abbreviated as NE) represents the symptom entity normalisation data, SE (Seed Entities, abbreviated as SE) represents the seed symptom entity data, NonSE (Non-seed Entities, abbreviated as NonSE) represents the non-seed symptom entity data, and SibSE (Sibling Seed Entity, abbreviated as SibSE) represents the category symptom entity data.
[0066] Obtain the TCM disease and syndrome training set T: In the TCM disease and syndrome training set T, each TCM disease and syndrome training sample data corresponds to a TCM disease and syndrome label and symptom description data. In some embodiments, the TCM disease and syndrome label can be represented as a combined name of a disease and a syndrome, such as "insomnia - syndrome of qi deficiency of heart and gallbladder"; in some embodiments, the symptom description data can be represented as the main complaint description and / or other symptom descriptions of a user, such as a patient, and the symptom description data can be stored in the form of free text.
[0067] Obtain the TCM diagnosis rule data R: It can be understood that in traditional Chinese medicine, there is a diagnosis theory based on the eight-principle differentiation. Each TCM disease and syndrome has its corresponding TCM diagnosis rule label, such as the eight-principle label, disease nature label, and disease location label. Therefore, in the embodiments of the present disclosure, under the classification of TCM diagnosis rule labels, such as each eight-principle label, disease nature label, and disease location label, there is corresponding TCM diagnosis rule data R for the TCM disease and syndrome.
[0068] In some embodiments, the TCM diagnosis rule data R can be stored in a tree-like data structure. Refer to Figure 2 The TCM diagnosis rule data shown in
[0069] Obtain the first symptom entity data E: It can be understood that the first symptom entity data E can be a sorted list of symptom entities. Standardized symptom entity data collected and sorted through external corpora (such as various channels including publicly available Internet datasets, crawler data, medical records, books, expert definitions, etc.), for example, stored as a list of symptom entities. This list of symptom entities includes various symptom entity data, among which the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules is a subset of it.
[0070] Obtain the labeled normalized data NE of symptom entities: The input is free text data including symptom description data, and the output is labeled data labeled by professional annotators. Extract the symptom description data from the free text data and establish a mapping relationship with the normalized symptom entity data to obtain the normalized data NE of symptom entities.
[0071] It can be understood that the above-mentioned traditional Chinese medicine disease and syndrome type training set T, traditional Chinese medicine diagnosis rule data R, first symptom entity data E, and normalized data NE of symptom entities can all be in the form of sets.
[0072] After obtaining the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and normalized data NE of symptom entities, perform data matching processing on the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, and normalized data NE of symptom entities to obtain seed symptom entity data. Then, based on the seed symptom entity data, perform data extraction processing on the first symptom entity data to obtain non-seed symptom entity data. And according to the seed symptom entity data, non-seed symptom entity data, traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and preset training rules, perform training processing on the preset prediction model to obtain the target matrix prediction model, where the target matrix prediction model is used to generate the target parameter matrix. Through training the target matrix prediction model to generate the target parameter matrix in the embodiments of the present disclosure, it can be understood that the target parameter matrix includes the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, normalized data NE of symptom entities, seed symptom entity data, non-seed symptom entity data, and the association relationships between the above data obtained according to the preset training rules. Thus, the embodiments of the present disclosure can improve the accuracy of symptom entity data according to the target parameter matrix to ensure the intelligent decision-making effect of traditional Chinese medicine diseases.
[0073] Refer to Figure 3 , in some embodiments, in step S120, performing data matching processing on the traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, and normalized data NE of symptom entities to obtain seed symptom entity data includes, but is not limited to, the following steps:
[0074] Step S121: Perform data matching processing on the symptom entity normalized data and the traditional Chinese medicine disease and syndrome training set to obtain the first matching data;
[0075] Step S122: Extract from the first matching data the second matching data included in the traditional Chinese medicine diagnosis rule data;
[0076] Step S123: Use the second matching data as the seed symptom entity data.
[0077] It can be understood that in step S121 of some embodiments, the symptom entity normalized sample data in the labeled symptom entity normalized data NE is obtained, and the symptom entity normalized sample data is the labeled data. For each symptom entity normalized sample data, data matching processing is performed with the traditional Chinese medicine disease and syndrome training sample data in the traditional Chinese medicine disease and syndrome training set T to obtain the first matching data.
[0078] In step S122 of some embodiments, for each first matching data, data matching processing is performed with the traditional Chinese medicine diagnosis rule data R to obtain the second matching data, that is, the second matching data indicates being included in the traditional Chinese medicine diagnosis rule data R.
[0079] In step S123 of some embodiments, the above-mentioned second matching data is used as the seed symptom entity data.
[0080] The seed symptom entity data of the embodiments of the present disclosure can be a seed symptom entity list.
[0081] Refer to Figure 4 , in some embodiments, in step S130, according to the seed symptom entity data, data extraction processing is performed on the first symptom entity data to obtain non-seed symptom entity data, including but not limited to the steps of:
[0082] Step S131: Determine candidate matching data from the first symptom entity data according to the seed symptom entity data;
[0083] Step S132: Perform data extraction processing on the candidate matching data to obtain non-seed symptom entity data.
[0084] It can be understood that for the symptom entity data that cannot be matched to the seed symptom entity data, the embodiments of the present disclosure perform data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data.
[0085] Specifically, candidate matching data is determined from the first symptom entity data, where the candidate matching data is the symptom entity data other than the seed symptom entity data, and data extraction processing is performed on the candidate matching data to obtain non-seed symptom entity data.
[0086] Reference Figure 5 Figure 5 , in some embodiments, in step S132, data extraction processing is performed on the candidate matching data to obtain non-seed symptom entity data, including at least one of the following:
[0087] Step S133, perform word segmentation extraction processing on the candidate matching data according to a preset word segmentation algorithm and word segmentation dictionary to obtain non-seed symptom entity data; or,
[0088] Step S134, perform data matching processing on the candidate matching data according to the symptom entity normalization data to obtain non-seed symptom entity data; or,
[0089] Step S135, perform classification extraction processing on the candidate matching data according to a preset entity recognition model to obtain non-seed symptom entity data.
[0090] It can be understood that non-seed symptom entity data is obtained through one or a combination of steps S133 to S135. The non-seed symptom entity data of the embodiments of the present disclosure may be a non-seed symptom entity list.
[0091] In step S133 of some embodiments, the word segmentation extraction processing may be:
[0092] Based on a preset word segmentation algorithm and word segmentation dictionary, perform data matching processing on the candidate matching data. The preset word segmentation algorithm may be: Hidden Markov Model or Conditional Random Field word segmentation algorithm. The word segmentation dictionary may be customized. For example, the current medical term library may be used as the word segmentation dictionary, or the annotation data annotated by professional annotators may be used as the word segmentation dictionary.
[0093] In some embodiments, there is a preset word segmentation dictionary D. The word segmentation dictionary D includes: a word segmentation entity set: {(E, tag)}, where tag represents the corresponding word tag. For example, "(abdominal distension and fullness, symptom)", where "abdominal distension and fullness" is a symptom. A preset word segmentation algorithm f. Input the candidate matching data, such as the symptom description t = "The patient feels abdominal distension and fullness...", then perform word segmentation processing on the candidate matching data to obtain the word segmentation result, which is represented as f(D, t) = "The patient feels abdominal distension and fullness...". If the word segmentation dictionary is not defined, the obtained word segmentation result may be: "The patient feels abdominal distension and fullness...". Therefore, according to the word segmentation dictionary D, the problem of word segmentation ambiguity can be solved, and the word segmentation accuracy probability in the field of traditional Chinese medicine can be improved. Since the tag of "abdominal distension and fullness" in this embodiment is "symptom", extraction processing can be performed therefrom to obtain candidate matching data with tag = "symptom".
[0094] In step S134 of some embodiments, the data matching process may be as follows: using the part-of-speech pattern corresponding to the normalized data NE of the labeled symptom entities, perform data matching on the candidate matching data to obtain non-seed symptom entity data.
[0095] In step S135 of some embodiments, a candidate method based on part-of-speech pattern: Assume a symptom set {E} (used to train the entity recognition model) labeled by professional annotators. For example, "a feeling of fullness and firmness below the heart" is a symptom, and the original training text corresponding to "a feeling of fullness and firmness below the heart" can be found: t = "The patient often feels a feeling of fullness and firmness below the heart after meals, accompanied by flatulence in the gastrointestinal tract...". Suppose there is a part-of-speech recognition algorithm F. We can find the part-of-speech recognition result corresponding to "a feeling of fullness and firmness below the heart" through F(t): heart(n), below(f), fullness and firmness(adj), where "heart" is a noun, "below" is a locative word, and "fullness and firmness" is an adjective. Thus, nfadj is a pattern (anatomical location + symptom description). We use this pattern to match all words with the same pattern in the unlabeled text as candidate symptom entities. For example, "a feeling of stuffiness and discomfort below the stomach" conforms to this pattern. However, this method may introduce unnecessary noise, and a round of manual intervention can be performed to remove incorrect matching results.
[0096] In some embodiments, as Figure 6 shown, preset training rules include at least one of the following:
[0097] Step S141, perform training on the seed symptom entity data and non-seed symptom entity data through a prediction model to obtain the traditional Chinese medicine disease and syndrome type labels corresponding to the seed symptom entity data and non-seed symptom entity data; or,
[0098] Step S142, perform training on the seed symptom entity data through a prediction model according to the traditional Chinese medicine diagnosis rule data to obtain the traditional Chinese medicine diagnosis rule labels corresponding to the seed symptom entity data, where the traditional Chinese medicine diagnosis rule data includes traditional Chinese medicine diagnosis rule labels, and the traditional Chinese medicine diagnosis rule labels include at least one of the eight principles labels, disease nature labels, or disease location labels; or,
[0099] Step S143, perform training on the seed symptom entity data and category symptom entity data through a prediction model according to the second symptom entity data to obtain the association data of the second symptom entity data corresponding to the seed symptom entity data and category symptom entity data, where the second symptom entity data represents the first symptom entity data belonging to the non-seed symptom entity data, and the category symptom entity data represents the traditional Chinese medicine diagnosis rule data having the same traditional Chinese medicine diagnosis rule label as the seed symptom entity data.
[0100] It can be understood that the acquisition of associated data of category symptom entity data: according to the seed symptom entity data and traditional Chinese medicine diagnosis rule data, obtain the traditional Chinese medicine diagnosis rule data with the same traditional Chinese medicine diagnosis rule label as the seed symptom entity data, and obtain the category symptom entity data. The traditional Chinese medicine diagnosis rule label is at least one of the eight-principle label, disease-nature label, or disease-location label. Through the embodiments of the present disclosure, the category symptom entity data at the same category hierarchy level of the traditional Chinese medicine diagnosis rule label as the seed symptom entity data is further obtained.
[0101] For example, referring to Figure 7 , a schematic diagram showing the seed symptom entity data and category symptom entity data of an embodiment. Figure 7 In which R represents the traditional Chinese medicine diagnosis rule data, heartache is the seed symptom entity data, palpitation, pericardial effusion, and enlarged cardiac silhouette are all category symptom entity data. The right side shows the traditional Chinese medicine diagnosis rule label and traditional Chinese medicine diagnosis rule data corresponding to heartache, that is, eight principles: deficiency and excess, disease-nature: blood stasis, disease-location: heart.
[0102] The preset training rules of the embodiments of the present disclosure include a first training rule, a second training rule, and a third training rule. Specifically, the preset training rules include, but are not limited to:
[0103] First training rule: Through the prediction model, perform training processing on the seed symptom entity data and non-seed symptom entity data to obtain the traditional Chinese medicine diseases and syndrome types labels corresponding to the seed symptom entity data and non-seed symptom entity data;
[0104] It can be understood that in the embodiments of the present disclosure, first obtain the first merged set of the seed symptom entity data SE and the non-seed symptom entity data NonSE, that is, the first merged set is: SE ∪ NonSE. Then, through the prediction model, predict the traditional Chinese medicine diseases and syndrome types labels corresponding to each seed symptom entity data SE and each non-seed symptom entity data NonSE in the first merged set, that is, obtain the traditional Chinese medicine diseases and syndrome types labels corresponding to the seed symptom entity data and non-seed symptom entity data, and further obtain the first association relationship between the seed symptom entity data and non-seed symptom entity data and the corresponding traditional Chinese medicine diseases and syndrome types labels.
[0105] For example, in some embodiments, when the seed symptom entity data SE = "abdominal distension" and the non-seed symptom entity data NonSE = "firm fullness below the heart", according to the first training rule, the traditional Chinese medicine diseases and syndrome types labels corresponding to the seed symptom entity data SE and the non-seed symptom entity data NonSE can be obtained, specifically: the traditional Chinese medicine diseases and syndrome types label = "gastric fullness - liver qi stagnation".
[0106] Second training rule: Through the prediction model, perform training processing on the seed symptom entity data according to the traditional Chinese medicine diagnosis rule data to obtain the traditional Chinese medicine diagnosis rule label corresponding to the seed symptom entity data;
[0107] It can be understood that in the embodiments of the present disclosure, a prediction model is used to predict the traditional Chinese medicine diagnosis rule labels in the traditional Chinese medicine diagnosis rule data R corresponding to the seed symptom entity data SE, that is, the traditional Chinese medicine diagnosis rule labels corresponding to the seed symptom entity data are obtained. Specifically, the traditional Chinese medicine diagnosis rule labels include the eight-principle label P, the disease nature label Prop, and the disease location label L. Through the embodiments of the present disclosure, a second association relationship between the seed symptom entity data SE and the traditional Chinese medicine diagnosis rule labels such as the eight-principle label P, the disease nature label Prop, and the disease location label L under the traditional Chinese medicine diagnosis rule is further obtained.
[0108] For example, in some embodiments, when the seed symptom entity data SE = "abdominal distension", according to the second training rule, the traditional Chinese medicine diagnosis rule labels corresponding to the seed symptom entity data SE can be obtained, specifically: the disease location label L = "spleen".
[0109] The third training rule: The prediction model performs training processing on the seed symptom entity data and the category symptom entity data according to the second symptom entity data to obtain the association data of the second symptom entity data corresponding to the seed symptom entity data and the category symptom entity data;
[0110] It can be understood that in the embodiments of the present disclosure, first, a second merged set of the seed symptom entity data SE and the category symptom entity data SibSE is obtained, that is, the second merged set is: SE ∪ SibSE. Then, the second symptom entity data is obtained, that is, the first symptom entity data E belonging to the non-seed symptom entity data NonSE is used as the second symptom entity data. Specifically, the second symptom entity data is expressed as: E ∈ NonSE; the prediction model predicts the association data of the second symptom entity data corresponding to the second merged set, that is, the association data of the second symptom entity data corresponding to the seed symptom entity data and the category symptom entity data is obtained. Through the embodiments of the present disclosure, a third association relationship between the seed symptom entity data SE and the category symptom entity data SibSE is further obtained;
[0111] It can be understood that in some embodiments, since there may be multiple category symptom entity data SibSE corresponding to a seed symptom entity data SE, for example, more than two, therefore, through the third training rule, when one of the category symptom entity data SibSE is successfully predicted, it can be indicated that the prediction is correct.
[0112] For example, in some embodiments, when the non-seed symptom entity data NonSE = "a feeling of fullness and firmness below the heart", the seed symptom entity data SE = "abdominal distension", and the category symptom entity data SibSE = "stuffy and uncomfortable feeling in the epigastric and abdominal regions", according to the third training rule, the associated data of the second symptom entity data corresponding to the seed symptom entity data and the category symptom entity data SibSE can be obtained. Specifically, "abdominal distension" of the seed symptom entity data SE and "stuffy and uncomfortable feeling in the epigastric and abdominal regions" of the category symptom entity data SibSE both belong to the level of the disease nature label Prop corresponding to "qi stagnation". Therefore, the seed symptom entity data SE and the category symptom entity data SibSE have a third association relationship. It can be understood that the category symptom entity data SibSE can be expressed as sibling symptom entity data.
[0113] It can be understood that for a set of standardized symptom entity data that has been collected and sorted out (such as the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules in the embodiments of the present disclosure) and the symptom entity data to be processed, it is also necessary to effectively mine the symptom entity data that originally does not conform to the preset traditional Chinese medicine rules, that is, does not belong to the definitions of traditional Chinese medicine theory and diagnostic rules, but has a high degree of correlation with the preset traditional Chinese medicine rules and has a high diagnostic value in actual applications. Furthermore, during the use of the online system, even if the symptom description input by the user does not conform to the preset traditional Chinese medicine rules (such as the third symptom entity data to be processed in the embodiments of the present disclosure), the mapping relationship between the third symptom entity data and the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules can be established through the data processing method of the embodiments of the present disclosure, so that the traditional Chinese medicine diagnostic rules can fully play their roles.
[0114] It can be understood that the above preset training rules are carried out synchronously during the training process of the prediction model to share the parameter matrix of the symptom entity data. Let the number of the original data, namely the seed symptom entity data, the non-seed symptom entity data, the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnostic rule data, and the first symptom entity data, be N, and the vector dimension corresponding to the target parameter matrix generated by the target matrix prediction model be M (which can be set or adjusted according to experience), then the target parameter matrix is represented as Param ∈ R N×M 。
[0115] It can be understood that according to the training parameter matrix, the prediction model can be iteratively trained until the target matrix prediction model is obtained, and the target parameter matrix is finally generated according to the target matrix prediction model.
[0116] For example, through the target matrix prediction model, the target parameter matrix Param ∈ R can be finally output N×M 。It can be understood that for an input symptom entity data E ∈ R 1×N , its vector representation function, that is, the lookup function, can be defined as lookup = R 1×N ×RN×M →R 1×M Similarly, for the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules, the lookup function such as the lookup function can be used to obtain an M-dimensional vector. Further, using the semantic similarity calculation method (such as cosine similarity) and sorting according to the similarity from high to low, the fifth symptom entity data corresponding to the first vector parameter data can be obtained, that is, the fifth symptom entity data can represent the mapping relationship with the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules.
[0117] For this mapping relationship, some of the fifth symptom entity data can be extracted and verified by experts in a round of marking to generate a new labeled data set, where the label of each sample is whether the symptom entity can be normalized to the symptom entity data that conforms to the preset traditional Chinese medicine rules. Thus, a supervised binary classification model is established.
[0118] Based on the binary classification model, the seed symptom entity data in the original data is further supplemented. For example, the original non-seed symptom entity data may be reclassified as seed symptom entity data after passing through the binary classification model. After supplementing the seed symptom entity data, repeat step S140 until the target matrix prediction model is obtained to finally generate the target parameter matrix Param ∈ R N×M In this way, iterative training is performed to update the training parameter matrix.
[0119] Refer to Figure 8 According to the data processing method of the second aspect of the embodiments of the present disclosure, applied to the target matrix prediction model trained by the model training method in the above embodiments, the target matrix prediction model is used to generate the target parameter matrix. The data processing method includes but is not limited to the steps:
[0120] Step S200, obtaining the third symptom entity data to be processed and the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules;
[0121] Step S210, according to the preset lookup function, screening the first vector parameter data corresponding to the third symptom entity data from the target parameter matrix;
[0122] Step S220, according to the lookup function, screening the second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix;
[0123] Step S230, performing similarity processing on the first vector parameter data and the second vector parameter data according to the semantic similarity calculation method to obtain the semantic similarity data of the first vector parameter data corresponding to the second vector parameter data;
[0124] Step S240: Screen the first vector parameter data according to the semantic similarity data to obtain fifth symptom entity data corresponding to the first vector parameter data. Among them, the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data represents non-compliance with the preset traditional Chinese medicine rules.
[0125] Step S250: Map the fifth symptom entity data into the fourth symptom entity data.
[0126] It can be understood that the target parameter matrix is represented as Param k , obtain the to-be-processed third symptom entity data and the fourth symptom entity data that comply with the preset traditional Chinese medicine rules. According to the target parameter matrix Param k , use a preset lookup function such as the lookup function to screen out the first vector parameter data corresponding to the third symptom entity data from the target parameter matrix, and screen out the second vector parameter data corresponding to the fourth symptom entity data; use a semantic similarity calculation method to perform similarity processing on the first vector parameter data and the second vector parameter data to obtain the semantic similarity data of the first vector parameter data corresponding to the second vector parameter data.
[0127] Perform similarity sorting on the semantic similarity data. For example, sort the semantic similarity data from high to low, or sort the semantic similarity data from low to high. This embodiment does not make specific limitations on this.
[0128] According to the preset threshold and the semantic similarity data, screen the first vector parameter data to obtain fifth symptom entity data corresponding to the first vector parameter data. For example, obtain the top N fifth symptom entity data after sorting, and map the fifth symptom entity data into the fourth symptom entity data.
[0129] It can be understood that an ultimate target parameter matrix Param can be output according to the target matrix prediction model in the embodiments of the present disclosure. k . According to the to-be-processed third symptom entity data, the fourth symptom entity data that comply with the preset traditional Chinese medicine rules, and the target parameter matrix Param k , use a lookup function such as the lookup function to obtain the corresponding vectorized representation, such as the first vector parameter data and the second vector parameter data.
[0130] It can be understood that the fifth symptom entity data represents non-compliance with the preset traditional Chinese medicine rules.
[0131] For example, refer to Figure 9As shown, it is a schematic diagram of the data processing method according to an embodiment of the present disclosure. By closely combining with preset traditional Chinese medicine rules, the embodiment of the present disclosure can more effectively enable the target matrix prediction model to learn the association relationships between symptom entity data and traditional Chinese medicine diseases and syndromes, between symptom entity data and traditional Chinese medicine diagnosis rules, and between symptom entity data and symptom entity data. The embodiment of the present disclosure also uses the candidate mapping relationships output by weak supervision semantic representation, combines expert annotations for supervised classification learning, and repeatedly performs weak supervision semantic representation learning, so that the semantic representation ability of symptom entity data is gradually enhanced, which not only reflects the diagnosis theory of traditional Chinese medicine but also reflects the marking thinking of traditional Chinese medicine experts.
[0132] The embodiment of the present disclosure also provides a model training device for training a target matrix prediction model. As Figure 10 shown, it can implement the above model training method. The device includes: a first acquisition module 310, a data matching module 320, a data extraction module 330, and a model training module 340.
[0133] Among them, the first acquisition module 310 is used to acquire a traditional Chinese medicine disease and syndrome training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data; the data matching module 320 is used to perform data matching processing on the traditional Chinese medicine disease and syndrome training set, traditional Chinese medicine diagnosis rule data, and symptom entity normalization data to obtain seed symptom entity data; the data extraction module 330 is used to perform data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data; the model training module 340 is used to train a preset prediction model according to the seed symptom entity data, non-seed symptom entity data, traditional Chinese medicine disease and syndrome training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and a preset training rule to obtain a target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix.
[0134] The model training device of the embodiment of the present disclosure is used to execute the model training method in the above embodiment, and its specific processing process is the same as that of the model training method in the above embodiment, and will not be elaborated here one by one.
[0135] As Figure 11 shown, the embodiment of the present disclosure also provides a data processing device applied to a target matrix prediction model trained by the model training method in the above embodiment. The target matrix prediction model is used to generate a target parameter matrix and can implement the above data processing method. The data processing device includes: a second acquisition module 410, a first search module 420, a second search module 430, a data calculation module 440, a data screening module 450, and a data mapping module 460.
[0136] Among them, the second acquisition module 410 is used to acquire the third symptom entity data to be processed and the fourth symptom entity data that conforms to the preset traditional Chinese medicine rules; the first search module 420 is used to screen and obtain the first vector parameter data corresponding to the third symptom entity data from the target parameter matrix according to the preset search function; the second search module 430 is used to screen and obtain the second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix according to the search function; the data calculation module 440 is used to perform similarity processing on the first vector parameter data and the second vector parameter data according to the semantic similarity calculation method to obtain the semantic similarity data of the first vector parameter data corresponding to the second vector parameter data; the data screening module 450 is used to perform screening processing on the first vector parameter data according to the semantic similarity data to obtain the fifth symptom entity data corresponding to the first vector parameter data, where the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data represents non-conformity to the preset traditional Chinese medicine rules; the data mapping module 460 is used to map the fifth symptom entity data into the fourth symptom entity data.
[0137] The data processing device of the embodiments of the present disclosure is used to execute the data processing method in the above embodiments, and its specific processing process is the same as that of the data processing method in the above embodiments, and will not be elaborated here one by one.
[0138] The embodiments of the present disclosure also provide a computer device, including:
[0139] At least one processor, and,
[0140] A memory communicatively connected to at least one processor; wherein,
[0141] The memory stores instructions, and the instructions are executed by at least one processor so that when the at least one processor executes the instructions, the methods in any one of the embodiments of the first aspect or the embodiments of the second aspect of the present application are implemented.
[0142] Next, in conjunction with Figure 12 The hardware structure of the computer device will be described in detail. The computer device includes: a processor 510, a memory 520, an input / output interface 530, a communication interface 540, and a bus 550.
[0143] The processor 510 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure;
[0144] The memory 520 can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), etc. The memory 520 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 520 and are called by the processor 510 to execute the model training method of the embodiments of the present disclosure or to execute the data processing method of the embodiments of the present disclosure;
[0145] The input / output interface 530 is used to implement information input and output;
[0146] The communication interface 540 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); and
[0147] The bus 550 transmits information between the various components of the device (such as the processor 510, the memory 520, the input / output interface 530, and the communication interface 540);
[0148] Among them, the processor 510, the memory 520, the input / output interface 530, and the communication interface 540 are communicatively connected to each other inside the device through the bus 550.
[0149] The embodiments of the present disclosure further provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the model training method of the embodiments of the present disclosure or to execute the data processing method of the embodiments of the present disclosure.
[0150] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0151] The embodiments described in the embodiments of the present disclosure are to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0152] Those skilled in the art can understand that Figure 1 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 8 the technical solutions shown in do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0154] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware and their appropriate combinations.
[0155] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0156] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0157] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0158] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0160] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0161] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the present disclosure shall be within the scope of rights of the present disclosure.
Claims
1. A training method for a model, characterized in that, used for training a target matrix prediction model, the training method comprising: obtaining a traditional Chinese medicine disease and syndrome type training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data; wherein, the traditional Chinese medicine disease and syndrome type training set contains traditional Chinese medicine disease and syndrome type training sample data, and the traditional Chinese medicine disease and syndrome type training sample data contains traditional Chinese medicine disease and syndrome type labels and symptom description data; performing data matching processing on the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, and the symptom entity normalization data to obtain seed symptom entity data; performing data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data; performing training processing on a preset prediction model according to the seed symptom entity data, the non-seed symptom entity data, the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, the first symptom entity data, and a preset training rule to obtain the target matrix prediction model, wherein the target matrix prediction model is used to generate a target parameter matrix; wherein, the preset training rule includes at least one of the following: performing training processing on the seed symptom entity data and the non-seed symptom entity data through the prediction model to obtain traditional Chinese medicine disease and syndrome type labels corresponding to the seed symptom entity data and the non-seed symptom entity data; or, performing training processing on the seed symptom entity data through the prediction model according to the traditional Chinese medicine diagnosis rule data to obtain traditional Chinese medicine diagnosis rule labels corresponding to the seed symptom entity data, wherein the traditional Chinese medicine diagnosis rule data includes the traditional Chinese medicine diagnosis rule labels, and the traditional Chinese medicine diagnosis rule labels include at least one of eight-principle labels, disease nature labels, or disease location labels; or, performing training processing on the seed symptom entity data and category symptom entity data through the prediction model according to second symptom entity data to obtain association data of the second symptom entity data corresponding to the seed symptom entity data and the category symptom entity data, wherein the second symptom entity data represents the first symptom entity data belonging to the non-seed symptom entity data, and the category symptom entity data represents traditional Chinese medicine diagnosis rule data having the same traditional Chinese medicine diagnosis rule label as the seed symptom entity data.
2. The model training method according to claim 1, characterized in that, the performing data matching processing on the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, and the symptom entity normalization data to obtain seed symptom entity data includes: performing data matching processing on the symptom entity normalization data and the traditional Chinese medicine disease and syndrome type training set to obtain first matching data; extracting second matching data included in the traditional Chinese medicine diagnosis rule data from the first matching data; using the second matching data as the seed symptom entity data.
3. The model training method according to claim 1, characterized in that, Performing data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data, including: Determining candidate matching data from the first symptom entity data according to the seed symptom entity data; Performing data extraction processing on the candidate matching data to obtain non-seed symptom entity data.
4. The training method of the model according to claim 3, characterized in that the performing data extraction processing on the candidate matching data to obtain non-seed symptom entity data includes at least one of the following: Performing word segmentation extraction processing on the candidate matching data according to a preset word segmentation algorithm and word segmentation dictionary to obtain the non-seed symptom entity data; Or, Performing data matching processing on the candidate matching data according to the symptom entity normalization data to obtain the non-seed symptom entity data; Or, Performing classification extraction processing on the candidate matching data according to a preset entity recognition model to obtain the non-seed symptom entity data.
5. A data processing method, characterized in that applied to a target matrix prediction model trained by the training method of the model according to any one of claims 1 to 4, the target matrix prediction model is used to generate a target parameter matrix, and the data processing method includes: Obtaining third symptom entity data to be processed and fourth symptom entity data that conform to preset traditional Chinese medicine rules; Screening and obtaining first vector parameter data corresponding to the third symptom entity data from the target parameter matrix according to a preset search function; Screening and obtaining second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix according to the search function; Performing similarity processing on the first vector parameter data and the second vector parameter data according to a semantic similarity calculation method to obtain semantic similarity data of the first vector parameter data corresponding to the second vector parameter data; Performing screening processing on the first vector parameter data according to the semantic similarity data to obtain fifth symptom entity data corresponding to the first vector parameter data, where the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data indicates non-conformity with the preset traditional Chinese medicine rules; Mapping the fifth symptom entity data into the fourth symptom entity data.
6. A training device for a model, characterized in that used to train a target matrix prediction model, and the training device includes: A first acquisition module, configured to acquire a traditional Chinese medicine disease and syndrome training set, traditional Chinese medicine diagnosis rule data, first symptom entity data, and symptom entity normalization data; wherein, the traditional Chinese medicine disease and syndrome training set includes traditional Chinese medicine disease and syndrome training sample data, and the traditional Chinese medicine disease and syndrome training sample data includes traditional Chinese medicine disease and syndrome labels and symptom description data; A data matching module, configured to perform data matching processing on the traditional Chinese medicine disease and syndrome training set, the traditional Chinese medicine diagnosis rule data, and the symptom entity normalization data to obtain seed symptom entity data; A data extraction module, configured to perform data extraction processing on the first symptom entity data according to the seed symptom entity data to obtain non-seed symptom entity data; A model training module, configured to perform training processing on a preset prediction model according to the seed symptom entity data, the non-seed symptom entity data, the traditional Chinese medicine disease and syndrome type training set, the traditional Chinese medicine diagnosis rule data, the first symptom entity data, and a preset training rule to obtain the target matrix prediction model, where the target matrix prediction model is used to generate a target parameter matrix; Wherein, the preset training rule includes at least one of the following: Performing training processing on the seed symptom entity data and the non-seed symptom entity data through the prediction model to obtain traditional Chinese medicine disease and syndrome type labels corresponding to the seed symptom entity data and the non-seed symptom entity data; Or, Performing training processing on the seed symptom entity data through the prediction model according to the traditional Chinese medicine diagnosis rule data to obtain a traditional Chinese medicine diagnosis rule label corresponding to the seed symptom entity data, where the traditional Chinese medicine diagnosis rule data includes the traditional Chinese medicine diagnosis rule label, and the traditional Chinese medicine diagnosis rule label includes at least one of an eight-principle label, a disease nature label, or a disease location label; Or, Performing training processing on the seed symptom entity data and the category symptom entity data through the prediction model according to the second symptom entity data to obtain association data corresponding to the second symptom entity data for the seed symptom entity data and the category symptom entity data, where the second symptom entity data represents the first symptom entity data belonging to the non-seed symptom entity data, and the category symptom entity data represents traditional Chinese medicine diagnosis rule data having the same traditional Chinese medicine diagnosis rule label as the seed symptom entity data.
7. A data processing device Characterized in that Applied to a target matrix prediction model trained by the model training method according to any one of claims 1 to 4, the target matrix prediction model is used to generate a target parameter matrix, and the data processing device includes: A second acquisition module, configured to acquire third symptom entity data to be processed and fourth symptom entity data that conform to a preset traditional Chinese medicine rule; A first search module, configured to screen and obtain first vector parameter data corresponding to the third symptom entity data from the target parameter matrix according to a preset search function; A second search module, configured to screen and obtain second vector parameter data corresponding to the fourth symptom entity data from the target parameter matrix according to the search function; A data calculation module, configured to perform similarity processing on the first vector parameter data and the second vector parameter data according to a semantic similarity calculation method to obtain semantic similarity data corresponding to the first vector parameter data for the second vector parameter data; A data screening module, configured to screen the first vector parameter data according to the semantic similarity data, so as to obtain fifth symptom entity data corresponding to the first vector parameter data, wherein the third symptom entity data includes the fifth symptom entity data, and the fifth symptom entity data represents non-compliance with the preset traditional Chinese medicine rules; A data mapping module, configured to map the fifth symptom entity data into the fourth symptom entity data.
8. A computer device, characterized in that, the computer device includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is configured to execute: the training method of the model according to any one of claims 1 to 4; or the data processing method according to claim 5.
9. A storage medium, which is a computer-readable storage medium, characterized in that, the computer-readable storage stores a computer program, and when the computer program is executed by a computer, the computer is configured to execute: the training method of the model according to any one of claims 1 to 4; or the data processing method according to claim 5.
Citation Information
Patent Citations
Traditional Chinese medicine standard symptom matching method and device
CN112331338A
Medical term automatic standardization system and method integrating self-supervision and active learning
CN113436698A