Training Sample Processing Method, Device, Storage Medium, and Computer Equipment

By using keyword matching, invalid data cleaning, feature clustering and free data cleaning in electronic medical records in the field of intelligent medical records, the problem of insufficient quality of electronic medical records is solved, and the accuracy and acquisition efficiency of training data are improved.

CN114255834BActive Publication Date: 2025-06-27BEIJING HUIJI ZHIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111346891.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-06-27
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

The quality of electronic medical records in the existing intelligent medical field is insufficient, resulting in low accuracy of training data. The existing rule matching methods cannot effectively handle the description of synonyms, synonyms and complex medical texts.

Method used

By obtaining the keywords of the target training task and the sample data in the initial sample set, invalid data cleaning process is performed, feature clustering process, and through free data cleaning process, the target training samples are selected to improve the accuracy of the training data.

Benefits of technology

On the basis of ensuring the accuracy of training samples, the efficiency of training samples acquisition is greatly improved and the accuracy and diversity of training data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255834B_ABST
    Figure CN114255834B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, storage medium and computer device for processing training samples. The method includes: performing invalid data cleaning processing on the sample data in the initial sample set according to the keywords of the target training task and the matching result between the sample data content and the keywords to obtain a candidate sample set; performing feature clustering processing on the sample data in the candidate sample set according to the data features of the sample data in the candidate sample set to obtain a sample data clustering result of the candidate sample set; performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and the preset outlier data determination condition to obtain candidate sample data clusters; selecting target training samples for the target training task from the candidate sample data clusters and outputting the target training samples. The present application greatly improves the efficiency of obtaining training samples while ensuring a certain accuracy rate of the training samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent medical technology, and particularly relates to a method and device for processing training samples, a computer-readable storage medium, and a computer device. Background Art

[0002] Currently, in the field of intelligent medical care, the overall quality of electronic medical records is insufficient, resulting in poor quality of large databases based on real medical data. The data in these databases cannot be directly used as training data for the training and testing of artificial intelligence systems, such as auxiliary diagnosis systems, medical record content rationality inspection systems, etc.

[0003] The method based on rule matching can select some training data from the database. For example, some electronic medical records with format errors can be excluded using simple field matching. For example, electronic medical records with a certain field (such as the current medical history) being an empty character, a placeholder character, or having too few words are deleted; or some features corresponding to diseases can be obtained using data mining methods. For example, the common symptoms of coronary atherosclerosis include chest tightness, chest pain, heartburn, etc. Appropriate electronic medical records can be selected based on these features. However, using field matching can only exclude some format errors in electronic medical records, and the method of data mining cannot handle synonyms, near-synonyms, and descriptions from another perspective, etc. Therefore, these rule matching methods have great limitations and cannot adapt to the complex characteristics of medical texts, resulting in low accuracy of training data.

[0004] To ensure the accuracy of training data, the method of manual indexing is often adopted. Professionals with professional knowledge check the entire database or part of the database, and select a part of representative and reasonable samples according to professional knowledge and personal experience. However, the method of manual indexing is time-consuming and laborious, greatly reducing the efficiency of obtaining training data. Summary of the Invention

[0005] Embodiments of this application provide a method and device for processing training samples, a computer-readable storage medium, and a computer device, which can greatly improve the efficiency of obtaining training samples on the basis of ensuring a certain accuracy rate of the training samples.

[0006] Embodiments of this application provide a method for processing training samples, including:

[0007] Obtain keywords of a target training task and sample data in an initial sample set corresponding to the target training task;

[0008] According to the matching result between the sample data content of the sample data and the keywords, perform invalid data cleaning processing on the sample data in the initial sample set to obtain a candidate sample set for the target training task;

[0009] Perform feature clustering processing on the sample data in the candidate sample set according to the data characteristics of the sample data in the candidate sample set, and obtain the sample data clustering result of the candidate sample set;

[0010] According to the sample data clustering result and the preset outlier data determination condition, perform outlier data cleaning processing on each sample data cluster in the sample data clustering result to obtain candidate sample data clusters;

[0011] Select the target training samples of the target training task from the candidate sample data clusters;

[0012] Output the target training samples of the target training task.

[0013] The embodiment of the present application further provides a training sample processing device, including:

[0014] An acquisition module, configured to acquire keywords of a target training task and sample data in an initial sample set corresponding to the target training task;

[0015] A first cleaning module, configured to perform invalid data cleaning processing on the sample data in the initial sample set according to the matching result between the sample data content of the sample data and the keywords, and obtain a candidate sample set of the target training task;

[0016] A clustering module, configured to perform feature clustering processing on the sample data in the candidate sample set according to the data characteristics of the sample data in the candidate sample set, and obtain the sample data clustering result of the candidate sample set;

[0017] A second cleaning module, configured to perform outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and the preset outlier data determination condition, and obtain candidate sample data clusters;

[0018] A selection module, configured to select the target training samples of the target training task from the candidate sample data clusters;

[0019] An output module, configured to output the target training samples of the target training task.

[0020] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the training sample processing method described in any one of the above embodiments.

[0021] The embodiments of the present application further provide a computer device, which includes a memory and a processor. A computer program is stored in the memory, and the processor executes the steps in the training sample processing method described in any of the foregoing embodiments by calling the computer program stored in the memory.

[0022] For the training sample processing method, device, computer-readable storage medium, and computer device provided by the embodiments of the present application, invalid data cleaning processing is performed on the sample data in the initial sample set through the keywords of the target training task and the matching result between the sample data content and the keywords to obtain a candidate sample set. Feature clustering processing is performed on the sample data in the candidate sample set according to the data characteristics of the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set. According to the clustering method, concentrated features in the candidate sample set can be found to improve the accuracy of the training samples. Then, according to the sample data clustering result and the preset free data determination condition, free data cleaning processing is performed on each sample data cluster in the sample data clustering result to obtain candidate sample data clusters, and the free data in each sample data cluster is cleaned to remove the free data with large differences from most of the sample data in the sample data cluster, improving the accuracy of the training samples. Finally, according to the preset training sample screening condition, target sample data is selected from the candidate sample data clusters as the training samples for the target training task, and the sample data in the candidate sample data clusters is selected again to improve the accuracy of the training samples. The embodiments of the present application automatically realize the acquisition of training samples. Compared with the manual indexing method, the efficiency of training sample acquisition is greatly improved on the basis of ensuring a certain accuracy rate of the training samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic flowchart of the training sample processing method provided by the embodiments of the present application.

[0025] Figure 2 It is a schematic diagram of symptom entity matching provided by the embodiments of the present application.

[0026] Figure 3 It is another schematic flowchart of the training sample processing method provided by the embodiments of the present application.

[0027] Figure 4a 、 Figure 4b 、 Figure 4c 、 Figure 4dIt is a schematic diagram of the clustering effect of the sample data provided by the embodiment of the present application.

[0028] Figure 5a 、 Figure 5b It is a schematic diagram of the verification result provided by the embodiment of the present application.

[0029] Figure 6 It is a schematic structural diagram of the training sample processing device provided by the embodiment of the present application.

[0030] Figure 7 It is a schematic structural diagram of the computer device provided by the embodiment of the present application. Specific embodiments

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0032] The embodiment of the present application provides a training sample processing method, device, computer-readable storage medium, and computer device. Specifically, the training sample processing method in the embodiment of the present application can be executed by a computer device, where the computer device can be a terminal or a server and other devices. The terminal can be a smart phone, a tablet computer, a notebook computer, a touch screen, a game console, a personal computer (PC), and other terminal devices. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services and cloud databases.

[0033] The following will separately elaborate on a training sample processing method, device, computer-readable storage medium, and computer device provided by the embodiment of the present application. It should be noted that the serial numbers of the following embodiments are not used to limit the preferred order of the embodiments.

[0034] At the same time, it should be noted that the training sample processing method in the embodiment of the present application can be applied to the acquisition of any kind of training sample. For the sake of easy understanding, the embodiment of the present application takes the target training task as the training task of a certain type of disease in the field of intelligent medicine as an example for illustration, such as coronary atherosclerosis, inguinal hernia, etc. Correspondingly, the initial sample set includes multiple electronic medical records related to this type of disease.

[0035] As Figure 1 shown, it is a schematic flowchart of the training sample processing method provided by the embodiment of the present application, and the training sample processing method includes the following steps.

[0036] 101. Obtain the keywords of the target training task and the sample data in the initial sample set corresponding to the target training task.

[0037] Among them, the target training task is a training task in any usage scenario, such as a training task for a certain type of disease in the field of intelligent healthcare, a training task for a certain type of virtual message recognition, and so on.

[0038] For a training task of a certain type of disease in the field of intelligent healthcare, the sample data in the initial sample set includes multiple electronic medical records, each electronic medical record corresponding to a sample data, and the sample data content corresponding to each electronic medical record includes fields such as patient name, gender, age, diagnosis result, diagnosis content, diagnosis department, diagnosis record, etc., as well as information such as the specific content corresponding to the fields.

[0039] The keywords of the target training task include entity information related to the target training task and / or the association relationship between entity information and / or label information, etc. This entity information can be the entity information in the knowledge data determined from the sample data content of the target training task. For example, when the target training task is a training task for a certain type / species of disease, the keywords of the target training task can be based on the medical entity information corresponding to this type of disease, including symptom entities, location entities, age, gender, diagnosis results, etc. information, and / or the association relationship between age, diagnosis results, etc., and / or label information corresponding to some diseases. This medical entity information can be the entity information in the knowledge data determined from the sample data content of the target training task.

[0040] Pre-determine the keywords corresponding to this type of disease. This keyword is determined based on the sample data content, not just based on the fields. It can include synonyms, near-synonyms, or words / sentences with similar semantics of the entities involved in the sample data content, and also includes the association relationship between multiple fields in the sample data content. In this way, the sample data with content errors in the sample data can be processed according to the keywords of the target training task, rather than just through the fields, improving the accuracy and rationality of the candidate sample set.

[0041] In one embodiment, the step of obtaining the keywords of the target training task includes: obtaining the diagnosis result in the knowledge information related to the target training task, the gender information matching the diagnosis result, and the age information matching the diagnosis result, and using the diagnosis result in the knowledge information, the gender information matching the diagnosis result, and the age information matching the diagnosis result as the keywords of the target training task.

[0042] In one embodiment, the step of obtaining the keywords of the target training task includes: obtaining the symptom entities in the pre-annotated expertise in the diagnosis corresponding to the target training task, for example, obtaining the symptom entities in the pre-annotated expertise in the diagnosis corresponding to a certain type of disease; obtaining the discriminative symptom entities in the diagnosis corresponding to each real electronic medical record extracted from the real electronic medical records corresponding to the target training task, and using the symptom entities in the expertise and the discriminative symptom entities as the keywords of the target training task. Understandably, in this embodiment, the keywords of the target training task include at least one symptom entity in the diagnosis corresponding to the target training task. In other embodiments, the keywords of the target training task include other entity data corresponding to the target training task.

[0043] Among them, the symptom entity is also called the symptom named entity, and includes medical terms related to describing diseases, disease incentives, laboratory indicators, etc., such as headache, vomiting, convulsion, unsteady walking, breast cancer, etc. Among them, the step of obtaining the discriminative symptom entities in the diagnosis corresponding to each real electronic medical record extracted from the real electronic medical records can be to obtain the discriminative symptom entities extracted from each real electronic medical record in advance, or directly determine and obtain the discriminative symptom entities in each real electronic medical record.

[0044] Among them, to determine the discriminative symptom entities in the sample data corresponding to the electronic medical record, the following steps can be used: First, use a structured tool to extract the symptom entities in the electronic medical record in a small number of high-quality real electronic medical records related to the target training task, and then use statistical methods, such as TF-IDF (term frequency-inverse document frequency), to determine the symptom entities that appear more frequently in the diagnosis corresponding to an electronic medical record and less frequently in other diagnoses, and use this symptom entity as the discriminative symptom entity in the sample data corresponding to this high-quality electronic medical record.

[0045] In one embodiment, the step of obtaining the keywords of the target training task includes: obtaining the location information in the diagnosis corresponding to the target training task; obtaining the location information, the upper-level location and the lower-level location of the location information from the preset location tree, and using the location information, the upper-level location and the lower-level location as the keywords of the target training task. This keyword can be understood as a location list for fuzzy matching.

[0046] For example, obtain all the location information involved in the training task corresponding to the disease diagnosed as acute upper respiratory tract infection. For example, the location information involved is "upper respiratory tract". Find this location information, as well as the upper and lower locations of this location information, from the preset location tree. For example, the upper location involved in "upper respiratory tract" is "respiratory tract", and the lower locations are "nasal cavity", "throat", etc. Use "upper respiratory tract", "respiratory tract", "nasal cavity", "throat", etc. as the keywords for the target training task. It can be understood that for the disease of "acute upper respiratory tract infection", when a doctor writes a medical record, they may only describe the conditions of locations such as "nasal cavity" and "throat", and not directly mention the location of "upper respiratory tract". It should be noted that this is only for illustrative purposes and for easy understanding, and does not constitute a limitation.

[0047] Among them, the preset location tree is set in advance and can be determined according to the corresponding location knowledge information.

[0048] In addition, the above-mentioned location information in the diagnosis corresponding to the target training task can be mainly obtained through the following methods: one is from the professional knowledge marked by the doctor corresponding to this diagnosis / disease, and the other is the discriminative location corresponding to each diagnosis in the real electronic medical record corresponding to this diagnosis / disease. Among them, the determination of the discriminative location can be to extract the symptom entities in the medical record using a structured tool, and then use the TF-IDF method to find the locations that appear more frequently within one diagnosis and less frequently in other diagnoses as the location information corresponding to this diagnosis, etc.

[0049] For example, for the disease diagnosed as "transient cerebral ischemic attack", although the cause of the disease is a lesion in the "brain", since the lesion in the brain cannot be directly observed, the locations involved in symptoms such as "limb" weakness and "ear" ringing are usually mentioned in the medical record. Therefore, the TF-IDF method can be used to obtain the locations that appear more frequently and are discriminative in this diagnosis as the location information corresponding to this diagnosis.

[0050] In one embodiment, the steps of obtaining the keywords of the target training task include: obtaining the electronic medical record pre-input into the neural network model and the label information corresponding to the electronic medical record. Use the electronic medical record and the corresponding label information as the keywords of the target training task; or use the symptom entities, location information, etc. in the electronic medical record and the label information corresponding to the electronic medical record as the keywords of the target training task. Among them, the label information includes the diagnosis result corresponding to the electronic medical record or the identification information corresponding to the diagnosis result, etc.

[0051] 102. According to the matching result between the sample data content of the sample data and the keywords, perform invalid data cleaning processing on the sample data in the initial sample set to obtain the candidate sample set of the target training task.

[0052] In one embodiment, the step 102 includes: obtaining the sample data content of each sample data in the initial sample set of the target training task; matching the sample data content with keywords to obtain a matching result; determining invalid sample data in the initial sample set according to the matching result; filtering the invalid sample data to perform invalid data cleaning on the sample data in the initial sample set, so as to obtain a candidate sample set for the target training task.

[0053] Matching the keywords of the target training task with each sample data in the initial sample set, and performing invalid data cleaning on the sample data in the initial sample set according to the matching result of the sample data content of each sample data and the keywords, can clean out the sample data that is not relevant to the target training task. Among them, using the keywords of the target training task to screen each electronic medical record in the initial sample set, excluding and filtering the electronic medical records that are not relevant to the target training task, that is, filtering out the unreasonable electronic medical records, so as to obtain a candidate sample set for the target training task.

[0054] In one embodiment, when the keywords of the target training task include the diagnosis result of the diagnosis corresponding to the target training task, the gender information matching the diagnosis result, and the age information matching the diagnosis result, correspondingly, the steps of matching the sample data content with the keywords to obtain a matching result, and determining invalid sample data in the initial sample set according to the matching result include: matching the target diagnosis result and target gender corresponding to the sample data content of each sample data in the initial sample set with the diagnosis result and gender in the keywords, if the match is unsuccessful, then determining the corresponding sample data as invalid sample data; and / or matching the target diagnosis result and target age corresponding to the sample data content of each sample data in the initial sample set with the diagnosis result and age in the keywords, if the match is unsuccessful, then determining the corresponding sample data as invalid sample data. Subsequently, filter the invalid sample data to obtain a candidate sample set.

[0055] In this embodiment, unreasonable sample data with conflicts between the diagnosis result in the diagnosis of the electronic medical record and the age and gender of the patient can be removed.

[0056] For example, check whether the diagnosis result and the patient's gender are consistent with those in the keywords. For example, patients with gynecological and obstetric diseases cannot be male, while patients with diseases such as prostate cannot be female. In this way, filter out the sample data whose diagnosis result and gender in the sample data content do not match the diagnosis result and gender in the keywords.

[0057] For example, check whether the age of the patient is within the common age range of the diagnosis result. Take the disease of neonatal pneumonia as an example. It is clear that the population of patients is neonates (generally referring to infants within 28 days after birth). If a 3-year-old child corresponds to this diagnosis, it indicates that there is a problem with the diagnosis or age in this electronic medical record. In this way, filter out the sample data content where the diagnosis result and age do not match the diagnosis result and age in the keyword.

[0058] In one embodiment, when the keyword of the target training task includes the symptom entity of the diagnosis corresponding to the target training task, correspondingly, the step of matching the sample data content with the keyword to obtain a matching result includes: extracting the target symptom entity in each sample data content in the initial sample set; segmenting the target symptom entity in each sample data content and converting it into a target symptom vector sequence, and determining the symptom vector sequence after segmenting and converting the symptom entity in the keyword; matching the target symptom entity in each sample data content with the symptom entity in the keyword, and matching the corresponding target symptom vector sequence in each sample data content with the symptom vector sequence in the keyword.

[0059] Among them, the method of extracting the target symptom entity can be implemented by using any structured tool. After extracting / determining the target symptom entity or the symptom entity, use a word segmentation tool to segment the target symptom entity or the symptom entity. For example, segment the target symptom entity of "right breast cancer" into two words, "right" and "breast cancer", and use a preset word vector tool to map the obtained words to the corresponding word vector sequence. The preset word vector tool includes Word2Vec (Word to Vector), BERT (Bidirectional Encoder Representations from Transformer, a language representation model), GloVe (Global Vectors for Word Representation, a word representation tool based on global word frequency statistics), etc.

[0060] In some cases, if the number of target symptom entities extracted from the sample data content is too small, for example, the target symptom entity is less than 3, then no matching is performed between the symptom entities, or the corresponding sample data content is directly filtered. In some electronic medical records mainly composed of reexaminations, follow-up consultations, or examination results, there are very few target symptom entities, and the corresponding electronic medical records can be directly filtered.

[0061] In one embodiment, the step of matching the target symptom entities in each sample data content with the symptom entities in the keywords, and the corresponding target symptom vector sequences in each sample data content with the symptom vector sequences in the keywords to obtain a matching result includes: matching the target symptom entities in each sample data content with each symptom entity in the keywords to obtain a first matching quantity; and / or determining the edit distance between the target symptom entities in each sample data content and each symptom entity in the keywords, and performing matching according to the edit distance to obtain a second matching quantity; and / or determining a third distance feature between the target symptom vector sequence corresponding to each sample data content and each symptom vector sequence in the keywords, and performing matching according to the third distance feature to obtain a third matching quantity.

[0062] As Figure 2 shown, it is a schematic diagram of symptom entity matching provided by an embodiment of the present application. Taking an electronic medical record with a diagnosis (disease) of coronary atherosclerosis as an example, after structuring the sample data content of the electronic medical record, the obtained target symptom entities include chest pain, mood abnormality, palpitations, hypertension, etc. The symptom entities in the keywords include chest pain, chest tightness, heartburn, tenderness, etc. Among them, the symptom entities in the keywords can come from the professional knowledge marked by doctors and the discriminative symptom entities from real electronic medical records. The discriminative symptom entities can also be called high-frequency symptom entities.

[0063] Among them, matching the target symptom entities in each sample data content with each symptom entity in the keywords to find the target symptom entities that are exactly the same in the sample data content and the keywords, and determining the number of exactly the same target symptom entities as the first matching quantity. For example, Figure 2 for "chest pain", "palpitations", "heartburn", and "acid reflux" in

[0064] Among them, to determine the edit distance between the target symptom entities in each sample data content and each symptom entity in the keywords, the Levenshtein Distance algorithm or other methods can be used to determine. Using the edit distance for matching, it is determined that the difference in the edit distance within the preset edit distance range (the edit distances are close) indicates a successful match between the corresponding target symptom entity and the corresponding symptom entity, and the number of successfully matched target symptom entities is determined as the second matching quantity. For example, Figure 2The "acute myocardial infarction" in the electronic medical record and the "acute anterior wall myocardial infarction" in the keywords are slightly different literally, but by calculating the edit distance, it is determined that their connotations are not very different, that is, the edit distances are similar and can be successfully matched. By means of the edit distance, two similar symptom entities can be found, such as synonyms and near-synonyms, without requiring the words corresponding to the two symptom entities to be exactly the same.

[0065] Among them, the third distance feature between the target symptom vector sequence (feature vector) corresponding to each sample data content and each symptom vector sequence (feature vector) in the keywords is determined. When the third distance feature is lower than the preset third feature threshold (such as 0.3), it can be understood that the feature vectors are similar, and it is determined that the matching is successful. The number of target symptom entities with successful matching is determined as the third matching number. Among them, the third distance can be the WMD distance (Word Mover’s Distance). The WMD distance can be used to well calculate the difference between two feature vectors of different lengths.

[0066] Suppose the two symptom entities are s=(t1,t2...t n ) and s′=(t′1,t′2...t′ n ), where t represents a word obtained by segmenting the symptom entity using a word segmentation tool. The words after segmenting the two symptom entities are mapped to the corresponding symptom vector sequences (ω1,ω2...ω n ) and (ω’1,ω’2...ω’ n ). The WMD distance is used to calculate the difference degree between the two symptom entities, and the specific calculation formula is shown in formula (1).

[0067]

[0068] Among them, it can be understood that a word in a symptom vector sequence is respectively used to calculate the distance from each word in another symptom vector sequence to find the minimum distance. Finally, the sum of the minimum distances between all the words in this symptom vector sequence and each word in another symptom vector sequence is the WMD distance.

[0069] Such as Figure 2 the examples of successful matching of "agitation" and "emotional abnormality", "chest tightness" and "palpitation" in

[0070] After obtaining the first matching quantity, the second matching quantity, and the third matching quantity, the invalid sample data in the initial sample set can be determined according to the first matching quantity and / or the second matching quantity and / or the third matching quantity in the matching result. In one case, for example, in the manner that the symptom entities are exactly the same, the feature vectors are similar, and the edit distances are similar, it is determined whether the first matching quantity reaches the first preset quantity, whether the second matching quantity reaches the second preset quantity, and whether the third matching quantity reaches the third preset quantity. If the matching quantities corresponding to the sample data of an electronic medical record do not reach the corresponding preset quantities in all three cases, it is determined that the sample data content does not match the keyword, and the sample data is determined as invalid sample data, and the invalid sample data is filtered.

[0071] It can be understood that the three methods of exactly the same symptom entities, similar feature vectors, and similar edit distances in this embodiment are from precise to fuzzy, not limited to complete field matching, can cover the matching of synonyms and near-synonyms, can match more diverse sample data, improve the accuracy of determining invalid sample data, improve the accuracy and diversity of the candidate sample set, and effectively avoid the problem of different expressions of the same noun that widely exist in the medical field.

[0072] In one embodiment, when the keyword of the target training task includes the location information in the diagnosis corresponding to the target training task, as well as the upper-level location and the lower-level location of the location information, correspondingly, the step of matching the sample data content with the keyword to obtain the matching result includes: extracting the target location information in each sample data content in the initial sample set; obtaining the target location information, the target upper-level location, and the target lower-level location of the target location information from the preset location tree to form the target location list of the sample data content; and matching the target location list with the keyword to find the intersection location between the target location list and the keyword, and obtaining the intersection location quantity. Among them, the target location list can be understood as the fuzzy matching list corresponding to the sample data.

[0073] For each sample data in the initial sample set, extract the location information in its sample data content, use the location information as the target location information, form the target location list of the sample data content according to the target location information, and determine the intersection between the target location list and the keyword to obtain the intersection location and determine the intersection location quantity.

[0074] Correspondingly, the step of determining the invalid sample data in the initial sample set according to the matching result includes: if the intersection location quantity corresponding to the sample data is zero, or the intersection location quantity is less than the preset location quantity threshold, it is determined that the sample data is invalid sample data. Filter the invalid sample data.

[0075] This embodiment can filter the site information corresponding to the diagnosis and the affected sites (including the upper-level sites and lower-level sites), and the invalid sample data that is inconsistent with the sites in the sample data.

[0076] In one embodiment, when the target training task includes electronic medical records and the corresponding label information, correspondingly, the keywords of the target training task and the sample data content in the initial sample set are input into the corresponding neural network model, and the neural network model is used for matching. According to the matching value of the matching, the invalid sample data in the initial sample set is determined, and the invalid sample data is filtered to perform invalid data cleaning processing on the sample data in the initial sample set to obtain a candidate sample set.

[0077] In the above embodiment, invalid data cleaning processing is performed on the initial sample set. During the invalid data cleaning processing, based on the sample data content of the sample data, methods such as synonyms, near-synonyms, different expressions of the same symptom, and fuzzy matching are integrated, and also include the association relationship between multiple fields in the sample data content. It can handle the content errors in the sample data to filter the invalid sample data with content errors in the sample data, rather than simply excluding the invalid sample data with field splicing errors or writing errors through fields, improving the accuracy and rationality of the candidate sample set, and ensuring the diversity of the candidate sample set.

[0078] 103. According to the data characteristics of the sample data in the candidate sample set, perform feature clustering processing on the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set.

[0079] In this step, the sample data in the candidate sample set is clustered to divide the diseases corresponding to a class of diseases in the candidate sample set of this class into multiple small categories. The purpose of clustering into multiple small categories is, on the one hand, to clean up some free samples in the candidate sample set that are too different from other sample data, and on the other hand, to facilitate obtaining better training samples from the candidate sample set. This training sample involves multiple small categories, improving the accuracy, rationality, and diversity of the training sample.

[0080] In one embodiment, the step 103 above includes: extracting entity data from the content of each sample data in the candidate sample set; constructing a data feature vector for each sample data according to the entity data; and using the mini-batch K-means clustering algorithm according to the data feature vector to perform feature clustering processing on the candidate sample set to obtain the sample data clustering result of the candidate sample set.

[0081] Among them, the entity data can be all the entity data involved in the sample data content. For example, all the medical entities involved in the electronic medical record content of the electronic medical record, including the symptom entities, site information, etc. mentioned above, and also include other medical entities, such as treatment indicators, treatment measures, etc.

[0082] After obtaining the entity data of the sample data content, construct the data feature vector of each sample data according to the entity data. For example, the entity data corresponding to this type of disease can be obtained. Suppose there are 1000 of them, then construct a 1000-dimensional data feature vector. Initially, the value of each dimension of the data feature vector corresponding to each sample data is zero. When the entity data in the sample data content of a certain sample data is obtained, then in the 1000-dimensional data feature vector, fill the dimension corresponding to the entity data of this sample data, that is, according to the entity data corresponding to the sample data, modify the value of the corresponding dimension in the data feature vector corresponding to this sample data, such as modifying it to 1, so as to obtain the data feature vector of each sample data in the candidate sample set.

[0083] After obtaining the data feature vector of each sample data in the candidate sample set, perform feature clustering processing on the candidate sample set according to the data feature vector. Any clustering algorithm can be used for feature clustering processing. For example, the K-Means algorithm, the Mini Batch K-Means algorithm, etc. can be used, or other clustering algorithms can also be used.

[0084] Among them, the Mini Batch K-Means algorithm is used to reduce the calculation time while still attempting to optimize the objective function. Since the amount of samples calculated is small, the running time will be greatly reduced, and the resulting accuracy degradation is generally within an acceptable range. Generally, the result produced by the Mini Batch K-Means algorithm is only slightly worse than that of the K-Means algorithm. Using the Mini Batch K-Means algorithm has a significant improvement in speed compared to the K-Means algorithm, etc., especially in the case of high-dimensional data. For example, taking a batch of electronic medical record data diagnosed with pneumonia as an example, using about 5000 sample data, the length of each sample data is about 80 words, and after being converted into a data feature vector, it is about 2000-dimensional. The K-Means algorithm takes about 6.53 seconds, while the Mini Batch K-Means algorithm only takes 1.42 seconds, about a 4-fold increase in speed.

[0085] In one embodiment, the step of performing feature clustering processing on the candidate sample set according to the data feature vector by using the Mini Batch K-Means clustering algorithm to obtain the sample data clustering result of the candidate sample set includes: setting the number of clustering sample data and the number of clustering categories for clustering; constructing the Mini Batch K-Means clustering algorithm; extracting batch candidate sample data with the same number as the number of clustering sample data from the candidate sample set, and calling the Mini Batch K-Means clustering algorithm to perform feature clustering processing on the batch candidate sample data according to the number of categories, so as to obtain the sample data clustering result of the candidate sample set.

[0086] Among them, the number of categories can be set to any positive integer, such as 4. The construction interface can be called to construct the mini-batch K-means clustering algorithm. Using the mini-batch K-means clustering algorithm, the batch candidate sample data is automatically clustered to obtain the sample data clustering result of the candidate sample set. The sample data clustering result includes information such as the sample data cluster or category to which each candidate sample data belongs, the number of categories, and the clustering center point of the sample data cluster.

[0087] It should be noted that the sample data clustering result obtained by automatic clustering can reveal the concentrated characteristics of the sample data. For a certain disease, there may be a large number of electronic medical records showing the same concentrated characteristics, and the concentrated characteristics correspond to a typical characteristic of this type of medical record. According to the clustering method, the concentrated characteristics in the candidate sample set can be found, improving the accuracy of training sample acquisition.

[0088] 104. According to the sample data clustering result and the preset outlier data determination condition, the outlier data cleaning process is performed on each sample data cluster in the sample data clustering result to obtain the candidate sample data cluster.

[0089] In this step, the outlier data cleaning process is performed on each sample data cluster according to the sample data clustering result to obtain the candidate sample data cluster. It can be understood that if a sample data is far from the clustering center points of all categories of a certain disease, it means that the sample data is far from all typical manifestation methods of this disease. Such sample data is very likely to be an unreasonable sample or a special medical record combining other difficult and complicated diseases. Such sample data has a great interference to the subsequent training of the artificial intelligence system and is not conducive to the subsequent training of the artificial intelligence system. Such sample data is used as outlier data, and these outlier data need to be cleaned to remove the sample data with large differences from most sample data in the sample data cluster, improving the accuracy of training sample acquisition.

[0090] In one embodiment, the above step 104 includes: for each sample data cluster in the sample data clustering result, obtaining the outlier feature of each sample data in the sample data cluster; determining the preset outlier data determination condition corresponding to the sample data cluster according to the outlier feature; and performing the outlier data cleaning process on the sample data cluster according to the preset outlier data determination condition to obtain the candidate sample data cluster.

[0091] In some cases, two features can be used to describe each sample point in the same category in the sample data clustering result. It should be noted that using two features to describe each sample point is for the convenience of determining outlier data on the one hand, and on the other hand, for the convenience of displaying the sample data clustering result, which will be described in detail later. These two features used to describe each sample point are used as outlier features.

[0092] For example, the outlier features include, in a high-dimensional data feature vector space, a first distance feature between the sample point corresponding to the sample data and the cluster center point, and a second distance feature between the data feature vector of the sample data and the central feature vector corresponding to the cluster center point in the high-dimensional data feature vector space. Among them, the first distance feature and the second distance feature are different distance features. For example, the first distance corresponds to the Euclidean distance, and the second distance corresponds to the cosine distance, etc.

[0093] In one embodiment, the step of obtaining the outlier features of each sample data in the sample data cluster includes: determining, according to the data feature vector of each sample data in the sample data cluster and the central feature vector of the corresponding cluster center point in the sample data cluster, a first distance feature from each sample data to the cluster center point and a second distance feature between the data feature vector of each sample data and the central feature vector; using the first distance feature and the second distance feature as the outlier features of each sample data in the sample data cluster.

[0094] Among them, the first distance feature and the second distance feature in the outlier features can be determined by the following formulas (2) and (3).

[0095]

[0096]

[0097] Among them, for each sample data in each sample data cluster, the x and y coordinates are used to represent the first distance feature and the second distance feature describing the sample data respectively. represents the data feature vector of the sample data, f i represents the value of each feature in the data feature vector, d represents the dimension of the data feature vector. represents the central feature vector of the cluster center point where the sample data is located. represents the value of each feature in the central feature vector of the cluster center point, and l2-norm represents the second norm.

[0098] After determining the outlier features corresponding to each sample data in the sample data cluster, determine the preset outlier data determination condition corresponding to the sample data cluster according to the outlier features. It should be noted that the values corresponding to the preset outlier data determination conditions corresponding to each sample data cluster are different.

[0099] In one embodiment, the step of determining the preset outlier data determination condition corresponding to the sample data cluster according to the outlier feature includes: determining the mean and standard deviation corresponding to the first distance feature of each sample data in the sample data cluster, and the mean and standard deviation corresponding to the second distance feature; determining a first outlier threshold according to the mean and standard deviation corresponding to the first distance feature, and determining a second outlier threshold according to the mean and standard deviation corresponding to the second distance feature; and determining the first outlier threshold and the second outlier threshold as the preset outlier data determination condition corresponding to the sample data cluster.

[0100] Among them, by calculating the mean and standard deviation corresponding to the first distance feature of each sample data in the sample data cluster, and calculating the mean and standard deviation corresponding to the second distance feature of each sample data in the sample data cluster, thus, the mean and standard deviation corresponding to the first distance feature and the mean and standard deviation corresponding to the second distance feature corresponding to each sample data cluster can be obtained.

[0101] For each sample data cluster, a first outlier threshold is determined according to the mean and standard deviation corresponding to the first distance feature, and a second outlier threshold is determined according to the mean and standard deviation corresponding to the second distance feature. Specifically, the first outlier threshold = the mean corresponding to the first distance feature + the standard deviation corresponding to the first distance feature * the standard deviation coefficient, and the second outlier threshold can be obtained in the same way. Thus, the preset outlier data determination condition corresponding to the sample data cluster is obtained.

[0102] In one embodiment, the step of performing outlier data cleaning processing on the sample data cluster according to the preset outlier data determination condition to obtain a candidate sample data cluster includes: determining the sample data with the first distance feature greater than the first outlier threshold or the second distance feature greater than the second outlier threshold in the sample data cluster as outlier sample data; and removing the outlier sample data in the sample data cluster to perform outlier data cleaning processing on the sample data cluster to obtain a candidate sample data cluster corresponding to the sample data cluster.

[0103] Among them, the first distance feature of the sample data in the sample data cluster is greater than the first outlier threshold, which means that the Euclidean distance between the sample point corresponding to the sample data and the clustering center point exceeds the average distance in the sample data cluster, that is, the sample point corresponding to the sample data has a large deviation from the clustering center point of the sample data cluster and does not conform to the core / typical features of the category corresponding to the sample data cluster. Therefore, this sample data is regarded as outlier data, and the sample data corresponding to this outlier data is deleted to perform outlier data cleaning. Similarly, the second distance feature of the sample data in the sample data cluster is greater than the second outlier threshold, which means that the cosine distance between the data feature vector corresponding to the sample data and the central feature vector of the clustering center point exceeds the average distance in the sample data cluster and does not conform to the typical features of the category corresponding to the sample data cluster. This sample data is regarded as outlier data, and the sample data corresponding to this outlier data is deleted to perform outlier data cleaning.

[0104] Clean the outlier data in each sample data cluster to remove the outlier data with large differences from most of the sample data in the sample data cluster, improving the accuracy of obtaining training samples.

[0105] After performing outlier data cleaning on each sample data cluster, each candidate sample data cluster is obtained.

[0106] In one embodiment, the preset outlier data determination condition further includes the preset number of sample data in the sample data cluster. Among them, the preset number of sample data can be 5, etc. Before the step of obtaining the outlier features of each sample data in the sample data cluster, it further includes: obtaining the number of sample data in each sample data cluster in the sample data clustering result; deleting the sample data clusters with the number of sample data less than the preset number to perform outlier data cleaning on the sample data clusters and obtain the remaining sample data clusters; correspondingly, the above-mentioned obtaining the outlier features of each sample data in the sample data cluster includes: obtaining the outlier features of each sample data in the remaining sample data clusters.

[0107] In this embodiment, the sample data clusters with too few sample data in the sample data clustering result can be removed in advance to retain typical sample data as much as possible and improve the accuracy of training samples.

[0108] 105. Select the target training samples for the target training task from the candidate sample data clusters.

[0109] Select the sample data in the candidate sample data clusters again to improve the accuracy of obtaining the target training samples.

[0110] In one embodiment, step 105 described above includes: determining the number of target samples to be selected in each candidate sample data cluster according to the total number of selections corresponding to each candidate sample data cluster, the number of samples in each candidate sample cluster, and the number of categories; and sampling target sample data corresponding to the number of target samples from the candidate sample data clusters as the target training samples for the target training task based on a preset screening method.

[0111] Among them, for the total number of selections corresponding to each candidate sample data cluster, for example, for the disease / diagnosis of inguinal hernia, the total number of selections is determined to be 100, the number of categories is determined to be 4, and the number of samples in each candidate sample cluster is 10, 120, 150, 350, etc.

[0112] In one embodiment, uniform sampling can be performed for each category. Performing uniform sampling for each category means that the number of sample data sampled for each category is the same.

[0113] In this embodiment, the number of samples corresponding to the candidate sample clusters is sorted from smallest to largest. Correspondingly, the number of target samples to be selected for each candidate sample cluster can be calculated according to the following formulas (4), (5), and (6).

[0114]

[0115]

[0116]

[0117] Among them, n i represents the number of target samples to be selected / sampled for the i-th category, S represents the total number of selections required for the disease, t represents the number of categories of a disease clustering, i represents the current category being sampled, and it is a number between 0 and t (excluding 0 and including t). represents the number of samples of the i-th category, that is, the actual number of samples. represents the number of samples that should be sampled for the i-th category without considering the actual number of samples of the i-th category, and it is only an intermediate value used in the calculation. r i represents the number of samples lacking when the number of samples in the current category is insufficient. That is to say, if the actual number of samples of a category is less than the number of target samples to be sampled then r i represents the difference between the number of target samples to be sampled for the i-th category and the actual number of samples, and its initial value r0 = 0.

[0118] Taking the corresponding quantity in the disease of inguinal hernia described above as an example, since the total number of selected samples is 100 and the number of categories is determined to be 4, the number of target samples to be sampled for each category is 25. For the first category, the actual number of samples is 10, so these 10 samples will all be sampled. Then the number of samples lacking in the first category is 25 - 10 = 15. Then the number of target samples to be sampled for the candidate sample clusters of the subsequent second, third, and fourth categories are 30 (25 + 15 / 3), 30, and 30 respectively.

[0119] Among them, the preset screening methods include random screening, sequential screening, etc. For example, when the preset screening method is random screening, then 30 samples need to be randomly selected from the sample quantities 120, 150, and 350 corresponding to the candidate sample clusters of the second, third, and fourth categories respectively. In this way, 100 target sample data are obtained, and these 100 target sample data are used as target training samples.

[0120] Perform uniform sampling for each category in order to retain target training samples with as much diversity as possible.

[0121] In one embodiment, the step of determining the number of target samples to be selected in each candidate sample data cluster according to the total number of selected samples corresponding to each candidate sample data cluster, the number of samples in each candidate sample cluster, and the number of categories can also be to determine the quantity ratio according to the number of samples in each candidate sample cluster, and determine the number of target samples to be selected in each sample cluster according to the quantity ratio and the total number of selected samples. For example, if the number of sample data in a sample data cluster is small, the number of target samples to be selected corresponding to this sample data cluster is small; if the number of samples in a sample data cluster is large, the number of target samples to be selected corresponding to this sample data cluster is large.

[0122] 106, output the target training samples of the target training task.

[0123] For example, the target training samples of the target training task can be output and displayed on a display interface, or output as a certain file, etc., or output to a training model / training module for processing, etc.

[0124] The above embodiments automatically realize the acquisition of training samples. Compared with the manual indexing method, on the basis of ensuring a certain accuracy rate of the training samples, the efficiency of obtaining training samples is greatly improved.

[0125] In one embodiment, after obtaining the sample data clustering result of the candidate sample set and before performing outlier data cleaning processing on each sample data cluster in the sample data clustering result, the following steps are further included: displaying the sample data clustering result through a display interface; after receiving a clustering result confirmation instruction, performing the step of performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and a preset outlier data determination condition. Among them, the clustering result confirmation instruction can be generated by a user triggering a clustering result confirmation control, can be automatically generated after a certain time has elapsed, or can be generated by other means.

[0126] In one embodiment, after obtaining the sample data clustering result of the candidate sample set and before performing outlier data cleaning processing on each sample data cluster in the sample data clustering result, the following steps are further included: obtaining a clustering result verification method; performing verification processing on the sample data clustering result through the clustering result verification method to obtain a verification result; displaying the verification result through a display interface; after receiving a verification result confirmation instruction, performing the step of performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and a preset outlier data determination condition. Among them, the verification result confirmation instruction can be generated by a user triggering a verification result confirmation control, can be automatically generated after a certain time has elapsed, or can be generated by other means.

[0127] In one embodiment, after displaying the sample data clustering result through a display interface, the following steps are further included: after receiving a clustering result verification instruction, obtaining a clustering result verification method; performing verification processing on the sample data clustering result through the clustering result verification method to obtain a verification result; displaying the verification result through a display interface; after receiving a verification result confirmation instruction, performing the step of performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and a preset outlier data determination condition.

[0128] Figure 3 It is another process schematic diagram of the training sample processing method provided by the embodiments of the present application. The training sample processing method includes the following steps.

[0129] 201, Obtain the keywords of the target training task and the sample data in the initial sample set corresponding to the target training task.

[0130] 202, According to the matching result between the sample data content of the sample data and the keywords, perform invalid data cleaning processing on the sample data in the initial sample set to obtain a candidate sample set for the target training task.

[0131] 203. According to the data characteristics of the sample data in the candidate sample set, perform feature clustering processing on the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set.

[0132] 204. Display the sample data clustering result through a display interface.

[0133] Since the sample data in the candidate sample set all correspond to high-dimensional data feature vectors, it is meaningless to display the sample data clustering result in a high-dimensional space, and it is also very difficult to display the sample data clustering result. Therefore, it is necessary to reduce the dimension of the sample data clustering result for display through the display interface.

[0134] In one embodiment, the step of 204 above includes: for each sample data cluster in the sample data clustering result, obtain the first distance feature and the second distance feature of each sample data in the sample data cluster, and in the coordinate system with the first distance feature and the second distance feature as the coordinate axes in the display interface, display the sample points corresponding to each sample data in the sample data cluster according to the first distance feature and the second distance feature. That is, the data corresponding to the sample points are the first distance feature and the second distance feature respectively.

[0135] Such as Figure 4a 、 Figure 4b 、 Figure 4c and Figure 4d shown, it is a schematic diagram of the sample data clustering result corresponding to the disease of inguinal hernia, and the number of clustering categories is 4. Figure 4a 、 Figure 4b 、 Figure 4c and Figure 4d correspond to the first category (class_0), the second category (class_1), the third category (class_2) and the fourth category (class_3) respectively. The x-axis corresponds to the Euclidean distance (distance) between the sample point corresponding to the sample data in the sample data cluster and the clustering center point in the sample data cluster, and the y-axis corresponds to the cosine distance (direction (cos of the vector and center vector)) between the data feature vector corresponding to the sample data in the sample data cluster and the center feature vector of the clustering center point.

[0136] In one embodiment, it further includes: for each sample data cluster, obtain the first outlier threshold determined according to the mean and standard deviation corresponding to the first distance feature, and the second outlier threshold determined according to the mean and standard deviation corresponding to the second distance feature; display the first straight line corresponding to the first outlier threshold and the second straight line corresponding to the second outlier threshold in the corresponding coordinate system.

[0137] Such as Figure 4a 、 Figure 4b 、Figure 4c and Figure 4d The vertical line in Figure 4d is the first line corresponding to the first dissociation threshold, and the horizontal line is the second line corresponding to the second dissociation threshold. The sample points on the right side of the first line indicate that the Euclidean distance from the cluster center point exceeds the distance threshold of this category, corresponding to the first dissociation threshold. The sample points above the second line indicate that the cosine distance from the cluster center point exceeds the distance threshold of this category, corresponding to the second dissociation threshold. Therefore, the sample data corresponding to the sample points on the right side of the first line and above the second line are determined as dissociation data.

[0138] By visually displaying the sample data clustering result on the display interface, users can intuitively feel the clustering result of the sample data and the dissociation data, improving the user experience.

[0139] It should be noted that if most of the sample points fall within the range of the dissociation data, the first dissociation threshold and the second dissociation threshold can be modified by modifying the standard deviation coefficient. The sample data clustering result can also be corrected by modifying the number of categories and the number of clustering samples used for clustering in the mini-batch K-means algorithm.

[0140] 205. After receiving the clustering result verification instruction, obtain the clustering result verification method.

[0141] Among them, the clustering result verification instruction can be generated by the user triggering the clustering result verification control, can also be automatically generated, or can be generated by other means. After receiving the clustering result verification instruction, obtain the clustering result verification method. Among them, the clustering result verification methods include PCA (Principal Component Analysis) verification method, tSNE (SNE is Stochastic Neighbor Embedding) verification method, and can also be other verification methods, etc.

[0142] 206. Verify and process the sample data clustering result through the clustering result verification method to obtain the verification result.

[0143] In one embodiment, the step of 206 above includes: reducing the dimensionality of the data feature vectors of the sample data in the candidate sample set through the clustering result verification method to obtain the low-dimensional feature vectors of each sample data, and mapping the low-dimensional feature vectors to the sample data points corresponding to the sample data; marking the sample data points according to the categories of the sample data clustering result to obtain the marking result; verifying and processing the sample data clustering result according to the marking result to obtain the verification result.

[0144] Dimensionality reduction is performed on the sample data in the candidate sample set for verification in a low-dimensional space, which is convenient for intuitively displaying the sample situation, while reducing the computational amount and improving the verification efficiency. After mapping the low-dimensional feature vectors to the sample data points corresponding to the sample data, the sample data points are marked according to the categories of the sample data clustering results. Among them, from the sample data corresponding to the sample data point, the category of the clustering corresponding to the sample data can also be known, and the sample data point is marked with the identifier corresponding to the category.

[0145] It should be noted that since the sample data all correspond to high-dimensional data feature vectors, when displaying the sample data, it is necessary to perform dimensionality reduction on the high-dimensional data feature vectors for display in the low-dimensional data feature vector space. Dimensionality reduction processing is performed on the high-dimensional data feature vectors of the sample data to facilitate the display of the marking results.

[0146] Correspondingly, when marking the sample data points, visual identifiers can be used for marking to facilitate subsequent visual verification and display; only identifiers such as 0, 1, 2, etc. can also be used for marking, and in this way, algorithms are used to verify the sample data clustering results. For example, algorithms are used to determine whether the sample data points in the same representation corresponding to multiple categories are clustered together, or whether most of them are clustered together, etc. If so, the verification is passed.

[0147] 207, the verification result is displayed through the display interface.

[0148] In an embodiment, the steps of displaying the verification result through the display interface include: after obtaining the sample data points, in the display interface, in the coordinate system formed by the low-dimensional feature vectors, the sample data points corresponding to the low-dimensional feature vectors are displayed; correspondingly, after obtaining the marking result, the marking result is displayed in the corresponding coordinate system to display the verification result. Among them, the marking styles of the sample data points of different categories in the marking result are different. Among them, the marking styles include marking colors and / or marking shapes, etc., and different marking results of different sample data points are facilitated to be distinguished through the marking styles.

[0149] As Figure 5a and Figure 5b shown, it is a schematic diagram of the marking result of marking the sample data clustering result after dimensionality reduction for inguinal hernia disease. Among them, Figure 5a is a schematic diagram of using the PCA verification method to verify the sample data clustering result, Figure 5b is a schematic diagram of using the tSNE verification method to verify the sample data clustering result. Among them, the sample data points of different categories in the sample data clustering result are marked with different-shaped marking styles. From Figure 5a and Figure 5bIt can be seen that most of the sample data points of a category are clustered together, and this verification result also supports the reliability of clustering to a certain extent.

[0150] The verification result is displayed through the display interface in a visual way, enabling users to intuitively feel the verification result and improving the user experience.

[0151] 208. After receiving the verification result confirmation instruction, according to the sample data clustering result and the preset free data determination condition, the free data cleaning process is performed on each sample data cluster in the sample data clustering result to obtain candidate sample data clusters.

[0152] Among them, the verification result confirmation instruction can be generated by the user triggering the verification result confirmation control, can also be automatically generated, or can be generated by other means. After receiving the verification result confirmation instruction, that is, after the verification passes, the subsequent free data cleaning process steps are executed. In this way, it is avoided that the sample data clustering result is unreasonable, resulting in unreasonable acquisition of training samples, and the accuracy of training samples can be improved.

[0153] 209. Select the target training samples of the target training task from the candidate sample data clusters.

[0154] 210. Output the target training samples of the target training task.

[0155] For the steps not detailed in this embodiment, please refer to the description of the corresponding steps in the above text and will not be elaborated here.

[0156] In the embodiment of the present application, the sample data clustering result is displayed in a visual way, and the verification result of the sample data clustering result is also displayed in a visual way. It can be ensured that subsequent operations are performed when the sample data clustering result is reasonable and the verification result is reasonable, further improving the rationality of training samples. At the same time, through the visual way, users can intuitively feel the sample data clustering result and the verification result, improving the user experience.

[0157] All the above technical solutions can be combined arbitrarily to form alternative embodiments of the present application, which will not be elaborated one by one here.

[0158] To facilitate better implementation of the training sample processing method in the embodiment of the present application, the embodiment of the present application also provides a training sample processing device. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the training sample processing device provided in the embodiment of the present application. The training sample processing device 300 may include an acquisition module 301, a first cleaning module 302, a clustering module 303, a second cleaning module 304, a selection module 305, and an output module 306.

[0159] An acquisition module 301, configured to acquire keywords of a target training task and sample data in an initial sample set corresponding to the target training task.

[0160] In one embodiment, when the acquisition module 301 executes the step of acquiring keywords of a target training task, it specifically executes: acquiring a diagnosis result in knowledge information related to the target training task, gender information matching the diagnosis result, and age information matching the diagnosis result, and using the diagnosis result, the gender information matching the diagnosis result, and the age information matching the diagnosis result in the knowledge information as the keywords of the target training task.

[0161] In one embodiment, when the acquisition module 301 executes the step of acquiring keywords of a target training task, it specifically executes: acquiring a symptom entity in a diagnosis corresponding to the target training task, and using the symptom entity as the keyword of the target training task.

[0162] In one embodiment, when the acquisition module 301 executes the step of acquiring keywords of a target training task, it specifically executes: acquiring part information in a diagnosis corresponding to the target training task; acquiring the part information, a superior part, and an inferior part of the part information from a preset part tree, and using the part information, the superior part, and the inferior part as the keywords of the target training task.

[0163] A first cleaning module 302, configured to perform invalid data cleaning processing on the sample data in the initial sample set according to a matching result between the sample data content of the sample data and the keywords, so as to obtain a candidate sample set of the target training task.

[0164] In one embodiment, the first cleaning module 302 includes a matching unit, an invalid data determination unit, and an invalid data cleaning processing unit. The matching unit is configured to match the sample data content with the keywords to obtain a matching result. The invalid data determination unit is configured to determine invalid sample data in the initial sample set according to the matching result. The invalid data cleaning processing unit is configured to filter the invalid sample data to perform invalid data cleaning processing on the sample data in the initial sample set, so as to obtain a candidate sample set.

[0165] In one embodiment, when the keywords of the target training task include the diagnosis result of the diagnosis corresponding to the target training task, the gender information matching the diagnosis result, and the age information matching the diagnosis result, correspondingly, when the matching unit executes the step of matching the sample data content with the keywords, it specifically executes: matching the target diagnosis result and the target gender corresponding to each sample data content in the initial sample set with the diagnosis result and the gender in the keywords; and / or matching the target diagnosis result and the target age corresponding to each sample data content in the initial sample set with the diagnosis result and the age in the keywords.

[0166] In one embodiment, when the keywords of the target training task include the symptom entities of the diagnosis corresponding to the target training task, correspondingly, when the matching unit executes the step of matching the sample data content with the keywords, it specifically executes: extracting the target symptom entities in each sample data content in the initial sample set; segmenting the target symptom entities in each sample data content and converting them into a target symptom vector sequence, and determining the symptom vector sequence after segmenting and converting the symptom entities in the keywords; matching the target symptom entities in each sample data content with the symptom entities in the keywords, and matching the corresponding target symptom vector sequence in each sample data content with the symptom vector sequence in the keywords.

[0167] In one embodiment, when the keywords of the target training task include the location information in the diagnosis corresponding to the target training task, as well as the upper-level location and the lower-level location of the location information, correspondingly, when the matching unit executes the step of matching the sample data content with the keywords, it specifically executes: extracting the target location information in each sample data content in the initial sample set; obtaining the target location information, the target upper-level location and the target lower-level location of the target location information from the preset location tree to form the target location list of the sample data content; matching the target location list with the keywords.

[0168] The clustering module 303 is used to perform feature clustering processing on the sample data in the candidate sample set according to the data features of the sample data in the candidate sample set, and obtain the sample data clustering result of the candidate sample set.

[0169] In one embodiment, the clustering module 303 includes a feature construction unit and a clustering unit. Among them, the feature construction unit is used to extract entity data from each sample data content in the candidate sample set; construct a data feature vector for each sample data according to the entity data. The clustering unit is used to perform clustering on the candidate sample set according to the data feature vector by using the mini-batch K-means clustering algorithm to obtain the sample data clustering result of the candidate sample set.

[0170] In one embodiment, the clustering unit is specifically configured to set the number of clustering sample data for clustering and the number of categories for clustering; construct a mini-batch K-means clustering algorithm; extract batch candidate sample data with the same number as the number of clustering sample data from the candidate sample set, and call the mini-batch K-means clustering algorithm to cluster the batch candidate sample data according to the number of categories, so as to obtain the sample data clustering result of the candidate sample set.

[0171] The second cleaning module 304 is configured to perform free data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and a preset free data determination condition, so as to obtain candidate sample data clusters.

[0172] In one embodiment, the second cleaning module 304 includes a free feature acquisition unit, a free condition determination unit, and a free data cleaning unit. Among them, the free feature acquisition unit is configured to, for each sample data cluster in the sample data clustering result, acquire the free feature of each sample data in the sample data cluster. The free condition determination unit is configured to determine the preset free data determination condition corresponding to the sample data cluster according to the free feature. The free data cleaning unit is configured to perform free data cleaning processing on the sample data cluster according to the preset free data determination condition, so as to obtain candidate sample data clusters.

[0173] In one embodiment, the free feature acquisition unit is specifically configured to determine, according to the data feature vector of each sample data in the sample data cluster and the central feature vector of the corresponding clustering center point in the sample data cluster, a first distance feature from each sample data to the clustering center point and a second distance feature between the data feature vector of each sample data and the central feature vector, where the first distance feature and the second distance feature are different distance features; and use the first distance feature and the second distance feature as the free feature of each sample data in the sample data cluster.

[0174] In one embodiment, the free condition determination unit is specifically configured to determine the mean and standard deviation corresponding to the first distance feature of each sample data in the sample data cluster and the mean and standard deviation corresponding to the second distance feature; determine a first free threshold according to the mean and standard deviation corresponding to the first distance feature, and determine a second free threshold according to the mean and standard deviation corresponding to the second distance feature; and determine the first free threshold and the second free threshold as the preset free data determination condition corresponding to the sample data cluster.

[0175] In one embodiment, the free data cleaning unit is specifically configured to: determine the sample data in the sample data cluster whose first distance feature is greater than the first free threshold or the second distance feature is greater than the second free threshold as free sample data; remove the free sample data in the sample data cluster to perform free data cleaning on the sample data cluster, so as to obtain a candidate sample data cluster corresponding to the sample data cluster.

[0176] In one embodiment, the preset free data determination condition further includes a preset number of sample data in the sample data cluster. Correspondingly, the second cleaning module 304 further includes a quantity acquisition unit. The quantity acquisition unit is further configured to acquire the number of samples in each sample data cluster in the sample data clustering result. The free data cleaning unit is further configured to delete the sample data clusters with the number of samples less than the preset number to perform free data cleaning on the sample data cluster, so as to obtain the remaining sample data clusters. Correspondingly, the free feature acquisition unit is further configured to acquire the free features of each sample data in the remaining sample data clusters.

[0177] The selection module 305 is configured to select target training samples for the target training task from the candidate sample data clusters.

[0178] In one embodiment, the selection module 305 includes a quantity determination unit and a sampling unit. The quantity determination unit is configured to determine the number of target samples to be selected in each candidate sample data cluster according to the total selection quantity corresponding to each candidate sample data cluster, the number of samples in each candidate sample cluster, and the number of categories. The sampling unit is configured to sample target sample data corresponding to the number of target samples from the candidate sample data clusters based on a preset screening method. The preset training sample screening condition includes the preset screening method and the number of target samples.

[0179] The output module 306 is configured to output the target training samples for the target training task.

[0180] In one embodiment, as Figure 6 shown, the training sample processing device 300 may further include a display module 307. The display module 307 is configured to display the sample data clustering result through a display interface after obtaining the sample data clustering result of the candidate sample set and before performing free data cleaning on each sample data cluster in the sample data clustering result.

[0181] In one embodiment, as Figure 6As shown in the figure, the training sample processing device 300 may further include a verification module 308. The verification module 308 is configured to obtain a clustering result verification method after obtaining the sample data clustering result of the candidate sample set and before performing outlier data cleaning processing on each sample data cluster in the sample data clustering result; verify the sample data clustering result through the clustering result verification method to obtain a verification result. Correspondingly, the display module 307 is further configured to display the verification result through a display interface.

[0182] Any combination of the above technical solutions can form an optional embodiment of the present application, and the beneficial effects that can be achieved can also be referred to the descriptions above, which will not be elaborated here one by one.

[0183] Correspondingly, an embodiment of the present application further provides a computer device, which can be a terminal or a server. As Figure 7 shown Figure 7 is a schematic structural diagram of the computer device provided by the embodiment of the present application. The computer device 400 includes a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, and a computer program stored in the memory 402 and executable on the processor. Among them, the processor 401 is electrically connected to the memory 402. Those skilled in the art can understand that the structural diagram of the computer device shown in the figure does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or arrange different components.

[0184] The processor 401 is the control center of the computer device 400, connecting various parts of the entire computer device 400 through various interfaces and lines, and executing various functions and processing data of the computer device 400 by running or loading software programs (computer programs) and / or modules stored in the memory 402, and calling the data stored in the memory 402, so as to monitor the computer device 400 as a whole.

[0185] In the embodiment of the present application, the processor 401 in the computer device 400 will load the instructions corresponding to the processes of one or more application programs into the memory 402 according to the following steps, and the processor 401 will run the application programs stored in the memory 402 to implement various functions:

[0186] Obtain the keywords of the target training task and the sample data in the initial sample set corresponding to the target training task; according to the matching result between the sample data content of the sample data and the keywords, perform invalid data cleaning processing on the sample data in the initial sample set to obtain the candidate sample set of the target training task; according to the data characteristics of the sample data in the candidate sample set, perform feature clustering processing on the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set; according to the sample data clustering result and the preset free data determination condition, perform free data cleaning processing on each sample data cluster in the sample data clustering result to obtain candidate sample data clusters; select the target training samples of the target training task from the candidate sample data clusters; output the target training samples of the target training task.

[0187] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.

[0188] Optionally, as Figure 7 shown, the computer device 400 further includes: a touch display screen 403, a radio frequency circuit 404, an audio circuit 405, an input unit 406, and a power supply 407. Among them, the processor 401 is electrically connected to the touch display screen 403, the radio frequency circuit 404, the audio circuit 405, the input unit 406, and the power supply 407 respectively. Those skilled in the art can understand that Figure 7 the computer device structure shown in

[0189] The touch display screen 403 can be used to display a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface. The touch display screen 403 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 401 to determine the type of touch event. Subsequently, the processor 401 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 403 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 403 can also be used as a part of the input unit 406 to implement the input function.

[0190] In the embodiments of the present application, the touch display screen 403 is used to present a graphical user interface and receive operation instructions generated by a user acting on the graphical user interface.

[0191] The radio frequency circuit 404 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other computer devices through wireless communication, and transmit and receive signals with the network device or other computer devices.

[0192] The audio circuit 405 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 405 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 405 and then converted into audio data. After the audio data is output to the processor 401 for processing, it is transmitted through the radio frequency circuit 404 to, for example, another computer device, or the audio data is output to the memory 402 for further processing. The audio circuit 405 may also include an earphone jack to provide communication between a peripheral earphone and the computer device.

[0193] The input unit 406 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, face information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0194] The power supply 407 is used to supply power to each component of the computer device 400. Optionally, the power supply 407 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 407 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0195] Although Figure 7 not shown in the figure, the computer device 400 can also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.

[0196] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0197] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0198] Therefore, the embodiments of the present application provide a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute the steps in any of the training sample processing methods provided by the embodiments of the present application. For example, the computer program can execute the following steps:

[0199] Obtain the keywords of the target training task and the sample data in the initial sample set corresponding to the target training task; according to the matching result between the sample data content of the sample data and the keywords, perform invalid data cleaning processing on the sample data in the initial sample set to obtain the candidate sample set of the target training task; according to the data characteristics of the sample data in the candidate sample set, perform feature clustering processing on the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set; according to the sample data clustering result and the preset outlier data determination condition, perform outlier data cleaning processing on each sample data cluster in the sample data clustering result to obtain candidate sample data clusters; select the target training samples of the target training task from the candidate sample data clusters; output the target training samples of the target training task.

[0200] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.

[0201] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0202] Since the computer program stored in the storage medium can execute the steps in any of the training sample processing methods provided in the embodiments of the present application, the beneficial effects achievable by any of the training sample processing methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments, which will not be elaborated herein.

[0203] The above has introduced in detail a training sample processing method, device, storage medium, and computer device provided in the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for processing training samples, characterized in that, Including: Obtain the keywords of the target training task and the sample data in the initial sample set corresponding to the target training task; The keywords of the target training task include the symptom entities in the diagnosis corresponding to the target training task; Match the sample data with the keywords, and perform invalid data cleaning on the sample data in the initial sample set according to the matching result between the sample data content of the sample data and the keywords, to obtain the candidate sample set of the target training task; The step of matching the sample data with the keywords includes: Obtain the sample data content of the sample data, and extract the target symptom entities in each sample data content in the initial sample set; Perform word segmentation on the target symptom entities in each sample data content and convert them into a target symptom vector sequence, and determine the symptom vector sequence after word segmentation of the symptom entities in the keywords; Match the target symptom entities in each sample data content with the symptom entities in the keywords, and match the corresponding target symptom vector sequence in each sample data content with the symptom vector sequence in the keywords; Perform feature clustering on the sample data in the candidate sample set according to the data characteristics of the sample data in the candidate sample set, to obtain the sample data clustering result of the candidate sample set; According to the sample data clustering result and the preset outlier data determination condition, perform outlier data cleaning on each sample data cluster in the sample data clustering result, to obtain candidate sample data clusters; the preset outlier data determination condition corresponding to the sample data cluster is determined according to the outlier characteristics of the sample data included in the sample data cluster, and the values corresponding to the preset outlier data determination conditions of different sample data clusters are not completely the same; Select the target training samples of the target training task from the candidate sample data clusters; Output the target training samples of the target training task.

2. The training sample processing method according to claim 1, wherein The step of performing outlier data cleaning on each sample data cluster in the sample data clustering result according to the sample data clustering result and the preset outlier data determination condition, to obtain candidate sample data clusters, includes: For each sample data cluster in the sample data clustering result, obtain the outlier characteristics of each sample data in the sample data cluster; Determine the preset outlier data determination condition corresponding to the sample data cluster according to the outlier characteristics; Perform outlier data cleaning on the sample data cluster according to the preset outlier data determination condition, to obtain candidate sample data clusters.

3. The training sample processing method according to claim 2, wherein, The step of obtaining the outlier characteristics of each sample data in the sample data cluster includes: According to the data feature vector of each sample data in the sample data cluster and the central feature vector of the corresponding clustering center point in the sample data cluster, determine the first distance feature of each sample data to the clustering center point, and the second distance feature between the data feature vector of each sample data and the central feature vector, where the first distance feature and the second distance feature are different distance features; Use the first distance feature and the second distance feature as the outlier features of each sample data in the sample data cluster.

4. The training sample processing method according to claim 2, wherein The step of determining the preset outlier data determination condition corresponding to the sample data cluster according to the outlier features includes: Determine the mean and standard deviation corresponding to the first distance feature of each sample data in the sample data cluster, and the mean and standard deviation corresponding to the second distance feature; Determine a first outlier threshold according to the mean and standard deviation corresponding to the first distance feature, and determine a second outlier threshold according to the mean and standard deviation corresponding to the second distance feature; Determine the first outlier threshold and the second outlier threshold as the preset outlier data determination condition corresponding to the sample data cluster.

5. The training sample processing method according to claim 2, wherein The outlier feature of each sample data includes a first distance feature and a second distance feature, and the preset outlier data determination condition includes a first outlier threshold and a second outlier threshold. The step of performing outlier data cleaning processing on the sample data cluster according to the preset outlier data determination condition to obtain a candidate sample data cluster includes: Determine the sample data in the sample data cluster whose first distance feature is greater than the first outlier threshold or the second distance feature is greater than the second outlier threshold as outlier sample data; Remove the outlier sample data in the sample data cluster to perform outlier data cleaning processing on the sample data cluster to obtain the candidate sample data cluster corresponding to the sample data cluster.

6. The training sample processing method according to claim 2, wherein The preset outlier data determination condition includes the preset number of sample data in the sample data cluster. Before the step of obtaining the outlier features of each sample data in the sample data cluster, it further includes: Obtain the number of sample data in each sample data cluster in the sample data clustering result; Delete the sample data clusters with the number of samples less than the preset number to perform outlier data cleaning processing on the sample data clusters to obtain the remaining sample data clusters; The step of obtaining the outlier features of each sample data in the sample data cluster includes: obtaining the outlier features of each sample data in the remaining sample data clusters.

7. The training sample processing method according to claim 1, wherein The step of performing invalid data cleaning processing on the sample data in the initial sample set according to the matching result between the sample data content of the sample data and the keyword to obtain the candidate sample set for the target training task includes: Obtain the sample data content of the sample data and match the sample data content with the keyword to obtain a matching result; Determine the invalid sample data in the initial sample set according to the matching result; Filter the invalid sample data to perform invalid data cleaning processing on the sample data in the initial sample set to obtain a candidate sample set.

8. The training sample processing method according to claim 7, wherein The steps of obtaining the keywords of the target training task include: obtaining the diagnosis result in the knowledge information related to the target training task, the gender information matched with the diagnosis result, and the age information matched with the diagnosis result, and using the diagnosis result, the gender information matched with the diagnosis result, and the age information matched with the diagnosis result in the knowledge information as the keywords of the target training task; The steps of matching the sample data content with the keywords include: matching the target diagnosis result and target gender corresponding to each sample data content in the initial sample set with the diagnosis result and gender in the keywords; and / or matching the target diagnosis result and target age corresponding to each sample data content in the initial sample set with the diagnosis result and age in the keywords.

9. The training sample processing method according to claim 1, wherein The steps of matching the target symptom entity in each sample data content with the symptom entity in the keywords, and the corresponding target symptom vector sequence in each sample data content with the symptom vector sequence in the keywords include: matching the target symptom entity in each sample data content with each symptom entity in the keywords; and / or determining the edit distance between the target symptom entity in each sample data content and each symptom entity in the keywords, and performing matching according to the edit distance; and / or determining the third distance feature between the target symptom vector sequence corresponding to each sample data content and each symptom vector sequence in the keywords, and performing matching according to the third distance feature.

10. The training sample processing method according to claim 7, wherein The steps of obtaining the keywords of the target training task include: obtaining the location information in the diagnosis corresponding to the target training task; obtaining the location information, the upper location and the lower location of the location information from a preset location tree, and using the location information, the upper location and the lower location as the keywords of the target training task; The steps of matching the sample data content with the keywords include: extracting the target location information in each sample data content in the initial sample set; obtaining the target location information, the target upper location and the target lower location of the target location information from a preset location tree to form a target location list of the sample data content; matching the target location list with the keywords.

11. The training sample processing method according to claim 1, wherein, One candidate sample data cluster corresponds to one category. The steps of selecting the target training sample of the target training task from the candidate sample data clusters include: determining the number of target samples to be selected in each candidate sample data cluster according to the total selection quantity corresponding to each candidate sample data cluster, the number of samples in each candidate sample cluster, and the number of categories; sampling, based on a preset screening method, the target sample data corresponding to the number of target samples from the candidate sample data clusters as the target training sample of the target training task.

12. The training sample processing method according to claim 1, wherein The step of performing feature clustering processing on the sample data in the candidate sample set according to the data characteristics of the sample data in the candidate sample set to obtain the sample data clustering result of the candidate sample set includes: Extract entity data from the content of each sample data in the candidate sample set; Construct a data feature vector for each sample data according to the entity data; According to the data feature vector, use the mini-batch K-means clustering algorithm to perform feature clustering processing on the candidate sample set to obtain the sample data clustering result of the candidate sample set.

13. The training sample processing method according to claim 12, wherein The step of using the mini-batch K-means clustering algorithm to perform feature clustering processing on the candidate sample set according to the data feature vector to obtain the sample data clustering result of the candidate sample set includes: Set the number of clustering samples and the number of clustering categories for clustering; Construct the mini-batch K-means clustering algorithm; Extract a batch of candidate sample data with the same number as the number of clustering samples from the candidate sample set, and call the mini-batch K-means clustering algorithm to perform feature clustering processing on the batch of candidate sample data according to the number of categories to obtain the sample data clustering result of the candidate sample set.

14. The training sample processing method according to claim 1, wherein After obtaining the sample data clustering result of the candidate sample set and before performing outlier data cleaning processing on each sample data cluster in the sample data clustering result, it further includes: Display the sample data clustering result through a display interface; After receiving the clustering result confirmation instruction, execute the step of performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and the preset outlier data determination condition.

15. The training sample processing method according to claim 14, wherein The step of displaying the sample data clustering result through the display interface includes: For each sample data cluster in the sample data clustering result, obtain the first distance feature and the second distance feature of each sample data in the sample data cluster; In the coordinate system with the first distance feature and the second distance feature as the coordinate axes in the display interface, display the sample points corresponding to each sample data in the sample data cluster according to the first distance feature and the second distance feature.

16. The training sample processing method according to claim 14, wherein After displaying the sample data clustering result through the display interface, it further includes: After receiving the clustering result verification instruction, obtain the clustering result verification method; Perform verification processing on the sample data clustering result through the clustering result verification method to obtain a verification result; Display the verification result through the display interface; After receiving the verification result confirmation instruction, execute the step of performing outlier data cleaning processing on each sample data cluster in the sample data clustering result according to the sample data clustering result and the preset outlier data determination condition.

17. The training sample processing method according to claim 16, wherein The step of performing verification processing on the sample data clustering result through the clustering result verification method to obtain a verification result includes: Through the clustering result verification method, reduce the dimension of the data feature vector of the sample data in the candidate sample set to obtain a low-dimensional feature vector for each sample data, and map the low-dimensional feature vector to the sample data point corresponding to the sample data; Label the sample data points according to the categories of the clustering results of the sample data to obtain a labeling result; Perform a verification process on the clustering results of the sample data according to the labeling result to obtain a verification result.

18. A training sample processing device, characterized in that, It includes: An acquisition module, configured to acquire keywords of a target training task and sample data in an initial sample set corresponding to the target training task; The keywords of the target training task include symptom entities in the diagnosis corresponding to the target training task; A first cleaning module, configured to match the sample data with the keywords, and perform invalid data cleaning processing on the sample data in the initial sample set according to the matching result between the sample data content of the sample data and the keywords, to obtain a candidate sample set of the target training task; The step of matching the sample data with the keywords includes: Obtain the sample data content of the sample data, and extract target symptom entities in each sample data content in the initial sample set; Segment the target symptom entities in each sample data content and convert them into a target symptom vector sequence, and determine the symptom vector sequence after segmentation and conversion of the symptom entities in the keywords; Match the target symptom entities in each sample data content with the symptom entities in the keywords, and match the corresponding target symptom vector sequence in each sample data content with the symptom vector sequence in the keywords; A clustering module, configured to perform feature clustering processing on the sample data in the candidate sample set according to the data features of the sample data in the candidate sample set, to obtain a clustering result of the sample data in the candidate sample set; A second cleaning module, configured to perform free data cleaning processing on each sample data cluster in the clustering result of the sample data according to the clustering result of the sample data and a preset free data determination condition, to obtain candidate sample data clusters; the preset free data determination condition corresponding to the sample data cluster is determined according to the free features of the sample data included in the sample data cluster, and the values corresponding to the preset free data determination conditions of different sample data clusters are not completely the same; A selection module, configured to select target training samples of the target training task from the candidate sample data clusters; An output module, configured to output the target training samples of the target training task.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the steps in the training sample processing method according to any one of claims 1-17.

20. A computer device, characterized in that, The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the steps in the training sample processing method according to any one of claims 1-17 by calling the computer program stored in the memory.

Citation Information

Patent Citations

  • Buddhist question and answer pair construction method and device, equipment and storage medium

    CN112988999A

  • Electronic device, picture sample set generation method, and computer readable storage medium

    WO2019237558A1