Data collection and cleaning method and system for large medical special disease model
The python toolkit analyzes PDF format documents, combined with pre-processing and post-processing steps, solves the normative problems of medical data collection and cleaning, improves data quality, and ensures the stability and accuracy of medical big models.
Patent Information
- Application Number
- CN202510656427.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-08
AI Technical Summary
The collection and cleaning process of traditional Chinese medicine data in the prior art lacks standardization, resulting in unstable performance of large models, noise data affects model accuracy, and lacks efficient data cleaning methods.
Use the Python toolkit to analyze PDF format documents, clean medical data through pre-processing and post-processing steps, including line spacing calculation, regular expressions to remove noise, identify and delete abnormal data, and ensure data quality.
It realizes efficient and accurate data collection and cleaning, provides a solid data foundation for medical big models, shortens the development cycle, and ensures the interpretability and accuracy of the results.
Smart Images

Figure CN120277172A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data processing, and particularly relates to a method and system for data collection and cleaning based on a large medical disease model. Background Art
[0002] With the rapid development of artificial intelligence technology, large models have gradually been introduced into the medical field to assist in diagnosis, disease prediction, and personalized treatment. For example, in the medical field, a large amount of clinical case data can be used to train a diagnostic large model, which helps doctors make more accurate diagnostic decisions. The effectiveness and accuracy of these large models largely depend on high-quality data sets.
[0003] However, in many current cases, the quality of the corpus is not high. For example, there are a large number of noisy data, which will affect the performance of the model. Therefore, before constructing the corpus, we need to preprocess and clean the original data to improve the data quality and ensure the reliability and accuracy of the data. With the wide application of large-scale data, the quality requirements for text data are getting higher and higher. Poor data quality may lead to training errors, such as problems with the output results of the NLP model being chaotic, incoherent, illogical, repetitive responses, etc. Therefore, high-quality text data cleaning methods have become particularly important.
[0004] Currently, research and practice on data cleaning mainly focus on the following aspects: handling missing values and anomalies: methods for handling missing values include directly deleting missing values, filling with the mean, filling with Bayesian regression, using maximum likelihood estimation, association rules, etc. Methods for handling outliers include outlier detection, outlier detection based on distance calculation, density-based local outlier factor detection, etc. However, there is no standardized operation process for the collection and cleaning process of medical data, resulting in instability in the performance and application effects of large models. Summary of the Invention
[0005] In view of the above deficiencies of the prior art, the present application provides a method and system for data collection and cleaning of a large medical disease model.
[0006] In a first aspect, the present application proposes a method for data collection and cleaning of a large medical disease model, including the following steps:
[0007] Determine the scope of target disease types, and collect core medical data according to the scope of target disease types. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specific disease, laws and regulations corresponding to the drugs, knowledge graphs corresponding to the specific disease, typical cases, papers, and clinical trials;
[0008] Taking a PDF format document with non - picture and selected text as the data collection standard for the core medical data, and obtaining the PDF text data to be processed according to the data collection standard;
[0009] Using a Python toolkit to analyze the PDF text data to be processed, taking the text data that is read normally as the first - type text data and performing pre - processing cleaning steps on it to obtain the cleaned first PDF text data; taking the text data that is read abnormally as the second - type text data and performing post - processing cleaning steps on it to obtain the cleaned second PDF text data;
[0010] Integrating the first PDF text data and the second PDF text data to obtain the final cleaned text data.
[0011] In some embodiments, the pre - processing cleaning steps include:
[0012] Step S1: Using a Python toolkit to read the first - type text data line by line, obtaining the corresponding attribute values for each line, where the attribute values include color, text content, and horizontal and vertical coordinate values;
[0013] Step S2: Selecting the vertical coordinate value of the text of any line (excluding the first line) in the first - type text data, and subtracting the text coordinate value of its previous line to obtain the line spacing;
[0014] Step S3: Judging the number of pages of the first - type text data. If the number of pages is less than 5, directly calculate the first mode of the line spacing; if the number of pages is greater than or equal to 5, delete the content of the first 20% and the last 20% of the pages and then obtain the line spacing, and calculate the first mode of the line spacing, and define the calculated first mode as the body text line spacing;
[0015] Step S4: Counting the mode of the body text font size of the first - type text data as the second mode;
[0016] Step S5: Judging whether the first - type text data is a Chinese document through function code. If the document is in Chinese, set the error value of the body text line spacing to 1.5; if the document is in English, set the error value of the full - text line spacing to 0.5, and process the first - type text data according to the error value of the full - text line spacing to obtain the PDF text data to be cleaned;
[0017] Step S6: Traversing the PDF text data to be cleaned, screening out the body text lines according to the body text line spacing, clearing the text data of non - body text lines, and at the same time judging whether the font meets the second mode standard. If it does not meet the standard, clear it to obtain the preliminary cleaning result;
[0018] Step S7: Use regular expressions to remove keywords, summaries, superscripts, and references from the preliminary cleaning result to obtain the first cleaned pdf text data.
[0019] In some embodiments, the post-processing cleaning step includes:
[0020] Traverse each line of the second type of text data, define the lines with less than 5 Chinese characters or lowercase letters as garbled lines. If the proportion of garbled lines in the total number of lines of the full text is greater than 0.5, define the pdf text file as a garbled file. At this time, directly delete the current garbled file, and transfer the remaining pdf text data to be processed to the pre-processing step for data cleaning.
[0021] In some embodiments, the post-processing cleaning step further includes:
[0022] Analyze whether the file size of the second type of text data is less than or equal to 3 kb. If so, directly identify it as a garbled file and delete it.
[0023] In a second aspect, the present application proposes a data collection and cleaning system for a medical specialty large model, including a specialty data collection module, a preliminary processing module, a cleaning module, and an integration module;
[0024] The specialty data collection module is used to determine the scope of target specialty disease types, and collect core medical data according to the scope of target specialty disease types. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specialty, laws and regulations corresponding to the drugs, knowledge graphs corresponding to the specialty, typical cases, papers, and clinical trials;
[0025] The preliminary processing module is used to use a non-picture and selected text pdf format document as the data collection standard for the core medical data, and obtain the pdf text data to be processed according to the data collection standard;
[0026] The cleaning module is used to analyze the pdf text data to be processed by using a python toolkit, regard the text data read normally as the first type of text data and process it by using the pre-processing cleaning step to obtain the first cleaned pdf text data; regard the text data read abnormally as the second type of text data and process it by using the post-processing cleaning step to obtain the second cleaned pdf text data;
[0027] The integration module is used to integrate the first pdf text data and the second pdf text data to obtain the final cleaned text data.
[0028] In some embodiments, the cleaning module includes a pre-treatment cleaning unit, and the pre-treatment cleaning unit is used to perform a pre-treatment cleaning step, and the pre-treatment cleaning step includes:
[0029] Step S1: using a python toolkit to read the first type of text data line by line, and obtain attribute values corresponding to each line, wherein the attribute values include color, text content, and horizontal and vertical coordinate values;
[0030] Step S2: Select the text vertical coordinate value of any line except the first line in the first type of text data, and subtract the text coordinate value of the previous line to obtain the line spacing;
[0031] Step S3: determining the number of pages of the first type of text data; if the number of pages is less than 5 pages, directly calculating the first mode of the line spacing; if the number of pages is greater than or equal to 5 pages, deleting the first 20% and the last 20% of the pages to obtain the line spacing, and calculating the first mode of the line spacing, and defining the calculated first mode as the text line spacing;
[0032] Step S4: Counting the mode of the font size of the first type of text data as the second mode;
[0033] Step S5: judging whether the first category of text data is a Chinese document through the function code, if the document is Chinese, setting the error value of the text line spacing to 1.5, if the document is English, setting the error value of the full text line spacing to 0.5, processing the first category of text data according to the error value of the full text line spacing, and obtaining the PDF text data to be cleaned;
[0034] Step S6: traverse the PDF text data to be cleaned, filter out the text lines according to the text line spacing, remove the text data of non-text lines, and determine whether the font meets the second mode standard. If not, remove it to obtain a preliminary cleaning result;
[0035] Step S7: removing keywords, summaries, superscripts and references from the preliminary cleaning results using regular expressions to obtain the cleaned first PDF text data.
[0036] In some embodiments, the cleaning module includes a post-processing cleaning unit, and the post-processing cleaning unit is used to perform a post-processing cleaning step, and the post-processing cleaning step includes:
[0037] Traverse each line of the second type of text data, and define the lines with less than 5 Chinese characters or lowercase letters as garbled lines. If the proportion of garbled lines to the total number of lines of the full text is greater than 0.5, the PDF text file is defined as a garbled file, and the current garbled file is directly deleted.
[0038] In some embodiments, the cleaning module includes a determination and deletion unit, which is configured to analyze whether the file size of the second type of text data is less than or equal to 3 kb. If so, it is directly determined as a garbled file and deleted.
[0039] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the above method.
[0041] Advantages of the present invention:
[0042] Taking the calculation of line spacing based on coordinate values as the basis, the idea of cleaning the text body is obtained, realizing efficient and accurate data collection and cleaning, providing a solid data foundation for the application of medical large models, shortening the development cycle of medical large models, and ensuring the interpretability and accuracy of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is the overall flowchart of the present invention.
[0044] Figure 2 It is the system principle block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0046] In a first aspect, the present application provides a method for data collection and cleaning of a medical specialty large model, as Figure 1 shown, including the following steps:
[0047] S100: Determine the scope of the target specialty disease, and collect core medical data according to the scope of the target specialty disease. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specialty disease, laws and regulations corresponding to the drugs, knowledge graphs corresponding to the specialty disease, typical cases, papers, and clinical trials;
[0048] Among them, it includes: 1. Select the disease scope:
[0049] Generally, for each specific disease, it is based on the set of sub - category codes under the category code of the disease in ICD - 10. Taking type 1 diabetes as an example, its ICD - 10 category code is E10, and the sub - category codes under it can vary according to the type of diabetes, complications, pathological features, etc. For example, the sub - category codes of type 1 diabetes may include: E10.0: Type 1 diabetes with ketoacidosis, E10.1: Type 1 diabetes with hypoglycemia, E10.9: Type 1 diabetes, unspecified, etc.
[0050] Translate all the disease types under this specialty into English by querying the medical dictionary to facilitate subsequent data collection in both Chinese and English respectively.
[0051] 2. Data collection:
[0052] (1) Core data
[0053] Guidelines and expert consensus: It is necessary to include all the latest authoritative institution - issued diagnosis and treatment guidelines for the corresponding specific disease.
[0054] Textbooks: Preferred textbooks published by People's Medical Publishing House related to the diagnosis, treatment, and case analysis of the disease. High - quality textbooks published by other publishers can be designated by experts.
[0055] Online Q&A: Experts provide keywords related to the disease to obtain Q&A data containing these keywords.
[0056] (2) Other data, including:
[0057] Drug instructions: Drugs closely related to the diagnosis and treatment process of this disease, such as insulin drugs;
[0058] Laws and regulations: Mainly laws and regulations regarding drug use;
[0059] Authoritative atlases: Knowledge bases and knowledge graphs unique to this specific disease;
[0060] Typical cases: Typical cases in textbooks, as well as de - sensitized diagnosis and treatment cases in the hospital diagnosis and treatment scenarios;
[0061] Papers: Lists of authoritative journals on the diagnosis and treatment of this specific disease, as well as various synonyms and sub - concept synonyms of the disease;
[0062] Clinical trials: Clinical trial records of this disease at home and abroad.
[0063] S200: Use non - picture and selectable text in pdf format documents as the data collection standard for the core medical data, and obtain the pdf text data to be processed according to the data collection standard;
[0064] Among them, for the documents involved in step S100, the collection standard is pdf format with non - picture and normal selectable text.
[0065] S300: Analyze the to-be-processed pdf text data using a python toolkit. Take the text data that is read normally as the first type of text data and process it using a preprocessing and cleaning step to obtain the cleaned first pdf text data. Take the text data that is read abnormally as the second type of text data and process it using a postprocessing and cleaning step to obtain the cleaned second pdf text data;
[0066] The following 3 examples are the attribute data of 3 random lines of a certain document. The step details in steps S1 - S7 in this embodiment are explained with the following examples for assistance:
[0067] Example 1:
[0068] {'spans':[{'size':10.0,'flags':4,'font':'STIX-Regular','color':-16777216,'ascender':1.0230000019073486,'descender':-0.4860000014305115,'text':'tionshadhypophosphatemia,suggestingthatthisincreasein','origin':(306.1394958496094,639.3757934570312),'bbox':(306.1394958496094,629.1458129882812,546.7841186523438,644.2357788085938)}],'wmode':0,'dir':(1.0,0.0),'bbox':(306.1394958496094,629.1458129882812,546.7841186523438,644.2357788085938)}。
[0069] Example 2:
[0070] {"spans":[{"size":10.0,"flags":4,"font":"STIX-Regular","color":-16777216,"ascender":1.0230000019073486,"descender":-0.4860000014305115,"text":"FGF23shouldbeintendedinthiscontextmoreasamarker","origin":(306.1394958496094,651.8757934570312),"bbox":(306.1394958496094,641.6458129882812,546.7869873046875,656.7357788085938)}],"wmode":0,"dir":(1.0,0.0),"bbox":(306.1394958496094,641.6458129882812,546.7869873046875,656.7357788085938)}。
[0071] Example 3:
[0072] {"spans":[{"size":10.0,"flags":4,"font":"STIX-Regular","color":-16777216,"ascender":1.0230000019073486,"descender":-0.4860000014305115,"text":"ofdiseaseprogressionratherthanaTIO-relatedcondition.","origin":(306.1394958496094,664.3757934570312),"bbox":(306.1394958496094,654.1458129882812,542.1321411132812,669.2357788085938)}],"wmode":0,"dir":(1.0,0.0),"bbox":(306.1394958496094,654.1458129882812,542.1321411132812,669.2357788085938)}。
[0073] In some embodiments, the pre-treatment cleaning step includes:
[0074] Step S1: Use the python toolkit to read the first type of text data line by line, and obtain the corresponding attribute values for each line. The attribute values include color, text content, and horizontal and vertical coordinate values.
[0075] Among them, the color attribute in the 3 samples is 'color'. There may be a certain gap between non-text content and text content (related to the specific file). The text content is 'text', which contains the text content of this recognized text block. The horizontal and vertical coordinates are included in the 'bbox' attribute. When the text is recognized by the fitz package, a bounding box will be obtained to get the location of the text of this line. This attribute contains four values, which respectively represent the specific xy positions of the upper left vertex and the lower right vertex of the bounding box of this text block.
[0076] Step S2: Select the text vertical coordinate value of any line except the first line in the first type of text data, and subtract the text coordinate value of its previous line to obtain the line spacing.
[0077] Among them, the second value in the bbox attribute of each line is the y value of the upper left vertex of the bounding box, which is the upper boundary of this text block. The calculation of the line spacing is to record the upper boundary of the text block with a different y value last time, and subtract the boundary value of the previous text block from the boundary value of this text block to obtain the line spacing between two lines.
[0078] Step S3: Judge the number of pages of the first type of text data. If the number of pages is less than 5 pages, directly calculate the first mode of the line spacing. If the number of pages is greater than or equal to 5 pages, delete the content of the first 20% and the last 20% of the pages, then obtain the line spacing, and calculate the first mode of the line spacing. Define the calculated first mode as the body text line spacing.
[0079] Among them, if the number of pages is too large, it is very likely to be the introduction of the first page or too many references at the end, which will affect the calculation of the mode of the body text line spacing. Therefore, deleting the last 20% and the first 20% eliminates the influence of useless data on the extraction of the mode of the body text line spacing.
[0080] Step S4: Statistically calculate the mode of the body text font size in the first type of text data as the second mode.
[0081] Among them, when obtaining the line attributes, the'size' attribute will be obtained. This attribute refers to the text font size of this text block. It may have different sizes in different text blocks. If different, take the average value, save this value, and statistically calculate the mode to obtain the mode of the font size in the body text. If the sum of the two fonts is less than 60% of the full text, it is considered that the font format of this document is chaotic and cannot be recognized.
[0082] Step S5: Determine whether the first type of text data is a Chinese document through function code. If the document is in Chinese, set the error value of the line spacing of the main text to 1.5. If the document is in English, set the error value of the line spacing of the full text to 0.5. Process the first type of text data according to the error value of the line spacing of the full text to obtain the PDF text data to be cleaned;
[0083] Among them, the Chinese content is set to 1.5, and the English content is recognized as 0.5. When set to 1.5, it can just cover the occasionally appearing English letters, and improve the recognition accuracy with the smallest possible error value. For English, 0.5 just covers the line spacing of English capital letters, covering both English uppercase and lowercase. Further, the difference between Chinese and lowercase English in common texts is about 1.5, and the difference between lowercase and uppercase English is about 0.5. Through testing and practice, try multiple error values to obtain the most suitable error value.
[0084] The function code for judging a Chinese document or an English document is:
[0085] def contains_chinese(s):
[0086] for ch in s:
[0087] if '\u4e00’ <= ch <= '\u9fa5':
[0088] return True:
[0089] return False;
[0090] Step S6: Traverse the PDF text data to be cleaned, screen out the main text lines according to the line spacing of the main text, clear the text data of non-main text lines, and at the same time judge whether the font meets the second mode standard. If not, clear it to obtain the preliminary cleaning result;
[0091] Among them, it is judged according to the line spacing between the current line and the previous line. Therefore, when judging the first line of a page or the first line of a paragraph, this line needs to be supplemented as the main text. Therefore, when judging that the current line is the main text, if the previous line is not, add the previous line;
[0092] Judge whether the font is the mode of the font size: The mode of the document font calculated and saved before is used as the font size of the main text. Through comparison, if they are equal, it is judged that the font of this line meets the main text rules, preventing incorrect addition of the main text.
[0093] Step S7: Use regular expressions to remove keywords, summaries, superscripts, and references in the preliminary cleaning result to obtain the first cleaned PDF text data.
[0094] Example:
[0095] patterns = ['Acknowledgements','References','references1','publisher'snote','supplementarymaterialsinteressenkonflikt','endorsingorganisation','[references]','references[1]',supplementaryinformation'.'orcidids','supplementalmaterial(s)','[references]','committee list (sorted by surname in pinyin)','conflicts of interest','co-authored',,'statement of conflict of interest','[references]','expert list (sorted by surname in pinyin)','list of experts in discussion:','references]','literatur','references','REFERENCES]].
[0096] Since the initial formats of PDFs are different, there are very few PDF files with special formats that cannot be read normally using the PDF package. The read text is garbled, which is considered a reading anomaly. For this reason, a post-processing method has been developed to delete these abnormal documents to prevent interference with the model.
[0097] In some embodiments, the post-processing cleaning step comprises:
[0098] Traverse each line of the second type of text data, and define the lines with less than 5 Chinese characters or lowercase letters as garbled lines. If the proportion of garbled lines to the total number of lines of the full text is greater than 0.5, the PDF text file is defined as a garbled file. At this time, directly delete the current garbled file, and use the remaining text data after deletion and cleaning as the second PDF text data.
[0099] The function code is:
[0100] def is_garbled_line(text:str)-> bool:
[0101] ##Judge whether a single line of text is garbled. If there are less than 5 readable characters, it is considered a garbled line.
[0102] readable_chars = sum(1 for ch in text if '\u4e00' <= ch <= '\u9fa5'or ('a' <= ch <= 'z'))return readable_chars<5.
[0103] ## If the number of readable characters is less than 5, the line is considered a garbled line.
[0104] In some embodiments, the post-processing cleaning step further includes:
[0105] Analyze whether the file size of the second type of text data is less than or equal to 3 kb. If so, directly identify it as a garbled file and delete it, and use the remaining text data after deletion and cleaning as the second pdf text data.
[0106] Furthermore, when the number of Chinese characters and lowercase letters in a line is less than 5, then the line is mostly symbols or garbled characters, so identify the line as a garbled line. Of course, there will also be errors. For example, there may be less than 5 readable characters at the end of a paragraph. So add the following function: when it is recognized that the garbled lines account for half of the total number of code lines, then identify the document as a garbled file and exit in advance.
[0107] S400: Integrate the first pdf text data and the second pdf text data to obtain the final cleaned text data.
[0108] In a second aspect, the present application proposes a data collection and cleaning system for a medical specialty large model, as Figure 2 shown, including a specialty data collection module, a preliminary processing module, a cleaning module, and an integration module;
[0109] The specialty data collection module is used to determine the scope of target specialty disease types, and collect core medical data according to the scope of target specialty disease types. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specialty, laws and regulations corresponding to the drugs, knowledge graphs corresponding to the specialty, typical cases, papers, and clinical trials;
[0110] The preliminary processing module is used to use a non-picture and selected-text pdf format document as the data collection standard for the core medical data, and obtain the to-be-processed pdf text data according to the data collection standard;
[0111] The cleaning module is used to analyze the to-be-processed pdf text data by using a python toolkit, regard the text data that is read normally as the first type of text data and perform pre-processing cleaning steps to obtain the cleaned first pdf text data; regard the text data that is read abnormally as the second type of text data and perform post-processing cleaning steps to obtain the cleaned second pdf text data;
[0112] The integration module is used to integrate the first pdf text data and the second pdf text data to obtain the final cleaned text data.
[0113] In some embodiments, the cleaning module includes a pre - processing cleaning unit, and the pre - processing cleaning unit is used to perform pre - processing cleaning steps, and the pre - processing cleaning steps include:
[0114] Step S1: Use a python toolkit to read the first - type text data line by line, and obtain the corresponding attribute values for each line. The attribute values include color, text content, and horizontal and vertical coordinate values.
[0115] Step S2: Select the text vertical coordinate value of any line except the first line in the first - type text data, and subtract the text coordinate value of its previous line to obtain the line spacing.
[0116] Step S3: Judge the number of pages of the first - type text data. If the number of pages is less than 5, directly calculate the first mode of the line spacing. If the number of pages is greater than or equal to 5, delete the content of the first 20% and the last 20% of the pages, then obtain the line spacing and calculate the first mode of the line spacing. Define the calculated first mode as the body text line spacing.
[0117] Step S4: Statistically calculate the mode of the body text font size in the first - type text data as the second mode.
[0118] Step S5: Use function code to judge whether the first - type text data is a Chinese document. If the document is in Chinese, set the error value of the body text line spacing to 1.5. If the document is in English, set the error value of the full - text line spacing to 0.5. Process the first - type text data according to the error value of the full - text line spacing to obtain the to - be - cleaned pdf text data.
[0119] Step S6: Traverse the to - be - cleaned pdf text data, screen out the body text lines according to the body text line spacing, clear the text data of non - body text lines, and at the same time judge whether the font meets the second mode standard. If it does not meet the standard, clear it to obtain a preliminary cleaning result.
[0120] Step S7: Use regular expressions to remove keywords, summaries, superscripts, and references in the preliminary cleaning result to obtain the cleaned first pdf text data.
[0121] In some embodiments, the cleaning module includes a post - processing cleaning unit, and the post - processing cleaning unit is used to perform post - processing cleaning steps, and the post - processing cleaning steps include:
[0122] Traverse each line of the second - type text data, and define the lines with less than 5 Chinese characters or lowercase letters as garbled lines. If the proportion of garbled lines in the total number of lines of the full text is greater than 0.5, define the pdf text file as a garbled file, and directly delete the current garbled file at this time.
[0123] In some embodiments, the cleaning module includes a determination and deletion unit, which is configured to analyze whether the file size of the second type of text data is less than or equal to 3 kb. If so, it is directly determined as a garbled file and deleted.
[0124] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0125] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0127] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0129] In the embodiments provided in the present disclosure, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0130] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0131] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0132] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present disclosure, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0133] The above is only the preferred embodiment of the present invention. It should be noted that for those skilled in the art, without departing from the technical solution of the present invention, several modified and improved technical solutions should also be regarded as falling within the scope protected by this solution.
Claims
1. A method for data collection and cleaning of a large medical model for specific diseases, characterized in that: Including the following steps: Determine the scope of target specific diseases, and collect core medical data according to the scope of target specific diseases. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specific diseases, laws and regulations corresponding to the drugs, knowledge graphs corresponding to the specific diseases, typical cases, papers and clinical trials; Use a non-picture and selected text PDF format document as the data collection standard for the core medical data, and obtain the PDF text data to be processed according to the data collection standard; Use a Python toolkit to analyze the PDF text data to be processed. Take the text data that is read normally as the first type of text data and process it using a pre-processing cleaning step to obtain the cleaned first PDF text data; take the text data that is read abnormally as the second type of text data and process it using a post-processing cleaning step to obtain the cleaned second PDF text data; Integrate the first PDF text data and the second PDF text data to obtain the final cleaned text data.
2. The method according to claim 1, wherein: The pre-processing cleaning step includes: Step S1: Use a Python toolkit to read the first type of text data line by line, and obtain the corresponding attribute values for each line. The attribute values include color, text content, and horizontal and vertical coordinate values; Step S2: Select the vertical coordinate value of the text of any line except the first line in the first type of text data, and subtract the vertical coordinate value of the previous line to obtain the line spacing; Step S3: Judge the number of pages of the first type of text data. If the number of pages is less than 5 pages, directly calculate the first mode of the line spacing; if the number of pages is greater than or equal to 5 pages, delete the content of the first 20% and the last 20% of the pages and then obtain the line spacing, and calculate the first mode of the line spacing. Define the calculated first mode as the body text line spacing; Step S4: Statistically calculate the mode of the body text font size of the first type of text data as the second mode; Step S5: Use function code to judge whether the first type of text data is a Chinese document. If the document is in Chinese, set the error value of the body text line spacing to 1.5; if the document is in English, set the error value of the full text line spacing to 0.
5. Process the first type of text data according to the error value of the full text line spacing to obtain the PDF text data to be cleaned; Step S6: Traverse the PDF text data to be cleaned, screen out the body text lines according to the body text line spacing, clear the text data of the non-body text lines, and at the same time judge whether the font meets the second mode standard. If it does not meet the standard, clear it to obtain a preliminary cleaning result; Step S7: Use regular expressions to remove keywords, summaries, superscripts, and references in the preliminary cleaning result to obtain the cleaned first PDF text data.
3. The method according to claim 2, characterized in that: The post-processing cleaning step includes: Traverse each line of the second type of text data, and define the lines with less than 5 Chinese characters or lowercase letters as garbled lines. If the proportion of garbled lines in the total number of lines of the full text is greater than 0.5, define the PDF text file as a garbled file, and directly delete the current garbled file at this time.
4. The method according to claim 3, wherein: The post-processing cleaning step also includes: Analyze whether the file size of the second type of text data is less than or equal to 3 kb. If so, directly identify it as a garbled file and delete it.
5. A data collection and cleaning system for a large medical model specialized in a particular disease, characterized in that: It includes a specialized disease data collection module, a preliminary processing module, a cleaning module, and an integration module; The specialized disease data collection module is used to determine the scope of target specialized disease types, and collect core medical data according to the scope of target specialized disease types. The core medical data includes medical guidelines, expert consensus, textbooks, online Q&A, drug instructions corresponding to the specialized disease, laws and regulations corresponding to the drug, knowledge graph corresponding to the specialized disease, typical cases, papers, and clinical trials; The preliminary processing module is used to use a non-picture and selected text pdf format document as the data collection standard for the core medical data, and obtain the pdf text data to be processed according to the data collection standard; The cleaning module is used to analyze the pdf text data to be processed by using a python toolkit, take the text data that is read normally as the first type of text data and perform pre-processing cleaning steps to obtain the cleaned first pdf text data; take the text data that is read abnormally as the second type of text data and perform post-processing cleaning steps to obtain the cleaned second pdf text data; The integration module is used to integrate the first pdf text data and the second pdf text data to obtain the final cleaned text data.
6. The system according to claim 5, characterized in that: The cleaning module includes a pre-processing cleaning unit, and the pre-processing cleaning unit is used to execute pre-processing cleaning steps. The pre-processing cleaning steps include: Step S1: Use a python toolkit to read the first type of text data line by line to obtain the corresponding attribute values for each line. The attribute values include color, text content, and horizontal and vertical coordinate values; Step S2: Select the text vertical coordinate value of any line except the first line in the first type of text data, subtract the text coordinate value of the previous line to obtain the line spacing; Step S3: Judge the number of pages of the first type of text data. If the number of pages is less than 5 pages, directly calculate the first mode of the line spacing; if the number of pages is greater than or equal to 5 pages, delete the content of the first 20% and the last 20% of the pages and then obtain the line spacing, and calculate the first mode of the line spacing. Define the calculated first mode as the body text line spacing; Step S4: Statistically calculate the mode of the body text font size of the first type of text data as the second mode; Step S5: Use function code to judge whether the first type of text data is a Chinese document. If the document is in Chinese, set the error value of the body text line spacing to 1.
5. If the document is in English, set the error value of the full text line spacing to 0.
5. Process the first type of text data according to the error value of the full text line spacing to obtain the pdf text data to be cleaned; Step S6: Traverse the pdf text data to be cleaned, screen out the body text lines according to the body text line spacing, clear the text data of non-body text lines, and at the same time judge whether the font meets the second mode standard. If it does not meet the standard, clear it to obtain the preliminary cleaning result; Step S7: Use regular expressions to remove keywords, abstracts, superscripts, and references from the preliminary cleaning result to obtain the first cleaned pdf text data.
7. The system according to claim 6, wherein: The cleaning module includes a post-processing cleaning unit for performing post-processing cleaning steps, and the post-processing cleaning steps include: Traverse each line of the second type of text data, define a line with less than 5 Chinese characters or lowercase letters as a garbled line. If the proportion of garbled lines in the total number of lines in the full text is greater than 0.5, define the pdf text file as a garbled file, and directly delete the current garbled file at this time.
8. The system according to claim 7, wherein: The cleaning module includes a determination and deletion unit for analyzing whether the file size of the second type of text data is less than or equal to 3 kb. If so, directly identify it as a garbled file and delete it.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-4.