Medical information classification method and system based on AI

Through the AI-based medical information classification method, the heterogeneity of multi-source physical examination data is solved, the automated processing of data cleaning and labeling models is realized, classification efficiency and diagnostic support capabilities are improved, and personalized medical and resource management are promoted.

CN120277482APending Publication Date: 2025-07-08BEIJING R&W ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326960.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Diversity and heterogeneity of medical information makes it difficult to effectively classify, affecting diagnosis, treatment decisions, personalized medical care, data security and resource management.

Method used

Using AI-based medical information classification method, through data cleaning and preprocessing, the physical examination data is classified and annotated using text and image annotation models, a unified standard template is constructed, and an artificial neural network is used for automatic annotation and classification.

Benefits of technology

It realizes unified formatting and efficient automatic classification of multi-source physical examination data, improves classification efficiency and effectiveness, and supports rapid diagnosis, personalized medical care and resource optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277482A_ABST
    Figure CN120277482A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to an AI-based medical information classification method and system, and aims to unify the format of multi-source data by collecting physical examination data of a plurality of data sources and cleaning the physical examination data of the plurality of data sources, and perform complementation, duplicate removal, obvious error correction and the like on the multi-source data. Then, an AI model is utilized to classify the preprocessed data, the AI model classifies the physical examination data by automatically recognizing numerical values, texts and images of the data, and therefore automatic classification on multiple classification dimensions is achieved. Besides, the formats of the multi-source data are subjected to unified processing, so that the data required by the target dimension can be extracted according to the position, and then automatic labeling is performed. The method has the advantages of high classification efficiency and good classification effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and specifically to an AI-based medical information classification method and system. Background Art

[0002] Medical information widely covers various data related to personal health conditions and medical service provision. Reasonable classification of medical information can help to quickly obtain relevant information, support more accurate diagnosis and treatment decisions, promote personalized medicine, enhance data security and privacy protection, support research and education, and optimize resource management.

[0003] However, there is a wide variety of medical information. Especially for physical examination data from different sources, their semantic expressions, content compositions, etc. are all different, so it is difficult to classify them. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an AI-based medical information classification method and system to solve the problems in the background art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] An AI-based medical information classification method of the present invention includes the steps of:

[0007] Obtain physical examination data from multiple data sources, where the physical examination data includes text data, table data, and picture data;

[0008] Clean the medical information to obtain preprocessed physical examination data, where the data cleaning methods include filling in missing data, deleting duplicate data, and correcting obvious errors;

[0009] Based on a pre-constructed AI annotation model, annotate the preprocessed physical examination data to obtain classification labels of the preprocessed physical examination data on multiple classification dimensions, where the AI annotation model includes a text annotation model and a picture annotation model;

[0010] Classify all the physical examination data based on the classification labels.

[0011] In an embodiment of the present application, cleaning the medical information to obtain preprocessed physical examination data includes:

[0012] Extract keywords from the text data, and perform type annotation on the data in each cell of the table data based on a pre-constructed type annotation template, where the type annotation template is constructed based on the data source;

[0013] Map the keyword to the corresponding position of the standard template based on the meaning of the keyword, map the tabular data to the standard template based on the data type in each cell, and map the picture data to the corresponding position of the standard template;

[0014] Fill in the missing data, delete the duplicate data, and correct the obvious errors in the standard template.

[0015] In one embodiment of the present application, label the preprocessed physical examination data based on a pre-constructed text annotation model to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, including:

[0016] Extract the target data related to the current classification dimension from the preprocessed physical examination data, and perform vectorization processing on the target data to obtain a target vector;

[0017] Input the target vector into the text annotation model of the current dimension to obtain a classification label.

[0018] In one embodiment of the present application, label the preprocessed physical examination data based on a pre-constructed image annotation model to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, including:

[0019] Grayscale and high-pass filter the image data to obtain a preprocessed image;

[0020] Extract the contour of the preprocessed image to obtain the contour features of the preprocessed image;

[0021] Input the contour features into a pre-constructed image annotation model to obtain a classification label.

[0022] In one embodiment of the present application, the construction method of the text annotation model includes:

[0023] Obtain physical examination sample data;

[0024] Split the training samples of multiple classification dimensions from the physical examination sample data, and label the training samples to obtain training data;

[0025] Train an artificial neural network based on the training data to obtain a text annotation model.

[0026] In one embodiment of the present application, the construction method of the image annotation model includes:

[0027] Obtain physical examination image sample data;

[0028] Grayscale and high-pass filter the physical examination image sample data to obtain a preprocessed image sample;

[0029] Extract the contours of the preprocessed image samples to obtain the contour features of the preprocessed image samples;

[0030] Use the contour features of the preprocessed image samples as image training samples, and label the image training data to obtain image training data;

[0031] Train an artificial neural network based on the image training data to obtain an image annotation model.

[0032] In an embodiment of the present application, the classification dimensions include user personal information, disease types, and disease severity.

[0033] The present application also provides an AI-based medical information classification system, including:

[0034] An acquisition module for acquiring physical examination data from multiple data sources, where the physical examination data includes text data, table data, and picture data;

[0035] A data cleaning module for cleaning the medical information to obtain preprocessed physical examination data, where the data cleaning methods include filling in missing data, deleting duplicate data, and correcting obvious errors;

[0036] A labeling module for labeling the preprocessed physical examination data based on a pre-constructed AI labeling model to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, where the AI labeling model includes a text labeling model and a picture labeling model;

[0037] A classification module for classifying all the physical examination data based on the classification labels.

[0038] The present application also provides an electronic device, including: a processor and a memory;

[0039] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes the method described above.

[0040] The present application also provides a computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, the method described above is implemented.

[0041] The beneficial effects of the present invention are as follows: A medical information classification method and system based on AI of the present invention collect physical examination data from multiple data sources, clean the physical examination data from multiple sources, so as to unify the formats of multi-source data, and supplement, de-duplicate and correct obvious errors in the multi-source data. Then, an AI model is used to classify the pre-processed data. The AI model classifies the physical examination data by automatically identifying the numerical values, texts and images of the data, so as to achieve automatic classification in multiple classification dimensions. In addition, since the present application first performs a unified processing on the formats of multi-source data, the data required for the target dimension can be extracted according to the position, and then automatic annotation can be performed. The present application has the advantages of high classification efficiency and good classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The present invention will be further described below in conjunction with the drawings and embodiments:

[0043] Figure 1 is a flowchart of a medical information classification method based on AI shown in an embodiment of the present application;

[0044] Figure 2 is a structural diagram of a medical information classification system based on AI shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0046] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the layers related to the present invention are shown in the drawings, rather than being drawn according to the number, shape and size of the layers in actual implementation. The type, quantity and ratio of each layer in actual implementation can be an arbitrary change, and the layer layout type may also be more complex.

[0047] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details.

[0048] Figure 1 is a flowchart of a medical information classification method based on AI shown in an embodiment of the present application, as Figure 1As shown in the figure, a medical information classification method based on AI in this embodiment may include the steps:

[0049] S110, Obtain physical examination data from multiple data sources, where the physical examination data includes text data, table data, and picture data;

[0050] For physical examination data from different data sources, due to different physical examination methods, report format requirements, and data processing methods of different physical examination units, when facing multi-source physical examination data, there will be a problem of being difficult to divide.

[0051] S120, Perform data cleaning on the medical information to obtain preprocessed physical examination data, where the data cleaning methods include filling in missing data, deleting duplicate data, and correcting obvious errors;

[0052] In this application, unified processing and cleaning are performed on multi-source physical examination data, so as to unify the formats of all data, and at the same time perform cleaning on the data, so as to standardize and unify the multi-source physical examination data.

[0053] The specific process includes:

[0054] S121, Extract keywords from the text data, and perform type annotation on the data in each cell of the table data based on a pre-constructed type annotation template, where the type annotation template is constructed based on the data source;

[0055] Among them, the text data in the physical examination report generally includes personal information, physical examination item descriptions, test result descriptions, etc. For the above three types of content, this application simplifies the text data by extracting keywords;

[0056] For example, the doctor's comment is: Your red blood cell count is slightly low. It is recommended to increase the intake of iron-rich foods and maintain good living habits. All other indicators are within the normal range.

[0057] The extracted keywords include: increasing iron-rich foods, good living habits.

[0058] Specifically, this application uses natural language processing technology to extract keywords.

[0059] S122, Map the keywords to the corresponding positions in the standard template based on the meanings of the keywords, map the table data to the standard template based on the data type in each cell, and map the picture data to the corresponding positions in the standard template;

[0060] The content in the table is generally the values of various physical indicators. Therefore, as long as the data type of each cell is known, the data in the table can be transferred to a unified template.

[0061] For example, the data type of the table header is health indicator A, and the numbers in the horizontal or vertical cells are the values of health indicator A. By constructing the statistical templates of each data source, the values in the table can be quickly transferred to the standard template.

[0062] S123. Fill in the missing data, delete the duplicate data, and correct the obvious errors in the standard template.

[0063] In the standard template, fill in the missing values with default values (such as 0 or empty strings), or fill in the missing values of numerical data with the mean, median, or mode. Detect and delete duplicate data entries through unique identifiers or other methods. Correct the obvious error data points according to the domain knowledge. For example, the age is negative or exceeds the human life span.

[0064] S130. Label the preprocessed physical examination data based on the pre-constructed AI annotation model to obtain the classification labels of the preprocessed physical examination data in multiple classification dimensions, where the AI annotation model includes a text annotation model and an image annotation model;

[0065] In this application, automatic annotation is achieved by constructing a text annotation AI model and an image annotation AI model. The text model mainly classifies numerical and character type data. For example, classify the physical examination data according to the user's personal information, disease type, and disease severity. The image annotation AI model classifies the photos for the physical examination types with CT and ultrasound examinations, so as to label the abnormal types of the user.

[0066] The construction method of the text annotation model includes:

[0067] (1) Obtain the physical examination sample data;

[0068] In this application, the physical examination sample data can be obtained from public medical databases, literature materials, or through cooperative medical institutions.

[0069] (2) Split the training samples of multiple classification dimensions from the physical examination sample data, and label the training samples to obtain the training data;

[0070] Since multiple-dimensional division of the physical examination data is required in this application, it is necessary to split the training samples of multiple classification dimensions from the physical examination sample data. For example, for gallbladder diseases, relevant physical examination parameters such as liver function parameters, white blood cell count, and C-reactive protein need to be extracted.

[0071] After extracting the training samples, label them according to the conclusion data in the physical examination sample data.

[0072] In addition, it is also necessary to process the training data into vectors and convert the text into a form that can be understood by machine learning algorithms. Common methods include the Bag of Words model, TF-IDF (Term Frequency-Inverse Document Frequency), and Word2Vec, etc.

[0073] (3) Train an artificial neural network based on the training data to obtain a text annotation model.

[0074] When training the model, divide the data set into a training set, a validation set, and a test set to avoid overfitting and evaluate the model performance. Use metrics such as accuracy, precision, recall, and F1 score to evaluate the performance of the model. Once the model reaches a satisfactory performance level, the required classification model can be obtained.

[0075] The method for constructing the image annotation model includes:

[0076] (1) Obtain physical examination image sample data;

[0077] In this application, the physical examination sample data can be obtained from public medical databases, literature materials, or through cooperative medical institutions.

[0078] (2) Grayscale and high-pass filter the physical examination image sample data to obtain preprocessed image samples;

[0079] Grayscale processing can reduce the information dimension in medical images, and high-pass filtering can retain high-frequency information and filter out low-frequency information with lower importance.

[0080] (3) Extract the contours of the preprocessed image samples to obtain the contour features of the preprocessed image samples;

[0081] Common methods for contour extraction include the Sobel operator, Prewitt operator, or Canny algorithm, etc. These methods determine the position of the edge by calculating the gradient of the image.

[0082] After extracting the contours, it is also necessary to delimit the regions of interest (such as organs, lesions, etc.) and use them as training data.

[0083] (4) Use the contour features of the preprocessed image samples as image training samples and label the image training data to obtain image training data;

[0084] Similarly, after extracting the training samples, they can be labeled according to the conclusion data in the physical examination sample data.

[0085] (5) Train an artificial neural network based on the image training data to obtain an image annotation model.

[0086] The training principle is similar to the above process and will not be elaborated here.

[0087] Based on the above model, the process of automatic classification includes:

[0088] (1) Extract target data related to the current classification dimension from the preprocessed physical examination data, and perform vectorization processing on the target data to obtain a target vector;

[0089] (2) Input the target vector into the text annotation model of the current dimension to obtain a classification label.

[0090] (3) Grayscale and high-pass filter the image data to obtain a preprocessed image;

[0091] (4) Extract the contour of the preprocessed image to obtain the contour features of the preprocessed image;

[0092] (5) Input the contour features into a pre-constructed image annotation model to obtain a classification label.

[0093] S140, classify all physical examination data based on the classification label.

[0094] After obtaining the classification label, it is possible to conveniently and quickly classify the physical examination data based on the classification label. Through an automated classification system, doctors can quickly identify abnormal results or patient groups that require special attention. This helps to shorten the diagnosis time and enables doctors to formulate treatment plans more quickly. Provide customized health advice and preventive measures for individuals according to the classification results of physical examination data. For example, provide specific lifestyle adjustment suggestions for people at risk of cardiovascular diseases. By effectively classifying physical examination data, it is easier to identify early signs of diseases. This is especially important for many chronic diseases (such as diabetes, hypertension, etc.) because early intervention can usually significantly improve the prognosis. In addition, medical institutions can reasonably arrange medical resources, including manpower, equipment, and drugs, according to the number and needs of different categories of patients, thereby improving service efficiency and service quality.

[0095] A medical information classification method based on AI of the present invention collects physical examination data from multiple data sources, cleans the physical examination data from multiple sources, so as to unify the formats of multi-source data, and perform complementing, de-duplicating, and obvious error correction on the multi-source data. Then, an AI model is used to classify the preprocessed data. The AI model classifies the physical examination data by automatically identifying the numerical values, texts, and images of the data, so as to achieve automatic classification in multiple classification dimensions. In addition, since the present application first performs unified processing on the formats of multi-source data, the data required for the target dimension can be extracted according to the position, and then automatic annotation can be performed. The present application has the advantages of high classification efficiency and good classification effect.

[0096] As Figure 2 shown, the present application further provides a medical information classification system based on AI, including:

[0097] An acquisition module, configured to acquire physical examination data from multiple data sources, wherein the physical examination data includes text data, table data, and picture data;

[0098] A data cleaning module, configured to perform data cleaning on the medical information to obtain preprocessed physical examination data, wherein the data cleaning method includes filling in missing data, deleting duplicate data, and correcting obvious errors;

[0099] A marking module, configured to mark the preprocessed physical examination data based on a pre-constructed AI marking model to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, wherein the AI marking model includes a text marking model and a picture marking model;

[0100] A classification module, configured to classify all physical examination data based on the classification labels.

[0101] A medical information classification system based on AI of the present invention collects physical examination data from multiple data sources, cleans the physical examination data from multiple sources, so as to unify the formats of multi-source data, and perform complementing, de-duplicating, and obvious error correction on the multi-source data. Then, an AI model is used to classify the preprocessed data. The AI model classifies the physical examination data by automatically identifying the numerical values, texts, and images of the data, so as to achieve automatic classification in multiple classification dimensions. In addition, since the present application first performs unified processing on the formats of multi-source data, the data required for the target dimension can be extracted according to the position, and then automatic annotation can be performed. The present application has the advantages of high classification efficiency and good classification effect.

[0102] This embodiment further provides an electronic terminal, including: a processor and a memory;

[0103] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the terminal executes any of the methods in this embodiment.

[0104] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to computer programs. The foregoing computer program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0105] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used for communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal executes each step of the above method.

[0106] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0107] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processing, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0108] In the above embodiments, although the present invention has been described in conjunction with specific embodiments of the present invention, according to the foregoing description, many substitutions, modifications, and variations of these embodiments will be obvious to those of ordinary skill in the art. The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.

[0109] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A medical information classification method based on AI, characterized in that Including the steps: Obtain physical examination data from multiple data sources, where the physical examination data includes text data, tabular data, and picture data; Perform data cleaning on the medical information to obtain preprocessed physical examination data, where the data cleaning methods include filling in missing data, deleting duplicate data, and correcting obvious errors; Based on a pre-constructed AI annotation model, annotate the preprocessed physical examination data to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, where the AI annotation model includes a text annotation model and a picture annotation model; Classify all the physical examination data based on the classification labels.

2. The method for classifying medical information based on AI according to claim 1, wherein Perform data cleaning on the medical information to obtain preprocessed physical examination data, including: Extract keywords from the text data, and perform type annotation on the data in each cell of the tabular data based on a pre-constructed type annotation template, where the type annotation template is constructed based on the data source; Map the keywords to the corresponding positions in the standard template based on the meanings of the keywords, map the tabular data to the standard template based on the data types in each cell, and map the picture data to the corresponding positions in the standard template; Fill in missing data, delete duplicate data, and correct obvious errors in the standard template.

3. The method for classifying medical information based on AI according to claim 1, characterized in that, Based on a pre-constructed text annotation model, annotate the preprocessed physical examination data to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, including: Extract target data related to the current classification dimension from the preprocessed physical examination data, and perform vectorization processing on the target data to obtain a target vector; Input the target vector into the text annotation model of the current dimension to obtain a classification label.

4. The method for classifying medical information based on AI according to claim 1, wherein, Based on a pre-constructed image annotation model, annotate the preprocessed physical examination data to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, including: Gray-scale the image data and perform high-pass filtering to obtain a preprocessed image; Extract the contours of the preprocessed image to obtain the contour features of the preprocessed image; Input the contour features into a pre-constructed image annotation model to obtain a classification label.

5. The medical information classification method based on AI according to claim 1, characterized in that, The construction method of the text annotation model includes: Obtain physical examination sample data; Split training samples of multiple classification dimensions from the physical examination sample data, and annotate the training samples to obtain training data; Train an artificial neural network based on the training data to obtain a text annotation model.

6. The method for classifying medical information based on AI according to claim 1, characterized in that, The construction method of the image annotation model includes: Obtain physical examination image sample data; Gray-scale the physical examination image sample data and perform high-pass filtering to obtain a preprocessed image sample; Extract the contours of the preprocessed image sample to obtain the contour features of the preprocessed image sample; Use the contour features of the preprocessed image sample as image training samples, and annotate the image training data to obtain image training data; Train an artificial neural network based on the image training data to obtain an image annotation model.

7. A medical information classification method based on AI according to claim 1, characterized in that, The classification dimensions include user personal information, disease type, and disease severity.

8. An AI-based medical information classification system, characterized in that, Including: An acquisition module for acquiring physical examination data from multiple data sources, wherein the physical examination data includes text data, table data, and picture data; A data cleaning module for cleaning the medical information to obtain preprocessed physical examination data, wherein the data cleaning methods include filling in missing data, deleting duplicate data, and correcting obvious errors; A labeling module for labeling the preprocessed physical examination data based on a pre-constructed AI labeling model to obtain classification labels of the preprocessed physical examination data in multiple classification dimensions, wherein the AI labeling model includes a text labeling model and a picture labeling model; A classification module for classifying all physical examination data based on the classification labels.

9. An electronic device, characterized in that, Comprising: A processor and a memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.