A clinical dataset classification method, electronic device and storage medium

By performing departmental pre-classification and symptom analysis on clinical data, extracting essential and secondary symptoms, and establishing a table to analyze possible misdiagnoses, this method solves the problem that existing clinical dataset classification methods cannot obtain corresponding data in a timely manner, thereby improving the accuracy of doctors' judgments and the effectiveness of treatment.

CN119361050BActive Publication Date: 2026-07-17SHANGHAI MEISI PHARM TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI MEISI PHARM TECH CO LTD
Filing Date
2024-10-25
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing clinical dataset classification methods cannot effectively address the problem that doctors cannot obtain relevant clinical data for reference in a timely manner after contacting patients, especially when the symptoms in the referenced clinical data are consistent with the patient's symptoms but the causes are different, making it impossible to make a timely judgment.

Method used

By acquiring clinical data to be classified, pre-classification is performed based on the corresponding departments. Clinical symptom analysis is used to extract necessary and secondary clinical symptoms, a symptom analysis table is established, misdiagnosable diseases are analyzed, and the clinical data is classified based on this.

Benefits of technology

It enables doctors to effectively obtain data that matches the patient's condition when referring to clinical data, preventing delays or errors in treatment and reducing the frequency of misdiagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119361050B_ABST
    Figure CN119361050B_ABST
Patent Text Reader

Abstract

This invention discloses a clinical dataset classification method, electronic device, and storage medium, relating to the field of data classification technology. The method includes: pre-classifying all clinical data; analyzing the clinical data using a clinical symptom analysis method to obtain necessary and secondary clinical symptoms, establishing a symptom analysis table, identifying potentially misdiagnosable clinical symptoms, and classifying all clinical data. This invention addresses the problem that existing classification methods for clinical datasets cannot effectively solve the issue of doctors being unable to promptly obtain corresponding clinical data for reference based on patient symptoms after contact with the patient, or where the symptoms in the referenced clinical data are consistent with the patient's symptoms but the underlying causes differ, thus hindering timely patient assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data classification technology, specifically to a clinical dataset classification method, electronic device, and storage medium. Background Technology

[0002] Clinical datasets are collected and organized clinical case data used to support research and applications in clinical protocol analysis. Clinical datasets contain various clinical data, such as patient basic information, hospitalization records, drug treatments, and laboratory test results. This data can support various types of clinical protocol analysis studies. Clinical datasets mainly include: safety analysis sets, intention-to-treat analysis sets, and protocol compliance analysis sets. These classifications are primarily to meet different needs in clinical trials, including evaluating treatment efficacy, assessing safety, and maintaining the integrity of the study design.

[0003] Existing methods for classifying clinical datasets typically involve sampling clinical data, adding different category labels to the data, and then training the dataset with labeled clinical data and other clinical data to improve classification performance. While this method can classify clinical data based on different categories, it fails to effectively process the content of the clinical data before classification. This leads to problems where, after doctors encounter patients, they cannot promptly obtain relevant clinical data for reference based on the patient's symptoms, or the referenced clinical data may contain symptoms consistent with the patient's but with different underlying causes, resulting in an inability to make timely judgments. For example, patent application CN110400610A discloses a method based on... This paper presents a small-sample clinical data classification method and system using multi-channel random forests. This approach uses labeled augmented data and clinical data together to form a training dataset, which is then used to train a classifier and improve its classification performance for clinical samples. Other improvements for classifying clinical datasets typically involve classifying different clinical datasets for different clinical trials, and the classification results are usually related to the medications used in the clinical trials. However, this approach still cannot effectively address the problem that doctors cannot obtain relevant clinical data for reference based on patient symptoms after contacting the patient, or that the symptoms in the referenced clinical data are consistent with the patient's symptoms but the etiology is different, thus leading to an inability to make timely judgments about the patient. Therefore, it is necessary to improve existing classification methods for clinical datasets. Summary of the Invention

[0004] This invention aims to at least partially solve one of the technical problems in the prior art by proposing a clinical dataset classification method, electronic device, and storage medium. This addresses the problem that existing classification methods for clinical datasets cannot effectively solve the problem of doctors being unable to obtain corresponding clinical data for reference based on patient symptoms after contacting the patient, or that the symptoms in the referenced clinical data are consistent with the patient's symptoms but the causes are different from the patient's, thus leading to the inability to make timely judgments about the patient.

[0005] To achieve the above objectives, in a first aspect, this application provides a clinical dataset classification method, comprising the following steps:

[0006] Acquire clinical data to be classified, and pre-classify all clinical data based on the departments corresponding to the clinical data. The pre-classification is used to record the departments corresponding to the clinical data as the source departments of the clinical data.

[0007] The clinical symptom analysis method is used to analyze clinical data, and the symptoms of patients in the clinical data are recorded as clinical symptoms. Based on the analysis results, the necessary clinical symptoms and secondary clinical symptoms in the clinical data are obtained.

[0008] Establish a symptom analysis table to identify potentially misdiagnosed symptoms based on clinically necessary and secondary symptoms from clinical data.

[0009] We analyze all clinical symptoms that could be misdiagnosed and categorize all clinical data based on the analysis results.

[0010] Furthermore, clinical symptom analysis methods include:

[0011] Set the number of analyses for all clinical data, and initially set the number of analyses for all clinical data to 0; for any clinical data with a number of analyses of 0, record the department where the clinical data is located as the clinical department;

[0012] All clinical data corresponding to the clinical departments are denoted as Analysis Data FX1 to Analysis Data FX1 respectively. n For any analysis data FX n1 FX data is obtained and analyzed based on text extraction. n1 All patient symptoms are listed and designated as clinical symptoms LB1 to LB1. k Where n represents the type of clinical data and k represents the type of patient symptoms.

[0013] Furthermore, clinical symptom analysis methods also include:

[0014] Establish a Cartesian coordinate system, denoted as the symptom differentiation coordinate system. The Y-axis of the symptom differentiation coordinate system is in units of quantity, and the coordinates to the right of the origin on the X-axis are filled with clinical symptom LB1 to clinical symptom LB1.k Extract patient symptoms from all analyzed data FX, and for any clinical symptom LB k1 Obtain LB containing clinical symptoms k1 The number of data points analyzed is denoted as p1, and the points in the symptom differentiation coordinate system (clinical symptom LB) are also considered. k1 p1) is recorded as the symptom point BZ. k1 ;

[0015] Obtain the disease points BZ corresponding to all clinical symptoms LB, and denote the straight line fitted by all disease points BZ as the disease line. The x-coordinates of the leftmost and rightmost points of the disease line are clinical symptom LB1 and clinical symptom LB1, respectively. k Draw a straight line parallel to the X-axis from the midpoint of the symptom line and denote it as the span dividing line.

[0016] Furthermore, clinical symptom analysis methods also include:

[0017] The symptom points BZ with the largest and smallest ordinates are recorded as the peak point and valley point, respectively. The line connecting the peak point and the valley point is recorded as the symptom crossing line. The symptom crossing line is translated so that the valley point coincides with the origin of the coordinate system. The sine of the angle between the symptom crossing line and the X-axis that is less than 90 degrees is recorded as the sine of the crossing line.

[0018] Furthermore, clinical symptom analysis methods also include:

[0019] Divide the ordinate of the peak point by n and record the peak percentage; when the peak percentage is greater than or equal to the sine of the cross-line, mark the trough point as the clinical data FX. n1 The clinically necessary symptoms are used to mark the abscissa of all symptom points BZ except for the trough points as clinical data FX. n1 Secondary clinical symptoms;

[0020] When the peak percentage is less than the sine value of the span, the abscissa of all symptom points BZ between the span dividing line and the X-axis is marked as clinical data FX. n1 Clinically necessary symptoms are identified, and symptom points BZ whose x-coordinates are not recorded as clinically necessary symptoms are designated as secondary points. The x-coordinates of these secondary points are then labeled as clinical data FX. n1 Secondary clinical symptoms;

[0021] Obtain the clinically necessary and secondary symptoms corresponding to all analysis data FX. When any analysis data FX obtains both clinically necessary and secondary symptoms, adjust the analysis count of analysis data FX to 1.

[0022] Furthermore, clinical symptom analysis methods also include:

[0023] Once the number of analyses for all clinical data is 1, obtain all distinct clinical symptoms (LB) corresponding to all clinical data and put them into set A;

[0024] Create a table with T rows and R columns, denoted as the misjudgment analysis table. In the top row of the misjudgment analysis table, except for the first cell, fill in all the clinical symptoms LB of set A from left to right. In the leftmost column of the misjudgment analysis table, except for the first cell, fill in all the clinical data from top to bottom.

[0025] Obtain two types of tags that can be used within the table, and denote them as follows: 1 and 2;

[0026] Within the misjudgment analysis table: For any given clinical data point, mark the cell where the row containing the clinical data intersects with the column containing all clinically necessary symptoms of that clinical data. 1; Mark the cell where the row containing the clinical data intersects with the column containing all the secondary clinical symptoms of the clinical data as... 2. Based on the clinically necessary and secondary symptoms corresponding to all clinical data, mark the rows containing all clinical data.

[0027] In the misclassification analysis table after labeling with γ1 and γ2: For any clinical symptom LB, the misclassification analysis coefficient corresponding to the clinical symptom LB is obtained using the misclassification analysis algorithm, which is as follows: Where F is the misjudgment analysis coefficient, and f1 is the clinical symptom LB marked as The number of cells is 1, and f2 is the number of clinical symptoms in the column where LB is marked as 1. The number of 2 squares.

[0028] Furthermore, clinical symptom analysis methods also include:

[0029] Obtain the misjudgment analysis coefficients (LBs) for all clinical symptoms; record the LBs of clinical symptoms with a misjudgment analysis coefficient of 1 as major influencing symptoms, the LBs of clinical symptoms with a misjudgment analysis coefficient of 0 as minor influencing symptoms, and the LBs of clinical symptoms with a misjudgment analysis coefficient less than 1 and greater than 0 as mixed influencing symptoms.

[0030] When the column that primarily affects the condition is marked as When the number of cells with a primary effect is 1, the primary effect disease is recorded as the unique disease; when the number of cells with a secondary effect disease is 1, the column containing the secondary effect disease is marked as... When the number of cells in a 2-cell pattern is 1, the secondary disease affecting the disease is recorded as the unique disease.

[0031] The number of major influencing diseases that are not recorded as unique diseases is denoted as A1, the number of minor influencing diseases that are not recorded as unique diseases is denoted as A2, and the number of mixed influencing diseases is denoted as A0; the value of A1 divided by A0 is recorded as the major judgment value, and the value of A2 divided by A0 is recorded as the minor judgment value.

[0032] For any mixed-effect condition, mark the column containing the mixed-effect condition as... The primary self-judgment value is calculated by dividing the number of cells with a value of 1 by the total number of marked cells in the column containing the mixed-effect disease. The cells marked as 1 in the column containing the mixed-effect disease are then considered as... The value of dividing the number of cells in 2 by the number of all marked cells in the column containing the mixed influence of the disease is recorded as the secondary self-judgment value;

[0033] When the primary self-judgment value is greater than or equal to the primary judgment value or the secondary self-judgment value is greater than or equal to the secondary judgment value, the mixed-effect disease is recorded as a disease that can be misjudged.

[0034] All major and minor influencing symptoms that are not recorded as unique symptoms are recorded as potentially misdiagnosable symptoms.

[0035] Furthermore, the possible misdiagnosable symptoms of all clinical conditions were analyzed, and all clinical data were categorized based on the analysis results, including:

[0036] For any given clinical data, if there is no misdiagnosable disease in the clinical disease LB corresponding to the clinical data, the clinical data is classified as a non-misdiagnosable disease type.

[0037] When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are major influencing diseases, the clinical data will be classified as the major disease misdiagnosis type.

[0038] When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are minor diseases, the clinical data will be classified as minor disease misdiagnosis type.

[0039] When the clinical data corresponds to clinical symptoms in the LB (Lesson Intake) and there are symptoms that can be misdiagnosed and symptoms that have mixed effects, the clinical data is classified into primary and secondary symptom misdiagnosis types.

[0040] Secondly, this application provides an electronic device including a processor and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the steps in the method described above are performed.

[0041] Thirdly, this application provides a storage medium on which a computer program is stored, which, when executed by a processor, performs the steps of the method described above.

[0042] The beneficial effects of this invention are as follows: First, the invention acquires clinical data to be classified. Based on the department corresponding to the clinical data, all clinical data are pre-classified. The pre-classification is used to record the department corresponding to the clinical data as the source department of the clinical data. Then, the clinical symptom analysis method is used to analyze the clinical data, and the patient's symptoms in the clinical data are recorded as clinical symptoms. Based on the analysis results, the necessary clinical symptoms and secondary clinical symptoms in the clinical data are obtained. The advantage of this is that it can effectively obtain the clinical symptoms that represent each clinical data and the clinical symptoms that commonly appear in other clinical data, i.e., necessary clinical symptoms and secondary clinical symptoms. This allows for the acquisition of differences between each clinical data, which helps doctors to effectively obtain clinical data that matches the patient's symptoms when using clinical data for reference, preventing problems such as delayed or incorrect treatment.

[0043] This invention also establishes a symptom analysis table to obtain potentially misdiagnosable symptoms of clinical diseases based on clinically necessary and secondary symptoms in clinical data. Finally, it analyzes all potentially misdiagnosable symptoms of clinical diseases and classifies all clinical data based on the analysis results. The advantage of this is that by obtaining potentially misdiagnosable symptoms of clinical diseases and classifying clinical data based on these symptoms, it can ensure that the classified clinical data can assist doctors in diagnosis and treatment while helping them to judge whether there is a possibility of misdiagnosis, thereby reducing the frequency of misdiagnosis by doctors and enabling effective treatment of patients. Attached Figure Description

[0044] Figure 1 This is a block diagram illustrating the principle of the method of the present invention;

[0045] Figure 2 This is a schematic diagram of the disease differentiation coordinate system of the present invention;

[0046] Figure 3 This is a schematic diagram illustrating the acquisition of the cross-line sine value according to the present invention;

[0047] Figure 4 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Example 1, First Aspect, Please refer to Figure 1As shown, this application provides a clinical dataset classification method, including the following steps:

[0050] Step S1: Obtain the clinical data to be classified. Pre-classify all clinical data based on the departments corresponding to the clinical data. Pre-classification is used to record the departments corresponding to the clinical data as the source departments of the clinical data. In the specific implementation process, considering that patients prioritize the department they visit when seeking medical treatment, the pre-classification in this embodiment uses the departments corresponding to the clinical data as the source departments of the clinical data. Thus, the clinical data is initially classified based on the source departments of all clinical data. The source departments may include departments within the hospital such as gastroenterology, urology, otolaryngology, pediatrics, and surgery. In the specific implementation, the clinical data can be pre-classified according to the actual classification needs, such as classification criteria based on the patient's trauma site, the patient's medical treatment time, and the frequency of the patient's symptoms in the clinical data.

[0051] Step S2: Analyze the clinical data using the clinical symptom analysis method, record the patient's symptoms in the clinical data as clinical symptoms, and obtain the necessary and secondary clinical symptoms in the clinical data based on the analysis results; establish a symptom analysis table, and obtain the possible misdiagnosable symptoms of clinical symptoms based on the necessary and secondary clinical symptoms in the clinical data.

[0052] The clinical symptom analysis method includes: step S201, setting the analysis number for all clinical data, with the initial analysis number for all clinical data being 0; for any clinical data with an analysis number of 0, recording the department where the clinical data is located as the clinical department;

[0053] Step S202: Record all clinical data corresponding to the clinical departments as analysis data FX1 to analysis data FX1 respectively. n For any analysis data FX n1 FX data is obtained and analyzed based on text extraction. n1 All patient symptoms are listed and designated as clinical symptoms LB1 to LB1. k Where n represents the type of clinical data and k represents the type of patient symptoms; in the specific implementation process, for example, in a data processing, if all patient symptoms in an analysis data FX include: dizziness, nausea, stomach pain and limb weakness, then the value of k is 4, and clinical symptoms LB1 to clinical symptoms LB4 are dizziness, nausea, stomach pain and limb weakness respectively.

[0054] In the specific implementation process, the "patient symptoms" in the analysis data FX or the patient symptoms filled in by doctors can be identified through text extraction. The text extraction can be trained according to the actual content composition of the analysis data FX before extraction to ensure the effectiveness of the obtained patient symptoms.

[0055] Step S203: Establish a Cartesian coordinate system, denoted as the symptom differentiation coordinate system. The Y-axis of the symptom differentiation coordinate system is in units of quantity, and the coordinates to the right of the origin on the X-axis are sequentially filled with clinical symptom LB1 to clinical symptom LB. k Extract patient symptoms from all analyzed data FX, and for any clinical symptom LB k1 Obtain LB containing clinical symptoms k1 The number of data points analyzed is denoted as p1, and the points in the symptom differentiation coordinate system (clinical symptom LB) are also considered. k1 p1) is recorded as the symptom point BZ. k1 ;

[0056] In the specific implementation process, for example, the coordinate system for differentiating symptoms obtained in a data processing operation, such as... Figure 2 As shown, where k is 5, all symptom points BZ are BZ1 to BZ5, and LB1 to LB5 are clinical symptom points LB1 to LB5. Analysis reveals that the resulting symptom line is LL1, the midpoint of LL1 is DD1, the span dividing line is LL2, the peak point is point BZ2, and the valley point is point BZ4. After separating the peak and valley points, please refer to [the provided text]. Figure 3 As shown, the symptom crossing line is LL3, and the angle between the symptom crossing line and the X-axis that is less than 90 degrees is 45 degrees. The sine value of the crossing line is... If the peak percentage is calculated to be 1, then the peak percentage is greater than... Let LB4, the abscissa of BZ4, be the clinically necessary symptom of clinical data FX, and let LB1, LB2, LB3 and LB5 be the clinically secondary symptom of clinical data FX.

[0057] Step S204: Obtain the symptom points BZ corresponding to all clinical symptoms LB, and denote the straight line fitted by all symptom points BZ as the symptom line. The x-coordinates of the leftmost and rightmost points of the symptom line are clinical symptom LB1 and clinical symptom LB2, respectively. k Draw a straight line parallel to the X-axis from the midpoint of the symptom line and denote it as the span dividing line;

[0058] Step S205: Record the disease point BZ with the largest and smallest ordinate among all disease points BZ as the peak point and valley point respectively; record the line connecting the peak point and valley point as the disease crossing line; translate the disease crossing line so that the valley point coincides with the origin of the coordinate system, and record the sine of the angle between the disease crossing line and the X-axis that is less than 90 degrees as the sine value of the crossing line.

[0059] Step S206: Divide the ordinate of the peak point by n and record the peak percentage value; when the peak percentage value is greater than or equal to the sine value of the cross line, mark the lateral coordinate of the valley point as the clinical data FX. n1The clinically necessary symptoms are used to mark the abscissa of all symptom points BZ except for the trough points as clinical data FX. n1 Secondary clinical symptoms;

[0060] Step S207: When the peak percentage is less than the sine value of the span, mark the abscissa of all symptom points BZ between the span dividing line and the X-axis as clinical data FX. n1 Clinically necessary symptoms are identified, and symptom points BZ whose x-coordinates are not recorded as clinically necessary symptoms are designated as secondary points. The x-coordinates of these secondary points are then labeled as clinical data FX. n1 Secondary clinical symptoms;

[0061] In the specific implementation process, by obtaining the necessary clinical symptoms and the secondary clinical symptoms, the differences between each clinical data can be obtained. This helps doctors to effectively obtain clinical data that matches the patient's condition when using clinical data for reference, and prevents problems such as delayed treatment or incorrect treatment.

[0062] Step S208: Obtain the clinically necessary symptoms and clinically secondary symptoms corresponding to all analysis data FX. When any analysis data FX obtains both clinically necessary symptoms and clinically secondary symptoms, adjust the analysis number of analysis data FX to 1.

[0063] Step S209: After the number of analyses for all clinical data is 1, obtain all distinct clinical symptoms LB corresponding to all clinical data and put them into set A;

[0064] Step S210: Create a table with T rows and R columns, denoted as the misjudgment analysis table. The top row of the misjudgment analysis table, excluding the first cell, is filled with all clinical symptoms (LB) from set A, from left to right. The leftmost column of the misjudgment analysis table, excluding the first cell, is filled with all clinical data, from top to bottom. In practice, T is the number of all clinical data plus 1, and R is the number of all clinical symptoms (LB) plus 1. For example, the misjudgment analysis table constructed in one analysis might look like this: Figure 1 The misjudgment analysis table is shown below;

[0065] Figure 1 Misjudgment Analysis Table

[0066]

[0067] Step S211: Obtain two types of markers that can be used within the table, and denote them as follows: 1 and 2;

[0068] Step S212, in the misjudgment analysis table: For any clinical data, mark the cell where the row containing the clinical data intersects with the column containing all the necessary clinical symptoms of the clinical data. 1; Mark the cell where the row containing the clinical data intersects with the column containing all the secondary clinical symptoms of the clinical data as... 2. Based on the clinically necessary and secondary symptoms corresponding to all clinical data, mark the rows containing all clinical data.

[0069] Step S213, in use 1 and 2. In the misclassification analysis table after labeling: For any clinical symptom LB, the misclassification analysis coefficient corresponding to the clinical symptom LB is obtained using the misclassification analysis algorithm. The misclassification analysis algorithm is as follows: Where F is the misjudgment analysis coefficient, and f1 is the clinical symptom LB marked as The number of cells is 1, and f2 is the number of clinical symptoms in the column where LB is marked as 1. The number of squares with a value of 2;

[0070] In the specific implementation process, for example, during a data processing session, the column containing the clinical symptoms (LB) is marked as... The number of cells with the value 1 is 2, and the column containing the clinical symptom LB is marked as... If the number of cells in a 2 is 1, then the misjudgment analysis coefficient can be calculated to be approximately 0.67. Clinical symptoms LB are denoted as mixed-effect symptoms.

[0071] Step S214: Obtain the misjudgment analysis coefficients corresponding to all clinical symptoms LB; record the clinical symptoms LB with a misjudgment analysis coefficient of 1 as the main influencing symptoms, record the clinical symptoms LB with a misjudgment analysis coefficient of 0 as the minor influencing symptoms, and record the clinical symptoms LB with a misjudgment analysis coefficient less than 1 and greater than 0 as the mixed influencing symptoms.

[0072] Step S215, when the column primarily affecting the symptom is marked as When the number of cells with a primary effect is 1, the primary effect disease is recorded as the unique disease; when the number of cells with a secondary effect disease is 1, the column containing the secondary effect disease is marked as... When the number of cells in a 2-cell pattern is 1, the secondary disease affecting the disease is recorded as the unique disease.

[0073] Step S216: Record the number of major influencing diseases that are not recorded as unique diseases as A1, the number of minor influencing diseases that are not recorded as unique diseases as A2, and the number of mixed influencing diseases as A0; record the value of A1 divided by A0 as the major judgment value, and the value of A2 divided by A0 as the minor judgment value.

[0074] Step S217: For any mixed-effect symptom, mark the column containing the mixed-effect symptom as... The primary self-judgment value is calculated by dividing the number of cells with a value of 1 by the total number of marked cells in the column containing the mixed-effect disease. The cells marked as 1 in the column containing the mixed-effect disease are then considered as... The value of dividing the number of cells in 2 by the number of all marked cells in the column containing the mixed influence of the disease is recorded as the secondary self-judgment value;

[0075] In the specific implementation process, for example, if A1 is 3, A2 is 1, and A0 is 5 in a data processing session, then the primary judgment value can be calculated to be 0.6 and the secondary judgment value to be 0.2. The primary and secondary self-judgment values ​​for mixed-effect diseases are 0.2 and 0.3, respectively. Therefore, through analysis, mixed-effect diseases can be recorded as misdiagnosable diseases. When mixed-effect diseases are recorded as misdiagnosable diseases, it means that mixed-effect diseases may be misdiagnosed as other diseases due to their primary or secondary clinical symptoms, thus making it impossible to treat patients in a timely manner.

[0076] Step S218: When the primary self-judgment value is greater than or equal to the primary judgment value or the secondary self-judgment value is greater than or equal to the secondary judgment value, the mixed-effect disease is recorded as a misjudgmentable disease.

[0077] Step S219: Record all major and minor influencing diseases that were not recorded as unique diseases as misdiagnosable diseases.

[0078] Step S3 involves analyzing all potentially misdiagnosable clinical symptoms and classifying all clinical data based on the analysis results. Step S3 includes:

[0079] Step S301: For any clinical data, if there is no misdiagnosable disease in the clinical disease LB corresponding to the clinical data, classify the clinical data as a non-misdiagnosable disease type.

[0080] Step S302: When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are major influencing diseases, the clinical data is classified as the major disease misdiagnosis type.

[0081] Step S303: When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are minor diseases, the clinical data is classified as a minor disease misdiagnosis type.

[0082] Step S304: When there are misdiagnosable diseases and mixed-effect diseases in the clinical disease LB corresponding to the clinical data, the clinical data is classified as the primary and secondary disease misdiagnosis type. In the specific implementation process, for example, in a data processing, if there are misdiagnosable diseases in all the clinical disease LB corresponding to the clinical data, and all the clinical disease LB are the primary influencing diseases, then the clinical data is classified as the primary misdiagnosis type.

[0083] In this embodiment, the clinical data classified as non-misdiagnosable disease types are all data that can be directly judged based on the patient's symptoms. The clinical data classified as primary disease misdiagnosis type, secondary disease misdiagnosis type, and primary and secondary disease misdiagnosis type are data that may be misdiagnosed due to the patient's symptoms. This helps doctors to judge whether there is a possibility of misdiagnosis when using clinical data for reference, thereby reducing the frequency of doctors' misdiagnosis and enabling effective treatment of patients.

[0084] Example 2, please refer to Figure 4 As shown, Figure 4 A schematic diagram of an electronic device is provided, which may include a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The memory stores computer-readable instructions, and the processor can call these instructions. When the processor executes a computer-readable instruction, it performs steps similar to those in a clinical dataset classification method to achieve the following functions: First, it acquires the clinical data to be classified. Based on the corresponding department, it pre-classifies all clinical data, using the pre-classification to designate the corresponding department as the source department. Then, it analyzes the clinical data using a clinical symptom analysis method, recording the patient's symptoms as clinical symptoms. Based on the analysis results, it obtains the necessary and secondary clinical symptoms from the clinical data. Furthermore, by establishing a symptom analysis table, it obtains the possible misdiagnosed symptoms based on the necessary and secondary clinical symptoms. Finally, it analyzes all the possible misdiagnosed symptoms and classifies all clinical data based on the analysis results.

[0085] Furthermore, when the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0086] Example 3: This application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a clinical dataset classification method provided by the above methods. The method includes: first, acquiring clinical data to be classified; pre-classifying all clinical data based on the departments corresponding to the clinical data, whereby the pre-classification is used to record the departments corresponding to the clinical data as the source departments of the clinical data; then, analyzing the clinical data using a clinical symptom analysis method, recording the patients' symptoms in the clinical data as clinical symptoms, and obtaining the necessary and secondary clinical symptoms in the clinical data based on the analysis results; furthermore, by establishing a symptom analysis table, obtaining the misdiagnosable symptoms of the clinical symptoms based on the necessary and secondary clinical symptoms of the clinical data; finally, analyzing the misdiagnosable symptoms of all clinical symptoms, and classifying all clinical data based on the analysis results.

[0087] Example 4: This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the steps of the above-described clinical dataset classification method to achieve the following functions: First, it acquires the clinical data to be classified; it pre-classifies all clinical data based on the departments corresponding to the clinical data, with the pre-classification used to record the departments corresponding to the clinical data as the source departments of the clinical data; then, it analyzes the clinical data using a clinical symptom analysis method, recording the patients' symptoms in the clinical data as clinical symptoms, and obtaining the necessary and secondary clinical symptoms in the clinical data based on the analysis results; it also establishes a symptom analysis table, obtaining the misdiagnosable symptoms of clinical symptoms based on the necessary and secondary clinical symptoms of the clinical data; finally, it analyzes the misdiagnosable symptoms of all clinical symptoms and classifies all clinical data based on the analysis results.

[0088] Based on the above description of the embodiments, the embodiments of the present invention can be provided as methods, systems, or computer program products. Based on this understanding, the above technical solutions, in essence or in terms of their contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or certain parts of the embodiments.

[0089] In the embodiments provided in this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces. The indirect coupling or communication connection between systems, modules, and units may be electrical, mechanical, or other forms.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A clinical dataset classification method, characterized in that, Includes the following steps: Acquire clinical data to be classified, and pre-classify all clinical data based on the departments corresponding to the clinical data. The pre-classification is used to record the departments corresponding to the clinical data as the source departments of the clinical data. The clinical symptom analysis method is used to analyze clinical data, and the symptoms of patients in the clinical data are recorded as clinical symptoms. Based on the analysis results, the necessary clinical symptoms and secondary clinical symptoms in the clinical data are obtained. Establish a symptom analysis table to identify potentially misdiagnosed symptoms based on clinically necessary and secondary symptoms from clinical data. Analyze all clinical symptoms that could be misdiagnosed, and classify all clinical data based on the analysis results; Clinical symptom analysis methods include: Set the number of analyses for all clinical data, and initially set the number of analyses for all clinical data to 0; for any clinical data with a number of analyses of 0, record the department where the clinical data is located as the clinical department; All clinical data corresponding to the clinical departments are denoted as Analysis Data FX1 to Analysis Data FX1 respectively. n For any analysis data FX n1 FX data is obtained and analyzed based on text extraction. n1 All patient symptoms are listed and designated as clinical symptoms LB1 to LB1. k Where n represents the type of clinical data and k represents the type of patient symptoms; Establish a Cartesian coordinate system, denoted as the symptom differentiation coordinate system. The Y-axis of the symptom differentiation coordinate system is in units of quantity, and the coordinates to the right of the origin on the X-axis are filled with clinical symptom LB1 to clinical symptom LB1. k Extract patient symptoms from all analyzed data FX, and for any clinical symptom LB k1 Obtain LB containing clinical symptoms k1 The number of data points analyzed is denoted as p1, and the points in the symptom differentiation coordinate system (clinical symptom LB) are also considered. k1 p1) is recorded as the symptom point BZ. k1 ; Obtain the disease points BZ corresponding to all clinical symptoms LB, and denote the straight line fitted by all disease points BZ as the disease line. The x-coordinates of the leftmost and rightmost points of the disease line are clinical symptom LB1 and clinical symptom LB1, respectively. k Draw a straight line parallel to the X-axis from the midpoint of the symptom line and denote it as the span dividing line; The disease points BZ with the largest and smallest ordinates among all disease points BZ are respectively recorded as the peak point and the valley point; the line connecting the peak point and the valley point is recorded as the disease crossing line; the disease crossing line is translated so that the valley point coincides with the origin of the coordinate system, and the sine of the angle between the disease crossing line and the X-axis that is less than 90 degrees is recorded as the sine value of the crossing line. Divide the ordinate of the peak point by n and record the peak percentage; when the peak percentage is greater than or equal to the sine of the cross-line, mark the trough point as the clinical data FX. n1 The clinically necessary symptoms are used to mark the abscissa of all symptom points BZ except for the trough points as clinical data FX. n1 Secondary clinical symptoms; When the peak percentage is less than the sine value of the span, the abscissa of all symptom points BZ between the span dividing line and the X-axis is marked as clinical data FX. n1 Clinically necessary symptoms are identified, and symptom points BZ whose x-coordinates are not recorded as clinically necessary symptoms are designated as secondary points. The x-coordinates of these secondary points are then labeled as clinical data FX. n1 Secondary clinical symptoms; Obtain the clinically necessary and secondary symptoms corresponding to all analysis data FX. When any analysis data FX obtains both clinically necessary and secondary symptoms, adjust the analysis count of analysis data FX to 1.

2. The clinical dataset classification method according to claim 1, characterized in that, Clinical symptom analysis methods also include: Once the number of analyses for all clinical data is 1, obtain all distinct clinical symptoms (LB) corresponding to all clinical data and put them into set A; Create a table with T rows and R columns, denoted as the misjudgment analysis table. In the top row of the misjudgment analysis table, except for the first cell, fill in all the clinical symptoms LB of set A from left to right. In the leftmost column of the misjudgment analysis table, except for the first cell, fill in all the clinical data from top to bottom. Obtain two types of labels that can be used within the table, and denote them as γ1 and γ2 respectively; In the misjudgment analysis table: For any clinical data, mark the cell where the row containing the clinical data intersects with the column containing all the necessary clinical symptoms of the clinical data as γ1; mark the cell where the row containing the clinical data intersects with the column containing all the minor clinical symptoms of the clinical data as γ2; mark the rows containing all the clinical data based on the necessary and minor clinical symptoms corresponding to all the clinical data. In the misclassification analysis table after labeling with γ1 and γ2: For any clinical symptom LB, the misclassification analysis coefficient corresponding to the clinical symptom LB is obtained using the misclassification analysis algorithm, which is as follows: , where F is the misclassification analysis coefficient, f1 is the number of cells marked as γ1 in the column containing clinical symptom LB, and f2 is the number of cells marked as γ2 in the column containing clinical symptom LB.

3. The clinical dataset classification method according to claim 2, characterized in that, Clinical symptom analysis methods also include: Obtain the misjudgment analysis coefficients (LBs) for all clinical symptoms; record the LBs of clinical symptoms with a misjudgment analysis coefficient of 1 as major influencing symptoms, the LBs of clinical symptoms with a misjudgment analysis coefficient of 0 as minor influencing symptoms, and the LBs of clinical symptoms with a misjudgment analysis coefficient less than 1 and greater than 0 as mixed influencing symptoms. When there is 1 cell marked as γ1 in the column containing the primary disease, the primary disease is recorded as the only disease; when there is 1 cell marked as γ2 in the column containing the secondary disease, the secondary disease is recorded as the only disease. The number of major influencing diseases that are not recorded as unique diseases is denoted as A1, the number of minor influencing diseases that are not recorded as unique diseases is denoted as A2, and the number of mixed influencing diseases is denoted as A0; the value of A1 divided by A0 is recorded as the major judgment value, and the value of A2 divided by A0 is recorded as the minor judgment value. For any mixed-effect disease, the value of dividing the number of cells marked as γ1 in the column containing the mixed-effect disease by the total number of all marked cells in the column containing the mixed-effect disease is recorded as the primary self-judgment value, and the value of dividing the number of cells marked as γ2 in the column containing the mixed-effect disease by the total number of all marked cells in the column containing the mixed-effect disease is recorded as the secondary self-judgment value. When the primary self-judgment value is greater than or equal to the primary judgment value or the secondary self-judgment value is greater than or equal to the secondary judgment value, the mixed-effect disease is recorded as a disease that can be misjudged. All major and minor influencing symptoms that are not recorded as unique symptoms are recorded as potentially misdiagnosable symptoms.

4. The clinical dataset classification method according to claim 3, characterized in that, The analysis included identifying potentially misdiagnosable symptoms across all clinical conditions, and classifying all clinical data based on the analysis results, including: For any given clinical data, if there is no misdiagnosable disease in the clinical disease LB corresponding to the clinical data, the clinical data is classified as a non-misdiagnosable disease type. When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are major influencing diseases, the clinical data will be classified as the major disease misdiagnosis type. When there are misdiagnosable diseases in the clinical disease LB corresponding to the clinical data and all of them are minor diseases, the clinical data will be classified as minor disease misdiagnosis type. When the clinical data corresponds to clinical symptoms in the LB (Lesson Intake) and there are symptoms that can be misdiagnosed and symptoms that have mixed effects, the clinical data is classified into primary and secondary symptom misdiagnosis types.

5. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-readable instructions that, when executed by the processor, perform the steps of the method as described in any one of claims 1-4.

6. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it performs the steps of the method as described in any one of claims 1-4.