A heat stroke medical record text classification method based on reinforcement learning
By optimizing the medical record text classification process using reinforcement learning, the problem of removing redundant information in existing technologies was solved, enabling efficient and accurate classification of heatstroke medical record texts. This improved the accuracy and applicability of the classification and ensured timely treatment for critically ill patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHANJIANG CENT PEOPLES HOSPITAL
- Filing Date
- 2025-10-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot remove redundant information when classifying medical record texts, resulting in poor accuracy in classifying heatstroke medical records and failing to meet the high requirements of precision and adaptability.
We employ a reinforcement learning-based approach, utilizing techniques such as parameter discrimination, key region matching, and replacement analysis to optimize the medical record text classification process. This includes parameter matching classification, key region matching classification, and replacement analysis. By leveraging indicators such as feature keywords, information density, and anomaly detection coefficients, we can accurately locate and optimize key regions, thereby improving the accuracy and efficiency of classification.
This improved the accuracy and efficiency of heatstroke medical record classification, ensured the efficiency of emergency response for critically ill patients, and enhanced the clinical applicability and reliability of medical record classification.
Smart Images

Figure CN121278096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reinforcement learning technology, and in particular to a method for classifying heatstroke medical records based on reinforcement learning. Background Technology
[0002] Medical records contain a wealth of crucial information related to disease diagnosis, treatment selection, and prognosis assessment. Efficient and accurate classification of these texts has become a critical need in medical data mining, clinical research, and public health control. Because medical records of heatstroke are highly similar to those of common heatstroke and infectious fever, existing technologies for classifying these texts do not incorporate the specific characteristics of heatstroke medical records into their design of targeted reinforcement learning strategies. This makes it difficult to meet the high requirements for accuracy and adaptability in heatstroke medical record classification. Therefore, improving the classification accuracy of medical records is a problem that urgently needs to be solved by those skilled in the art.
[0003] Chinese Patent Publication No. CN117556047A discloses a method, device, and storage medium for classifying medical record information. The method includes: acquiring target medical record data and classification instruction information for guiding the classification of the target medical record data; extracting text from the target medical record data to obtain target medical record text; extracting classification information from the classification instruction information to obtain preliminary disease categories, category description information, and medical record screening information; encoding medical record features in the target medical record text to obtain initial medical record features; encoding category description information into category feature codes to obtain category description features; encoding screening information into screening feature codes to obtain medical record screening features; extracting features from the initial medical record features based on the medical record screening features to obtain target medical record features; and classifying the target medical record data based on the category description features and target medical record features. However, the above solution has the following problem: it cannot remove redundant information in the medical records, resulting in poor classification accuracy. Summary of the Invention
[0004] To address this issue, the present invention provides a heatstroke medical record text classification method based on reinforcement learning, which overcomes the problem in existing technologies that cannot remove redundant information in medical records, resulting in poor classification accuracy.
[0005] To achieve the above objectives, this invention provides a heatstroke medical record text classification method based on reinforcement learning, comprising:
[0006] Obtain the medical record text to be analyzed;
[0007] Based on the parameter discrimination, determine whether to perform parameter matching classification or key region matching classification;
[0008] In the key region matching classification, the initial key region is determined based on the total number of matching texts in the medical record text to be analyzed, and either the line frequency matching degree or the information density parameter; wherein, the information density parameter is determined based on the number of feature keywords.
[0009] The key area anomaly identification coefficient determines whether to perform area optimization to determine the final key area. When performing area optimization, the number of effective keywords is adjusted according to the anomaly comparison value, or the key area or the mapped area is increased according to the distribution coefficient of unclassified keywords.
[0010] Based on the matching degree of key areas, determine whether to perform direct classification or replacement analysis;
[0011] In the replacement analysis, keywords are replaced for replaceable keywords, and regions are classified based on the distinguishability of the replaced text.
[0012] Furthermore, parameter matching and classification are performed on the medical record texts to be analyzed where the parameter discrimination is equal to the standard parameter discrimination.
[0013] For medical record texts to be analyzed where the parameter discrimination is less than the standard parameter discrimination, key region matching and classification are performed.
[0014] Furthermore, for medical record texts to be analyzed where the total amount of matched text is greater than or equal to the preset total amount of matched text, the initial key region is determined based on the line frequency matching degree.
[0015] For medical record texts to be analyzed where the total amount of matched text is less than the preset total amount of matched text, the initial key region is determined based on the information density parameter.
[0016] The information density parameter is positively correlated with the number of feature keywords.
[0017] Furthermore, the methods for identifying characteristic keywords include:
[0018] For keywords whose frequency of appearance is greater than or equal to the preset frequency of appearance, determine whether they are feature keywords based on the inverse frequency reference value;
[0019] For keywords that appear less than the preset number of times, determine whether they are feature keywords based on their semantic influence.
[0020] Furthermore, for the medical record text to be analyzed where the anomaly identification coefficient of the key area is greater than or equal to the first preset anomaly identification coefficient, regional optimization is performed, and the number of effective keywords is increased based on the anomaly comparison value.
[0021] Furthermore, for medical record texts to be analyzed where the anomaly identification coefficient of the key area is less than the first preset anomaly identification coefficient but greater than or equal to the second preset anomaly identification coefficient, regional optimization is performed, and key areas or mapped areas are added based on the distribution coefficient of unclassified keywords.
[0022] Furthermore, for medical record texts to be analyzed where the distribution coefficient of unclassified keywords is greater than or equal to the preset distribution coefficient of unclassified keywords, key areas are added;
[0023] As key areas are being added, the number of rows in key areas is being increased based on the anomaly identification coefficient of the key areas.
[0024] The increase in the number of rows in the key area is positively correlated with the anomaly identification coefficient of the key area.
[0025] Furthermore, for medical record texts to be analyzed where the distribution coefficient of unclassified keywords is less than the preset distribution coefficient of unclassified keywords, the mapping region is increased;
[0026] As the mapping region is being added, the mapping rows are determined based on text similarity and region complementarity coefficients, and the reference value for the number of mapping rows to be added is determined based on the abnormal span value.
[0027] Furthermore, for medical record texts to be analyzed that have a key region identification matching degree greater than or equal to a preset key region identification matching degree, direct classification is performed.
[0028] Furthermore, for medical record texts to be analyzed where the key region identification matching degree is less than the preset key region identification matching degree, replacement analysis is performed;
[0029] In the replacement analysis, keywords are replaced for replaceable keywords, and the replacement texts are categorized based on their distinguishability.
[0030] Compared with the prior art, the beneficial effects of the present invention are that, in the technical solution of the present invention, the parameter discrimination degree effectively reflects the ability of the detection indicators of the medical records to be analyzed and historical medical records to effectively distinguish heatstroke classification. Then, based on the parameter discrimination degree, parameter matching classification or key area matching classification is adaptively performed, making the choice of classification method more in line with the actual application scenario. This is conducive to balancing the efficiency and accuracy of heatstroke medical record classification, thereby improving the clinical applicability and reliability of the classification results.
[0031] Furthermore, this invention matches the total amount of text to be analyzed in order to ensure the overall correlation between the total amount of text and the historical valid medical record text. Then, based on the total amount of text matched, key regions are adaptively determined according to the line frequency matching degree or information density parameter. This helps to balance the reliability and adaptability of key region positioning, thereby providing accurate and high-quality feature input for subsequent heatstroke medical record text classification, and ultimately improving the accuracy of the classification results.
[0032] Furthermore, in this invention, the frequency of appearance of target keywords in the heatstroke medical record text to be analyzed is effectively reflected by the number of appearances. Then, based on the number of appearances, the inverse frequency reference value or semantic influence degree is adaptively selected to determine the feature keywords, making the determination of feature keywords more in line with the actual application scenario. This is conducive to improving the accuracy and targeting of feature extraction from heatstroke medical record texts, thereby improving the reliability of heatstroke medical record classification.
[0033] Furthermore, in this invention, the key region anomaly identification coefficient effectively reflects the degree of deviation between the proportion of categorized keywords in the key regions of the medical record text to be analyzed and the ideal state, accurately judges the information validity level of the key regions, and then adaptively selects different region optimization methods according to the key region anomaly identification coefficient. This is conducive to matching differentiated strategies for the degree of information loss, ensuring that the key regions always have high-quality information to support subsequent classification, and improving the accuracy and efficiency of medical record text classification.
[0034] Furthermore, this invention effectively reflects the density of unclassified keywords in the initial key region by using the distribution coefficient of unclassified keywords, accurately determining the potential location of effective information. Then, based on the distribution coefficient of unclassified keywords, the key region or the mapping region is adaptively increased, which helps to avoid information redundancy or omission caused by indiscriminate expansion. It matches efficient information mining methods for different distribution characteristics, thereby accurately filling the information gaps in the key region, ensuring the integrity and accuracy of medical record text feature extraction, and providing high-quality data support for subsequent classification. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the heatstroke medical record text classification method based on reinforcement learning according to the present invention;
[0036] Figure 2 This is a flowchart illustrating the process of determining parameter matching classification or key region matching classification based on parameter discrimination in this invention.
[0037] Figure 3 This is a flowchart illustrating the process of adding key regions or mapping regions based on the distribution coefficient of unclassified keywords in this invention.
[0038] Figure 4 This is a flowchart illustrating the process of direct classification or replacement analysis based on the matching degree of key area identification in this invention. Detailed Implementation
[0039] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0040] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0041] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0042] Please see Figures 1 to 4 As shown, this invention provides a heatstroke medical record text classification method based on reinforcement learning, including:
[0043] Obtain the medical record text to be analyzed;
[0044] Based on the parameter discrimination, determine whether to perform parameter matching classification or key region matching classification;
[0045] In the key region matching classification, the initial key region is determined based on the total number of matching texts in the medical record text to be analyzed, and either the line frequency matching degree or the information density parameter; wherein, the information density parameter is determined based on the number of feature keywords.
[0046] The key area anomaly identification coefficient determines whether to perform area optimization to determine the final key area. When performing area optimization, the number of effective keywords is adjusted according to the anomaly comparison value, or the key area or the mapped area is increased according to the distribution coefficient of unclassified keywords.
[0047] Based on the matching degree of key areas, determine whether to perform direct classification or replacement analysis;
[0048] In the replacement analysis, keywords are replaced for replaceable keywords, and regions are classified based on the distinguishability of the replaced text.
[0049] The application scenario of this invention is to classify heatstroke medical record texts through reinforcement learning. This invention uses the "key area localization - unclassified keyword replacement - dynamic classification" process of the reinforcement learning model to quickly process heatstroke medical record texts, ensuring that critically ill patients receive priority treatment and improving emergency response efficiency.
[0050] This invention includes several historical records. Each historical record records at least one instance of heatstroke medical record text classification using reinforcement learning, including the number of times the record appeared, the reference value of the inverse frequency, the semantic influence, and the distribution coefficient of unclassified keywords. Each historical record also has a corresponding qualified marker, which records whether the heatstroke medical record text classification process using reinforcement learning meets the user's needs. The qualified marker can be recorded manually. It is understood that the user can determine whether the heatstroke medical record text classification process using reinforcement learning meets the requirements based on self-defined indicators. Self-defined indicators can be, but are not limited to, the number of anomalies, which will not be elaborated here. The number of anomalies refers to the number of times the heatstroke medical record text was classified as a non-heatstroke medical record text.
[0051] The medical record text to be analyzed is the electronic medical record text of a single patient.
[0052] Specifically, parameter matching and classification are performed on medical record texts to be analyzed where the parameter discrimination is equal to the standard parameter discrimination.
[0053] For medical record texts to be analyzed where the parameter discrimination is less than the standard parameter discrimination, key region matching and classification are performed.
[0054] Specifically, heatstroke medical record texts in the historical records that meet the user's needs are recorded as valid medical record texts, while medical record texts in the historical records that do not meet the user's needs but are not heatstroke medical record texts are recorded as invalid medical record texts.
[0055] The parameter matching degree between any two medical record texts is the minimum value among the indicator matching degrees of each detection indicator in the medical record text to be analyzed; the indicator matching degree of a single detection indicator is the difference between 1 and the ratio reference value. The absolute value of the difference between the monitoring parameters corresponding to the detection indicator in the two medical record texts is recorded as the first reference value, the larger value among the detection parameters corresponding to the detection indicator in the two medical record texts is recorded as the second reference value, and the ratio of the first reference value to the second reference value is recorded as the ratio reference value.
[0056] The detection indicators in the medical record text include, but are not limited to, body temperature, blood glucose and central venous oxygen saturation. Each detection indicator has corresponding detection parameters, which are the actual measured values corresponding to the detection indicators.
[0057] Valid medical record texts with a parameter matching degree greater than the preset parameter matching degree are recorded as valid medical record texts, and invalid medical record texts with a parameter matching degree greater than the preset parameter matching degree are recorded as invalid medical record texts.
[0058] The method for confirming parameter discrimination is as follows:
[0059] The parameter discrimination is 1 when only valid medical record text or only invalid medical record text exists.
[0060] The parameter discrimination is 0 when neither a valid nor an invalid medical record text exists, or when both a valid and an invalid medical record text exists.
[0061] The standard parameter discrimination index is 1;
[0062] When performing parameter matching and classification, if only valid matching medical record texts exist, the medical record text to be analyzed is heatstroke medical record text; if only invalid matching medical record texts exist, the medical record text to be analyzed is non-heatstroke medical record text.
[0063] Understandably, parameter discrimination effectively reflects the ability of the detection indicators in the analyzed medical records and historical medical records to differentiate between heatstroke classification. When the parameter discrimination equals the standard parameter discrimination, it indicates that the core indicators clearly support the classification conclusion, and further analysis based on text content is unnecessary; therefore, parameter matching classification is performed. When the parameter discrimination is less than the standard parameter discrimination, it indicates that the core indicators cannot support the classification conclusion, and it is necessary to delve into the key areas of the medical record text to further explore the classification basis through information density, keyword distribution, etc.; therefore, key area matching classification is performed.
[0064] Specifically, for medical record texts to be analyzed where the total amount of matched text is greater than or equal to the preset total amount of matched text, the initial key region is determined based on the line frequency matching degree.
[0065] For medical record texts to be analyzed where the total amount of matched text is less than the preset total amount of matched text, the initial key region is determined based on the information density parameter.
[0066] The information density parameter is positively correlated with the number of feature keywords.
[0067] Specifically, the matching coefficient between the medical record text to be analyzed and a single valid medical record text is the ratio of the number of keywords in the same position to the number of keywords in the medical record text to be analyzed; the keywords in the same position between the medical record text to be analyzed and a single valid medical record text are keywords that are located in the same line and are completely identical. The identical keywords are keywords that are in the same line number position in the medical record text to be analyzed and the valid medical record text. For any medical record text, the line number position is set sequentially according to the natural arrangement order of the line units, i.e., line 1, line 2, line 3, ..., line n (n is the total number of lines);
[0068] The preset matching coefficient value can be determined by the user based on the actual application scenario. The greater the user's need to improve the accuracy of medical record text classification, the larger the preset matching coefficient value should be. One preset matching coefficient value is provided: 0.75.
[0069] The matched text is the valid medical record text whose matching coefficient with the medical record text to be analyzed is greater than the preset matching coefficient; the total number of matched texts is the number of valid medical record texts whose matching coefficient with the medical record text to be analyzed is greater than the preset matching coefficient.
[0070] Each matching text is a medical record text in the history that has been classified and identified as heatstroke medical record text. The corresponding final key region of each matching text has been determined. The line number of the final key region of all matching texts is extracted. For a single line number, the line frequency matching degree corresponding to the line number is the ratio of the number of matching texts that take the line number as the final key region to the total number of matching texts.
[0071] When determining the initial key region based on line frequency matching degree, the text of the line corresponding to the line number with a line frequency matching degree greater than the preset line frequency matching degree is taken as the initial key region;
[0072] The preset value of the total amount of matched text can be determined by the user according to the actual application scenario. The greater the user's need for classification speed, the smaller the preset value of the total amount of matched text. A method for setting the preset value of the total amount of matched text is provided: detect the historical records of the user's determination of the initial key area based on the line frequency matching degree, and record the average value of the total amount of matched text corresponding to the historical records that can meet the user's needs as the preset total amount of matched text.
[0073] When determining the initial key region based on the information density parameter, the text of the line corresponding to the line number whose information density parameter is greater than the preset information density parameter is taken as the initial key region.
[0074] The information density parameter corresponding to a single line number is the ratio of the number of feature keywords in the line corresponding to that line number in the medical record text to be analyzed to the total number of keywords in the line corresponding to that line number in the medical record text to be analyzed;
[0075] Keywords in medical record texts are extracted using NLP algorithms, a common technique used by those skilled in the art, and will not be elaborated upon here.
[0076] Users can determine the values of the preset line frequency matching degree and preset information density parameters according to the actual application scenario. The greater the user's need for improving the accuracy of text classification, the higher the values of the preset line frequency matching degree and preset information density parameters will be. One preset line frequency matching degree and preset information density parameter value is provided: preset line frequency matching degree is 0.6 and preset information density parameter is 0.5.
[0077] Understandably, the total amount of matched text effectively reflects the overall sufficiency of the correlation between the medical record to be analyzed and historical valid medical record texts. When the total amount of matched text is greater than or equal to the preset total amount of matched text, it indicates that the core information of the medical record to be analyzed (such as causes, symptoms, and signs) has a high degree of overlap with a sufficient number of historically highly correlated medical records. The key areas of historical medical records have sufficient reference value. Therefore, it is not necessary to rely on the information of the medical record to be analyzed itself. The initial key area can be determined directly based on the line frequency matching degree, which is both efficient and reduces the risk of information bias in a single medical record. When the total amount of matched text is less than the preset total amount of matched text, it indicates that the core information of the medical record to be analyzed has a low degree of overlap with historical valid medical records. Therefore, it is necessary to turn to the medical record to be analyzed itself and focus on the area with dense feature keywords through the information density parameter to ensure that the key area positioning does not rely on scarce historical references but is based on the core information of the medical record itself. Therefore, the initial key area is determined based on the information density parameter.
[0078] Specifically, the methods for identifying characteristic keywords include:
[0079] For keywords whose frequency of appearance is greater than or equal to the preset frequency of appearance, determine whether they are feature keywords based on the inverse frequency reference value;
[0080] For keywords that appear less than the preset number of times, determine whether they are feature keywords based on their semantic influence.
[0081] Specifically, for a single keyword, the keyword is recorded as the target keyword, and the number of times the target keyword appears is the number of times the target keyword appears in the medical record text to be analyzed;
[0082] The value of the preset number of appearances can be determined by the user based on the actual application scenario. The greater the user's need for improving the accuracy of heatstroke medical record text classification, the smaller the value of the preset number of appearances. A method for determining the value of the preset number of appearances is provided, which is the average number of appearances of each keyword in the history of the user's determination of whether it is a feature keyword based on semantic influence.
[0083] When determining whether a keyword is a feature keyword based on the inverse frequency reference value, keywords with an inverse frequency reference value less than the preset inverse frequency reference value are recorded as feature keywords;
[0084] When determining whether a keyword is a feature keyword based on its semantic influence, keywords with a semantic influence greater than the preset semantic influence are recorded as feature keywords.
[0085] The formula for calculating the inverse frequency reference value I corresponding to the target keyword is: I=log[(a-a0+k) / (a0+k)], where a is the number of valid medical record texts, a0 is the number of valid medical record texts in the final key area where the target keyword appears, and k is the smoothing coefficient, with a value of 0.5;
[0086] Other keywords in the medical record text to be analyzed, excluding the target keyword, are recorded as reference keywords. The semantic influence of the target keyword is the ratio of the number of related keywords to the total number of reference keywords.
[0087] The number of relevant keywords is the number of reference keywords whose semantic similarity to the target keyword is greater than the preset semantic similarity.
[0088] For any two keywords, the word2vec model is used to convert the keywords into word vectors. The cosine similarity between the two word vectors is recorded as the semantic similarity between the two keywords. The range of the semantic similarity is [0, 1].
[0089] The preset semantic similarity value can be determined by the user according to the actual application scenario. The greater the user's need for improving the accuracy of medical record text classification, the higher the preset semantic similarity value. One preset semantic similarity value is provided, with a preset semantic similarity of 0.7.
[0090] Users can determine the preset inverse frequency reference value and preset semantic influence value according to the actual application scenario. The greater the user's requirement for the clustering accuracy of heatstroke medical record text, the smaller the preset inverse frequency reference value and the larger the preset semantic influence value. A method for determining the preset inverse frequency reference value and preset semantic influence value is provided, which takes the average value of the inverse frequency reference value and the average value of the semantic influence value corresponding to each feature keyword in the historical record that can meet the user's needs as the preset inverse frequency reference value and preset semantic influence value, respectively.
[0091] Understandably, the frequency of appearance effectively reflects the occurrence frequency of target keywords in the heatstroke medical records to be analyzed. When the frequency of appearance is greater than or equal to the preset frequency of appearance, it indicates that the keyword appears frequently in the medical records. However, there may be a problem that the keyword is described too many times but is not specific to the characteristics of heatstroke. It is necessary to further determine its universality in the key areas of the medical records. At this time, it is necessary to judge its specific value by using the inverse frequency reference value. The inverse frequency reference value can quantify the coverage of the keyword in all valid medical records. The narrower the coverage, the higher the inverse frequency reference value. Therefore, it is necessary to determine whether it is a feature keyword based on the inverse frequency reference value.
[0092] When the number of occurrences is less than the preset number of occurrences, it indicates that the keyword appears in the medical record in a low frequency and cannot directly reflect the universal characterization ability of heatstroke characteristics through the frequency of occurrence. In this case, it is necessary to verify its feature value through semantic relevance. Semantic influence can measure the strength of the association between the keyword and other keywords in the medical record. The stronger the association, the more prominent its supplementary characterization role in the disease characteristics. Therefore, it is determined whether it is a feature keyword based on semantic influence.
[0093] Specifically, for medical record texts to be analyzed where the anomaly identification coefficient in key areas is greater than or equal to the first preset anomaly identification coefficient, regional optimization is performed, and the number of effective keywords is increased based on the anomaly comparison value.
[0094] Specifically, for medical record texts to be analyzed where the anomaly identification coefficient of the key area is less than the second preset anomaly identification coefficient, no area optimization is required, and the initial key area is directly recorded as the final key area.
[0095] The first preset anomaly detection coefficient is greater than the second preset anomaly detection coefficient;
[0096] The key area anomaly identification coefficient is the difference between the standard ratio and the quantity ratio. The standard ratio is 1, and the quantity ratio is the ratio of the number of categorized keywords in the initial key area to the total number of keywords in the initial key area.
[0097] This invention includes a keyword library for storing several keywords, including but not limited to anhidrosis, convulsions, dizziness and headache, nausea and vomiting, fatigue, heat cramps, dehydration, and electrolyte imbalance. If the target keyword is the same as any keyword in the keyword library, the target keyword is recorded as a categorized keyword. If the target keyword is different from all keywords in the keyword library, the target keyword is recorded as an uncategorized keyword.
[0098] The values of the first preset anomaly recognition coefficient and the second preset anomaly recognition coefficient can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of text classification, the greater the values of the first preset anomaly recognition coefficient and the second preset anomaly recognition coefficient. One possible value for the first preset anomaly recognition coefficient and the second preset anomaly recognition coefficient is 0.7 and 0.5.
[0099] The anomaly comparison value is the difference between the anomaly identification coefficient of the key area and the first preset anomaly identification coefficient;
[0100] When adjusting the number of effective keywords, rows containing reference feature keywords are marked as rows to be added. Rows to be added are selected and added to the key area in ascending order of row number intervals until the key area reaches the increased value for the number of effective keywords. The text corresponding to each added row and each row in the initial key area is recorded as the final key area. For medical record texts to be analyzed with anomaly comparison values greater than or equal to the preset anomaly comparison value, the increase in the number of effective keywords is 0.37 times the initial number of effective keywords. For medical record texts to be analyzed with anomaly comparison values less than the preset anomaly comparison value, the increase in the number of effective keywords is 0.25 times the initial number of effective keywords.
[0101] Effective keywords are the feature keywords that appear in the initial key area, while reference feature keywords are the feature keywords that have not appeared in the initial key area; the initial number of effective keywords is the number of different feature keywords that appear in the initial key area.
[0102] The row interval for a single row is the number of rows between that row and the row with the smallest interval in the initial key area.
[0103] It should be noted that after adjusting the number of valid keywords based on the anomaly comparison value, if the adjusted key area anomaly identification coefficient is greater than or equal to the second preset anomaly identification coefficient, then the key area or the mapping area will be increased based on the distribution coefficient of unclassified keywords.
[0104] Specifically, for medical record texts to be analyzed where the anomaly identification coefficient of the key area is less than the first preset anomaly identification coefficient but greater than or equal to the second preset anomaly identification coefficient, the region is optimized, and the key area or the mapped area is increased based on the distribution coefficient of the unclassified keywords.
[0105] Specifically, the distribution coefficient of unclassified keywords is the average of the interval reference values corresponding to each unclassified keyword in the initial key area. For a single unclassified keyword in the initial key area, the unclassified keyword is recorded as the target unclassified keyword. Other unclassified keywords in the initial key area, excluding the target unclassified keyword, are recorded as reference unclassified keywords. The interval reference value corresponding to the target unclassified keyword is the number of rows between the row containing the target unclassified keyword and the row containing the nearest reference unclassified keyword.
[0106] Understandably, when the distribution coefficient of unclassified keywords in the medical record text to be analyzed is greater than or equal to the preset distribution coefficient of unclassified keywords, it indicates that the distribution of unclassified keywords in the initial key area is relatively sparse. Since electronic medical records are typically distributed by themes such as chief complaint, present illness, past medical history, and progress notes, and the logical connections between these thematic paragraphs are close, information within the same scenario is often concentrated in adjacent areas. In this case, the sparse and scattered unclassified keywords are mostly fragmented representations of marginally related information, and the core effective content is more likely to be located in paragraphs adjacent to the initial key area. Expanding the surrounding area is sufficient to uncover useful information, hence the addition of key areas. When the distribution coefficient of unclassified keywords in the medical record text to be analyzed is less than the preset distribution coefficient of unclassified keywords, it indicates that the distribution of unclassified keywords in the initial key area is relatively dense, and effective information is more likely to exist in paragraphs related to the themes of the key area in the medical record. Therefore, the mapping area is added.
[0107] Specifically, for medical record texts to be analyzed where the distribution coefficient of unclassified keywords is greater than or equal to the preset distribution coefficient of unclassified keywords, key areas are added.
[0108] As key areas are being added, the number of rows in key areas is being increased based on the anomaly identification coefficient of the key areas.
[0109] The increase in the number of rows in the key area is positively correlated with the anomaly identification coefficient of the key area.
[0110] Specifically, the value of the preset unclassified keyword distribution coefficient can be determined by the user based on the actual application scenario. The greater the user's need for improving the accuracy of medical record text feature extraction, the smaller the value of the preset unclassified keyword distribution coefficient should be. A method for determining the value of the preset unclassified keyword distribution coefficient is provided by detecting the user's historical records of adding key areas and recording the average value of the unclassified keyword distribution coefficients corresponding to the historical records that meet the user's needs as the preset unclassified keyword distribution coefficient.
[0111] The increase in the number of rows in the critical area = (critical area anomaly identification coefficient - second preset anomaly identification coefficient) / second preset anomaly identification coefficient × initial number of rows in the critical area; the initial number of rows in the critical area is the number of rows in the initially determined critical area;
[0112] When adding key regions, add reference rows to the key regions in order of increasing average interval until the number of added reference rows reaches the increase value of the number of rows in the key regions; record the text corresponding to each added reference row and each row in the initial key region as the final key region;
[0113] The rows outside the initial key area are designated as reference rows. The average interval of a single reference row is the average number of interval rows between the reference row and each row in the initial key area. The number of interval rows between a single reference row and a single row in the initial key area is the number of rows between the two rows.
[0114] Specifically, for medical record texts to be analyzed where the distribution coefficient of unclassified keywords is less than the preset distribution coefficient of unclassified keywords, the mapping region is increased;
[0115] As the mapping region is being added, the mapping rows are determined based on text similarity and region complementarity coefficients, and the reference value for the number of mapping rows to be added is determined based on the abnormal span value.
[0116] Specifically, the text similarity of a single reference line is the maximum value of the cosine similarity between the text vector corresponding to that reference line and the text vectors corresponding to each line in the reference text paragraph;
[0117] The text includes, but is not limited to, the text corresponding to the reference line and the text corresponding to each line in the reference text paragraph. For any text, the word vectors corresponding to each keyword in the text are determined by the word2vec model, the weight of each keyword in the text is determined by the TF-IDF algorithm, the word vector is multiplied by the corresponding TF-IDF weight, and then the average of all weighted word vectors is taken to obtain a text vector with the same dimension as the word vector. In addition, there are other technical means commonly used by those skilled in the art, which will not be elaborated on in detail.
[0118] The regional complementarity coefficient corresponding to a single reference row = the number of complementary keywords in that reference row / the average number of complementary keywords in all reference rows + the average reference ratio corresponding to each complementary keyword in that reference row / the average reference ratio corresponding to each complementary keyword in the medical record text to be analyzed;
[0119] The reference text paragraphs are the text in the lines containing uncategorized keywords in the initial key area;
[0120] For a single keyword in the reference row and a single valid medical record text, if the keyword and any keyword in the reference text paragraph appear in the same line of the final key area in the valid medical record text, then the valid medical record text is recorded as a reference valid text. The ratio of the number of reference valid texts corresponding to a single keyword in the reference row to the number of valid medical record texts is recorded as the reference ratio. If the reference ratio is greater than the preset ratio, then the keyword is recorded as a complementary keyword.
[0121] The user can determine the value of the preset ratio according to the actual application scenario. The greater the user's need for improving the accuracy of medical record text classification, the larger the value of the preset ratio. A method for determining the preset ratio is provided, with a preset ratio of 0.62.
[0122] When determining the mapping row based on text similarity and region complementarity coefficient, reference rows with text similarity greater than the preset text similarity or region complementarity coefficient greater than the preset region complementarity coefficient are used as mapping rows.
[0123] Users can determine the preset values of text similarity and region complementarity coefficient according to the actual application scenario. The greater the user's need for improving the accuracy of medical record text classification, the higher the values of preset text similarity and region complementarity coefficient will be. One preset value of text similarity and region complementarity coefficient is provided: preset text similarity is 0.68 and preset region complementarity coefficient is 0.72.
[0124] The abnormal span value is the ratio of the total number of unclassified keywords in the initial key region to the number of rows in the initial key region containing unclassified keywords;
[0125] When adding mapping regions, the mapping lines are recorded as added mapping lines in descending order of comprehensive evaluation value until the number of added mapping lines reaches the reference value for the number of added mapping lines, and each added mapping line is recorded in the key region; the text corresponding to each added mapping line and each line in the initial key region is recorded as the final key region;
[0126] For medical record texts to be analyzed with abnormal span values greater than or equal to the preset abnormal span value, the reference value for the number of mapping lines added is 0.8 times the total number of mapping lines; for medical record texts to be analyzed with abnormal span values less than the preset abnormal span value, the reference value for the number of mapping lines added is 0.5 times the total number of mapping lines.
[0127] Understandably, the key area anomaly identification coefficient effectively reflects the degree to which the proportion of categorized keywords in the key areas of the medical record text being analyzed deviates from the ideal state;
[0128] When the anomaly identification coefficient of the key area is greater than or equal to the first preset anomaly identification coefficient, it indicates that the proportion of heatstroke-related classification keywords in the initial key area is extremely low, the useful information is seriously insufficient, the deviation from the ideal state is high, and it can no longer provide effective support for classification. Therefore, the area is optimized, and the number of effective keywords is increased according to the anomaly comparison value. The information quality of the key area is improved by supplementing heatstroke feature keywords.
[0129] When the anomaly identification coefficient of the key area is less than the first preset anomaly identification coefficient but greater than or equal to the second preset anomaly identification coefficient, it indicates that the proportion of categorized keywords in the initial key area is moderate, the useful information is insufficient but not severe, and the deviation from the ideal state is moderate. It may miss atypical symptoms (such as "muscle spasms" in uncategorized keywords) or lack symptom-related information (such as only mentioning "nausea and vomiting" without mentioning "high temperature exposure triggers"). If only core keywords are added, the classification requirements may still not be met due to incomplete information dimensions. Therefore, area optimization is required. If only core keywords are added, the accurate classification requirements may still not be met due to missing information dimensions. Therefore, it is necessary to select to increase the key area or the mapping area according to the distribution coefficient of uncategorized keywords, and to explore more useful information by expanding the coverage of the key area.
[0130] When the anomaly identification coefficient of the key area is less than the second preset anomaly identification coefficient, it means that the proportion of categorized keywords in the initial key area is already high, there is enough useful information, and the information quality of the initial key area can support subsequent classification. No additional optimization is needed, and the process can proceed directly to the next classification step.
[0131] Specifically, for medical record texts to be analyzed that have a key area identification matching degree greater than or equal to a preset key area identification matching degree, direct classification is performed.
[0132] Specifically, the key region identification matching degree is the difference between the effective matching coefficient and the invalid matching coefficient; the effective matching coefficient is the minimum value of the matching coefficients between the medical record text to be analyzed and each effective medical record text, and the invalid matching coefficient is the maximum value of the matching coefficients between the medical record text to be analyzed and each invalid medical record text.
[0133] The matching coefficient between the medical record text to be analyzed and the single medical record text is the ratio of the first quantity to the second quantity. The number of categorized keywords that appear in the final key areas corresponding to both medical record texts is recorded as the first quantity, and the larger value of the number of keywords that appear in the final key areas corresponding to both medical record texts is recorded as the second quantity.
[0134] When performing direct classification, the medical record text to be analyzed is directly recorded as heatstroke medical record text;
[0135] Specifically, for medical record texts to be analyzed where the key area identification matching degree is less than the preset key area identification matching degree, replacement analysis is performed;
[0136] In the replacement analysis, keywords are replaced for replaceable keywords, and the replacement texts are categorized based on their distinguishability.
[0137] Specifically, replaceable keywords are unclassified keywords in the final key area of the text to be analyzed that have a semantic similarity greater than the semantic similarity threshold or a pinyin matching degree greater than the pinyin matching degree threshold with any keyword in the keyword library; the semantic similarity threshold and the pinyin matching degree threshold are both 0.9;
[0138] For a single replaceable keyword, replace the replaceable keyword with the keyword in the keyword library that has the highest replacement match value for that replaceable keyword;
[0139] The replacement matching value for any two keywords is the larger of the semantic similarity and the pinyin matching degree of the two keywords.
[0140] The method for determining the pinyin matching degree between any two keywords is as follows: if the two keywords have different numbers of characters, the pinyin matching degree is 0; if the two keywords have the same number of characters, the pinyin matching degree is the minimum value among the pinyin similarity of characters corresponding to each character in the same position. Characters in the same position are characters in the same order in the two keywords. The pinyin similarity of a single character in the same position = the number of the same initials and finals in the two characters / the average value of the pinyin volume of the two characters. The pinyin volume of a single character is the number of initials and finals in the pinyin corresponding to that character.
[0141] Valid matching text refers to valid medical record texts whose matching coefficient with the medical record text to be analyzed after keyword replacement is greater than the preset matching coefficient; invalid matching text refers to invalid medical record texts whose matching coefficient with the medical record text to be analyzed after keyword replacement is greater than the preset matching coefficient.
[0142] When classifying based on the distinguishability of the replacement text, the distinguishability of the replacement text is 1 when there is only valid text or only invalid text. When there is only valid text, the category of the medical record text to be analyzed is recorded as heatstroke medical record text. When there is only invalid text, the category of the medical record text to be analyzed is recorded as non-heatstroke medical record text.
[0143] When neither valid nor invalid text exists, or when both valid and invalid text exist, the replacement text has a discrimination score of 0, and a classification error alert is sent to the user for manual classification.
[0144] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A heatstroke medical record text classification method based on reinforcement learning, characterized in that, include: Obtain the medical record text to be analyzed; Based on the parameter discrimination, determine whether to perform parameter matching classification or key region matching classification; In the key region matching classification, the initial key region is determined based on the total number of matching texts in the medical record text to be analyzed, and the initial key region is determined based on the line frequency matching degree or information density parameter; wherein, the information density parameter is determined based on the number of feature keywords; The key area anomaly identification coefficient determines whether to perform area optimization to determine the final key area. When performing area optimization, the number of effective keywords is increased based on the anomaly comparison value, or the key area or the mapped area is increased based on the distribution coefficient of unclassified keywords. Based on the matching degree of key areas, determine whether to perform direct classification or replacement analysis; In the replacement analysis, keywords are replaced for replaceable keywords, and regions are classified according to the distinguishability of the replaced text; When determining the initial key region based on line frequency matching degree, the text of the line corresponding to the line number with a line frequency matching degree greater than the preset line frequency matching degree is taken as the initial key region; Among them, the line frequency matching degree corresponding to a single line number is the ratio of the number of matching texts that take the line number as the final key area to the total number of matching texts; the matching text is a valid medical record text whose matching coefficient with the medical record text to be analyzed is greater than the preset matching coefficient; the matching coefficient between the medical record text to be analyzed and a single valid medical record text is the ratio of the number of keywords in the same position to the number of keywords in the medical record text to be analyzed. When adjusting the number of effective keywords based on the anomaly comparison value, rows containing reference feature keywords are recorded as rows to be added. Rows to be added are selected in ascending order of row number interval value and recorded in the key area until the key area reaches the increase value of the number of effective keywords. The text corresponding to each added row and each row in the initial key area is recorded as the final key area. Among them, the reference feature keywords are feature keywords that have not appeared in the initial key area. For medical record texts to be analyzed where the distribution coefficient of unclassified keywords is greater than or equal to the preset distribution coefficient of unclassified keywords, key areas are added. For medical record texts to be analyzed where the distribution coefficient of unclassified keywords is less than the preset distribution coefficient of unclassified keywords, the mapping region is increased; The key area anomaly identification coefficient is the difference between the standard ratio and the quantity ratio. The standard ratio is 1, and the quantity ratio is the ratio of the number of categorized keywords in the initial key area to the total number of keywords in the initial key area. The anomaly comparison value is the difference between the anomaly identification coefficient of the key area and the first preset anomaly identification coefficient; The distribution coefficient of unclassified keywords is the average of the interval reference values corresponding to each unclassified keyword in the initial key area. For a single unclassified keyword in the initial key area, the unclassified keyword is recorded as the target unclassified keyword. The other unclassified keywords in the initial key area, excluding the target unclassified keyword, are recorded as reference unclassified keywords. The interval reference value corresponding to the target unclassified keyword is the number of rows between the row containing the target unclassified keyword and the row containing the nearest reference unclassified keyword.
2. The heatstroke medical record text classification method based on reinforcement learning according to claim 1, characterized in that, For medical record texts to be analyzed where the parameter discrimination is equal to the standard parameter discrimination, parameter matching and classification are performed. For medical record texts to be analyzed where the parameter discrimination is less than the standard parameter discrimination, key region matching and classification are performed.
3. The heatstroke medical record text classification method based on reinforcement learning according to claim 2, characterized in that, For medical record texts to be analyzed where the total amount of matched text is greater than or equal to the preset total amount of matched text, the initial key region is determined based on the line frequency matching degree. For medical record texts to be analyzed where the total amount of matched text is less than the preset total amount of matched text, the initial key region is determined based on the information density parameter. The information density parameter is positively correlated with the number of feature keywords.
4. The heatstroke medical record text classification method based on reinforcement learning according to claim 3, characterized in that, Methods for identifying key features include: For keywords whose frequency of appearance is greater than or equal to the preset frequency of appearance, determine whether they are feature keywords based on the inverse frequency reference value; For keywords that appear less than the preset number of times, determine whether they are feature keywords based on their semantic influence.
5. The heatstroke medical record text classification method based on reinforcement learning according to claim 1, characterized in that, For medical record texts to be analyzed where the anomaly identification coefficient in key areas is greater than or equal to the first preset anomaly identification coefficient, regional optimization is performed, and the number of effective keywords is increased based on the anomaly comparison value.
6. The heatstroke medical record text classification method based on reinforcement learning according to claim 5, characterized in that, For medical record texts to be analyzed where the anomaly identification coefficient of the key area is less than the first preset anomaly identification coefficient but greater than or equal to the second preset anomaly identification coefficient, the region is optimized, and the key area or the mapped area is increased according to the distribution coefficient of the unclassified keywords.
7. The heatstroke medical record text classification method based on reinforcement learning according to claim 6, characterized in that, For medical record texts to be analyzed where the distribution coefficient of unclassified keywords is greater than or equal to the preset distribution coefficient of unclassified keywords, key areas are added. As key areas are being added, the number of rows in key areas is being increased based on the anomaly identification coefficient of the key areas. The increase in the number of rows in the key area is positively correlated with the anomaly identification coefficient of the key area.
8. The heatstroke medical record text classification method based on reinforcement learning according to claim 6, characterized in that, For medical record texts to be analyzed where the distribution coefficient of unclassified keywords is less than the preset distribution coefficient of unclassified keywords, the mapping region is increased; As the mapping region is being added, the mapping rows are determined based on text similarity and region complementarity coefficients, and the reference value for the number of mapping rows to be added is determined based on the abnormal span value.
9. The heatstroke medical record text classification method based on reinforcement learning according to claim 1, characterized in that, For medical record texts to be analyzed that have a key area identification matching degree greater than or equal to the preset key area identification matching degree, direct classification is performed.
10. The heatstroke medical record text classification method based on reinforcement learning according to claim 9, characterized in that, Replacement analysis is performed on medical record texts whose key region identification matching degree is less than the preset key region identification matching degree; In the replacement analysis, keywords are replaced for replaceable keywords, and the replacement texts are categorized based on their distinguishability.