Historical inquiry data flexible screening method and system

By normalizing the historical consultation data and evaluating importance, determining the order of data deletion, the problem of unreasonable data deletion in the existing technology is solved, flexible screening is realized, and storage costs are reduced.

CN120183737AActive Publication Date: 2025-06-20THE 3RD AFFILIATED HOSPITAL OF CHANGCHUN UNIVERSITY OF CHINESE MEDICINE
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510653966.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The prior art lacks the importance of considering data when deleting historical consultation data, resulting in high cost and unreasonable data deletion methods.

Method used

The consultation data is normalized based on the preset medical record template, a consultation array is constructed, and the data deletion order is determined through a comprehensive assessment of local importance and global importance, so as to achieve flexible screening.

Benefits of technology

It realizes the importance of data based on the predictability of data and the degree of retention requirements, and deletes data according to the order of importance, avoiding the irrationality of traditional regular deletion methods and being more in line with the actual situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183737A_ABST
    Figure CN120183737A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and particularly discloses a historical inquiry data flexible screening method and system, and the method comprises the steps: querying an identity label of any patient, reading an inquiry array of the patient according to the identity label, and carrying out the sorting based on a time label, and obtaining an inquiry array sequence; determining the local importance of each data according to the inquiry array sequence; all the inquiry arrays are clustered, a mean value array of each type of inquiry arrays is determined, and the global importance degree of each inquiry array is determined according to the mean value array; determining a data deletion sequence according to the global importance of the inquiry array and the local importance of each piece of data; according to the method, the importance of the data is jointly determined according to the predictability and the reserved demand degree, and the data deletion sequence is determined according to the ascending order of the importance, so that the unimportant data is deleted firstly, and compared with a traditional regular deletion mode, the method is softer and more consistent with the actual situation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and specifically to a flexible screening and deletion method and system for historical medical consultation data. Background Art

[0002] With the development of society and the progress of technology, a lot of data in existing hospital areas has been converted into digital data. These data can be stored in a database with a very long storage time, and the current memory cost is very low, having sufficient storage capacity to store these data. However, the cost of the memory still exists, including its own cost, occupied space cost, and maintenance cost. If data is to be stored all the time, these costs will become higher and higher. And some data in the medical consultation data may not be so important and can be deleted regularly. The existing deletion methods are all hard deletion processes according to time, without much consideration of the importance of the data. Therefore, how to provide a more suitable screening and deletion scheme based on the data itself is the technical problem that the technical solution of the present invention wants to solve. Summary of the Invention

[0003] The purpose of the present invention is to provide a flexible screening and deletion method and system for historical medical consultation data to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions: A flexible screening and deletion method for historical medical consultation data, the method comprising: Performing normalization processing on all medical consultation data based on a preset medical record template to obtain data types and their quantization values, and constructing a medical consultation array; wherein, the medical consultation array has the same identity identifier and time identifier as the medical consultation data; For any patient, querying their identity identifier, reading the medical consultation array of the patient according to the identity identifier and sorting it based on the time identifier to obtain a medical consultation array sequence; Determining the local importance degree of each data according to the medical consultation array sequence; Clustering all medical consultation arrays, determining the mean array of each type of medical consultation array, and determining the global importance degree of each medical consultation array according to the mean array; Determining the data deletion order according to the global importance degree of the medical consultation array and the local importance degree of each data, and performing the deletion process regularly and quantitatively based on the data deletion order.

[0005] As a further solution of the present invention: The step of performing normalization processing on all medical consultation data based on a preset medical record template to obtain data types and their quantization values, and constructing a medical consultation array includes: Sequentially reading type tags in the preset medical record template; Locate data in the medical interview data based on the type tags, and perform non-dimensionalization processing on the located data to obtain quantization values; Statistically analyze the quantization values according to the order of the type tags to obtain a medical interview array; each element in the medical interview array corresponds to a type tag; Query the identity identifier and time identifier in the medical interview data, and copy and insert them into the medical interview data; Among them, when the type tags in the medical record template have priorities, adjust the reading order of the type tags based on the priorities; statistically analyze the positioning results corresponding to each type tag in real time, calculate the mean value of the positioning results, and when a certain positioning process fails, use the mean value as the positioning result and synchronously insert a prompt tag.

[0006] As a further solution of the present invention: the step of querying the identity identifier of any patient, reading the medical interview array of this patient according to the identity identifier and sorting it based on the time identifier to obtain a medical interview array sequence includes: For any patient, query their identity identifier; Match the medical interview array of this patient in the medical interview array library according to the identity identifier; Sort the medical interview array of this patient according to the time identifier to obtain a medical interview array sequence.

[0007] As a further solution of the present invention: the step of determining the local importance of each data according to the medical interview array sequence includes: Randomly select a preset proportion of medical interview arrays in the medical interview array sequence as a feature set, and at the same time use all the remaining other medical interview arrays as a test set; Loop and execute a preset number of times to obtain a preset number of pairs of feature sets and test sets; For each pair of feature sets and test sets, train a regression model based on the feature set and determine the average error of calculating each data based on the test set; For each data, calculate its average error in each training process and calculate the mean value of the average error; Determine the local importance of each data of this patient based on the mean value of the average error.

[0008] As a further solution of the present invention: the step of clustering all medical interview arrays, determining the mean array of each class of medical interview arrays, and determining the global importance of each medical interview array according to the mean array includes: For all medical interview arrays within a preset time range, compare the medical interview arrays pairwise and calculate the similarity; Cluster each class of medical interview arrays based on the similarity. After clustering is completed, determine the mean array of each class of medical interview arrays at the same time; Input the mean array into a preset evaluation model to determine the evaluation value of the mean array; Adjust the evaluation value according to the number of arrays in each type of interrogation array to determine the global importance.

[0009] As a further solution of the present invention: The step of determining the data deletion order according to the global importance of the interrogation array and the local importance of each data, and performing the deletion process regularly and quantitatively based on the data deletion order includes: For each data of each patient, query the local importance of the data; Query the global importance of the interrogation array where the data is located; Sum the local importance and the global importance to obtain the comprehensive importance; Take the ascending order of the comprehensive importance as the data deletion order, and perform the deletion process regularly and quantitatively based on the data deletion order.

[0010] The technical solution of the present invention also provides a flexible screening system for historical interrogation data, and the system includes: An interrogation array construction module, configured to perform normalization processing on all interrogation data based on a preset medical record template to obtain data types and their quantization values, and construct an interrogation array; wherein, the interrogation array has the same identity identifier and time identifier as the interrogation data; An interrogation array sorting module, configured to query the identity identifier of any patient, read the interrogation array of the patient according to the identity identifier, and sort it based on the time identifier to obtain an interrogation array sequence; A local feature extraction module, configured to determine the local importance of each data according to the interrogation array sequence; A global feature extraction module, configured to cluster all interrogation arrays, determine the mean array of each type of interrogation array, and determine the global importance of each interrogation array according to the mean array; A data deletion module, configured to determine the data deletion order according to the global importance of the interrogation array and the local importance of each data, and perform the deletion process regularly and quantitatively based on the data deletion order.

[0011] As a further solution of the present invention: The interrogation array construction module includes: A type label reading unit, configured to sequentially read type labels in a preset medical record template; A data preprocessing unit, configured to locate data in the interrogation data based on the type label, and perform dimensionless processing on the located data to obtain a quantization value; A quantization value statistics unit, configured to count the quantization values according to the order of the type labels to obtain an interrogation array; each element in the interrogation array corresponds to a type label; An identifier insertion unit, configured to query the identity identifier and time identifier in the interrogation data, and copy and insert the interrogation data; Among them, when the type tags in the medical record template contain priorities, the reading order of the type tags is adjusted based on the priorities; the positioning results corresponding to each type tag are statistically counted in real time, the mean value of the positioning results is calculated, and when a certain positioning process fails, the mean value is used as the positioning result, and a prompt tag is inserted synchronously.

[0012] As a further solution of the present invention: the interrogation array sorting module includes: An identity identification query unit, which is used to query the identity identification of any patient; An array matching unit, which is used to match the interrogation array of the patient in the interrogation array library according to the identity identification; A sorting execution unit, which is used to sort the interrogation array of the patient according to the time identification to obtain an interrogation array sequence.

[0013] As a further solution of the present invention: the local feature extraction module includes: A feature set construction unit, which is used to randomly select a preset proportion of interrogation arrays in the interrogation array sequence as the feature set, and at the same time use all the remaining other interrogation arrays as the test set; A loop execution unit, which is used to loop a preset number of times to obtain a preset number of pairs of feature sets and test sets; A training unit, which is used to train a regression model based on the feature set for each pair of feature sets and test sets, and determine the average error of each data based on the test set; An error analysis unit, which is used to calculate the average error of each data in each training process and calculate the mean value of the average error; An error application unit, which is used to determine the local importance of each data of the patient based on the mean value of the average error.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention analyzes the interrogation data to determine the predictability of each data, and at the same time identifies the interrogation data itself to determine the degree of retention required. The importance of the data is jointly determined according to the predictability and the degree of retention required, and the data deletion order is determined according to the ascending order of importance, so that unimportant data is deleted first. Compared with the traditional regular deletion method, it is softer and more in line with the actual situation. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0016] Figure 1 It is a flowchart of a flexible screening method for historical interrogation data.

[0017] Figure 2 It is a block diagram of the composition structure of the flexible screening system for historical medical interview data. Specific implementation mode

[0018] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] Figure 1 It is a flow chart of the flexible screening method for historical medical interview data. In an embodiment of the present invention, a flexible screening method for historical medical interview data, the method includes: Step S100: Normalize all medical interview data based on a preset medical record template to obtain the data type and its quantization value, and construct a medical interview array; wherein, the medical interview array has the same identity identifier and time identifier as the medical interview data; Medical interview data is the interaction information between patients and medical staff in the hospital area. In the same hospital area, the medical record template is generally the same, but with the passage of time, the medical record template will also be updated. At this time, there will be a distinction between "old" and "new" medical interview data. Old data is the data statistically obtained using the previous template, and new data is the data statistically obtained using the latest template. When the templates are different, the order and format of these data may be different. Therefore, before processing the data, it is necessary to first perform a format unification process on all medical interview data. The way of the format unification process is to convert all medical interview data based on a unified medical record template, and unified data can be obtained. The unified medical record template generally adopts the latest medical record template.

[0020] After normalizing all medical interview data, a data conversion process is introduced to quantify each data in the template, mainly converting different strings into quantization values, counting each data type and its quantization value, and obtaining an array called the medical interview array. Each element in the medical interview array corresponds to a data type. In addition, the medical interview array has the same identity identifier and time identifier as the medical interview data, which is used to represent which patient's medical interview data the medical interview array is at what time.

[0021] Step S200: For any patient, query its identity identifier, read the medical interview array of the patient according to the identity identifier and sort it based on the time identifier to obtain a medical interview array sequence; For any patient, query their identity identifier. The simplest identity identifier is their ID number. All the medical consultation arrays are stored in the database. According to the patient's identity identifier, the medical consultation array of this patient can be retrieved. Sort the medical consultation array. The sorting order is chronological. Since the medical consultation array contains time identifiers, the sorting process is not complex. After sorting, the sorted medical consultation array of each patient is obtained, which is called the medical consultation array sequence.

[0022] Step S300: Determine the local importance of each data according to the medical consultation array sequence; Analyze the medical consultation array sequence of each patient, and the importance of each data can be evaluated. The analysis process is to compare each data with its corresponding data at different times. The obtained importance is called the local importance. Each data of each patient has a local importance.

[0023] Step S400: Cluster all the medical consultation arrays, determine the mean array of each type of medical consultation array, and determine the global importance of each medical consultation array according to the mean array; Cluster all the medical consultation arrays in the database, and multiple types of medical consultation arrays can be obtained. For each type of medical consultation array, calculate the mean array of all its medical consultation arrays, and analyze the mean array to determine the importance of this type of medical consultation array. At this time, the importance is the importance of a type of medical consultation array, which is called the global importance.

[0024] Step S500: Determine the data deletion order according to the global importance of the medical consultation array and the local importance of each data, and perform the deletion process regularly and quantitatively based on the data deletion order; It can be known from the above that each medical consultation array corresponds to a global importance, and each data in the medical consultation array corresponds to a local importance. For each data, take the global importance of the medical consultation array where it is located as the global importance of this data. Combining the global importance and the local importance, a comprehensive importance can be obtained. Determine the data deletion order from the comprehensive importance. The data deletion order is the reverse order of the comprehensive importance. The lower the comprehensive importance, the earlier the data deletion order. The meaning of performing the deletion process regularly and quantitatively based on the data deletion order is to delete a preset amount of data every preset time period.

[0025] It should be noted that in the actual application of the technical solution of the present invention, a time span will be determined in advance, such as a quarter. At this time, the present invention is applied to the data within a quarter, and the obtained medical consultation data is the medical consultation data within this quarter.

[0026] Regarding step S100, the steps of normalizing all the medical consultation data based on a preset medical record template to obtain the data type and its quantization value, and constructing the medical consultation array include: Read type tags in sequence from a preset medical record template; Locate data in the interrogation data based on the type tags, and perform non-dimensionalization processing on the located data to obtain quantization values; Statistically calculate the quantization values according to the order of the type tags to obtain an interrogation array; each element in the interrogation array corresponds to a type tag; Query the identity identifier and time identifier in the interrogation data, and copy and insert them into the interrogation data.

[0027] Read the latest medical record template, and read type tags in sequence from the latest medical record template. The type tags are the names of each data type. Locate data in the interrogation data within a preset time range (within one quarter) based on the type tags, and perform non-dimensionalization processing on the located data to obtain quantization values; actually, the non-dimensionalization processing of the located data is divided into two steps. The first step is to convert the located data into a numerical format. For strings, generally, the staff needs to pre-encode all the values of this type of string in advance to convert it into a numerical format; after the first step is completed, the second step is to convert the numerical value into a unitless data. The ways to achieve this include obtaining the maximum value and the minimum value of the numerical value, for any numerical value, calculating the difference between it and the minimum value as the actual difference, then calculating the difference between the maximum value and the minimum value as the span, and calculating the ratio of the actual difference to the span to obtain the non-dimensionalized numerical value; finally, statistically calculate the quantization values according to the order of the type tags to obtain an interrogation array. After obtaining the interrogation array, query the identity identifier and time identifier in the interrogation data, and copy and insert them into the interrogation data.

[0028] It is worth mentioning that when the type tags in the medical record template have priorities, adjust the reading order of the type tags based on the priorities, which affects the type tags corresponding to each element in the interrogation array; in addition, statistically calculate the positioning results corresponding to each type tag in real time, and calculate the mean value of the positioning results. When a certain positioning process fails, it means that the data of this type tag is not queried in the interrogation data of a certain patient. At this time, use the mean value as the positioning result and synchronously insert a prompt tag to inform the manual terminal, and the manual terminal will perform inspections later, which will not be elaborated here.

[0029] Regarding step S200, the steps of querying the identity identifier of any patient, reading the interrogation array of this patient according to the identity identifier, and sorting it based on the time identifier to obtain an interrogation array sequence include: For any patient, query their identity identifier; Match the interrogation array of this patient in the interrogation array library according to the identity identifier; Sort the interrogation array of this patient according to the time identifier to obtain an interrogation array sequence.

[0030] For any patient, query his identity identifier, match the patient's interrogation array in the interrogation array library according to the identity identifier. The interrogation array contains time identifiers. Determine the time order of the interrogation array according to the time identifiers, and count the interrogation array based on the time order to obtain an interrogation array sequence.

[0031] Regarding step S300, the step of determining the local importance of each data according to the interrogation array sequence includes: Randomly select a preset proportion of the interrogation arrays in the interrogation array sequence as the feature set, and at the same time use all the remaining other interrogation arrays as the test set; Execute the preset number of times in a loop to obtain the preset number of pairs of feature sets and test sets; For each pair of feature sets and test sets, train a regression model based on the feature set, and determine the average error of calculating each data based on the test set; For each data, calculate its average error in each training process, and calculate the mean value of the average error; Determine the local importance of each data of the patient based on the mean value of the average error.

[0032] The above content provides a specific calculation process of local importance. It uses the idea of a random forest to randomly select a preset proportion of the interrogation arrays in the interrogation array sequence. For example, assume that there are 20 interrogation arrays in the interrogation array sequence and the selection proportion is 80%. At this time, 16 interrogation arrays can be selected. These interrogation arrays are used as the feature set, and at the same time all the remaining other interrogation arrays are used as the test set (4 interrogation arrays); the selection process is random. By executing the preset number of times in a loop, different splitting methods can be obtained. That is, if one splitting result is used as a pair of feature sets and test sets, then multiple execution processes can obtain multiple pairs of feature sets and test sets; for each pair of feature sets and test sets, train a regression model based on the feature set, and with the help of the test set, the regression model can be tested to obtain the error rate; it should be noted that the regression model is used to process each data in the interrogation array. If the above example is used as a reference, the test set contains 4 interrogation arrays. For each data, it can be tested 4 times, and its mean value is calculated as the average error of the data; at this time, in each training process, each data has an average error, and the training process has been looped the preset number of times, and the obtained average errors have the preset number. Then calculate the mean value of its average error in each training process, and determine the local importance of each data of the patient according to the mean value of the average error.

[0033] It should be noted that for the random selection process, a condition may also need to be set, that is, each interrogation array needs to be used as an element in the test set at least once to prevent the situation where some data is only used as the training benchmark of the regression model.

[0034] The mean of the mean errors represents the predictability of the data (the smaller the mean of the mean errors, the higher the predictability). If the predictability is higher, the present application considers its importance to be lower because even if it is deleted, the accuracy of predicting it based on the remaining data is relatively high. Therefore, the local importance is directly proportional to the mean of the mean errors. Further, in the medical consultation array sequence of a patient, each type label corresponds to a set of data. Taking the above example as a reference, for the second element in 20 medical consultation arrays, 20 data can be extracted. For the local importance of any one of these 20 data, as can be known from the above content, it is directly proportional to the mean of the mean errors. On this basis, a condition is also introduced, that is, it is inversely proportional to the total number of remaining retained data among these 20 data, which means that if there are more data of the same type, there are more original data that can predict this data, and at this time, this data is less important. Generally speaking, for any data, the local importance is inversely proportional to the remaining quantity of data with the same type label. Regarding step S400, the steps of clustering all medical consultation arrays, determining the mean array of each category of medical consultation arrays, and determining the global importance of each medical consultation array according to the mean array include: For all medical consultation arrays within a preset time range, compare the medical consultation arrays pairwise and calculate the similarity; Cluster each category of medical consultation arrays based on the similarity. After clustering is completed, simultaneously determine the mean array of each category of medical consultation arrays; Input the mean array into a preset evaluation model to determine the evaluation value of the mean array; Adjust the evaluation value according to the number of arrays in each category of medical consultation arrays to determine the global importance.

[0035] Assume the time range is one quarter (it can also be longer, such as one year). For all medical consultation arrays within a preset time range, compare the medical consultation arrays pairwise and calculate the similarity. Cluster each category of medical consultation arrays based on the similarity. After clustering is completed, simultaneously determine the mean array of each category of medical consultation arrays. These processes can be applied with the simplest array operations. Input the mean array into a preset evaluation model to determine the evaluation value of the mean array. The evaluation model is an evaluation model for various indicators, and existing data importance evaluation criteria can be adopted. The output of the evaluation model is the evaluation value of the mean array, which is used to characterize the importance degree of a category of medical consultation arrays and is called the global importance.

[0036] On this basis, if the number of arrays in a type of interrogation array is larger and this situation is considered to be more valuable for analysis, a regulation coefficient greater than one can be provided and multiplied by the global importance to amplify the global importance. On the contrary, if it is considered that in this situation, the data samples are sufficient and their importance is not high, a regulation coefficient less than one can be provided and multiplied by the global importance to reduce the global importance.

[0037] It is worth mentioning that the same medical record template is adopted in the generation process of the interrogation array, and the dimensions of the interrogation arrays are the same, which enables the above array calculation process to proceed smoothly.

[0038] Regarding step S500, the steps of determining the data deletion order according to the global importance of the interrogation array and the local importance of each data and performing the deletion process regularly and quantitatively based on the data deletion order include: For each data of each patient, query the local importance of this data; Query the global importance of the interrogation array where this data is located; Sum the local importance and the global importance to obtain the comprehensive importance; Use the ascending order of the comprehensive importance as the data deletion order, and perform the deletion process regularly and quantitatively based on the data deletion order.

[0039] In an example of the technical solution of the present invention, for each data of each patient, query the local importance of this data, query the global importance of the interrogation array where this data is located, add the local importance and the global importance to obtain the comprehensive importance, use the ascending order of the comprehensive importance as the data deletion order, and perform the deletion process regularly and quantitatively based on the data deletion order, so that the less important data is deleted first. Among them, the specific parameters of regular and quantitative are determined by the staff according to the specific situation.

[0040] It is worth mentioning that since for any data, the local importance is inversely proportional to the remaining quantity of data with the same type of label, this makes the local importance of the remaining data change somewhat every time some data is deleted. Correspondingly, its comprehensive importance also changes, thus building a recursive adjustment process.

[0041] Figure 2 For the structural block diagram of the composition of the historical interrogation data flexible screening system, in an embodiment of the present invention, a historical interrogation data flexible screening system, the system 10 includes: An interrogation array construction module 11, configured to perform normalization processing on all interrogation data based on a preset medical record template to obtain the data type and its quantization value, and construct an interrogation array; wherein, the interrogation array has the same identity identifier and time identifier as the interrogation data; The medical interview array sorting module 12 is used to query the identity identifier of any patient, read the medical interview array of the patient according to the identity identifier, and sort it based on the time identifier to obtain a medical interview array sequence; The local feature extraction module 13 is used to determine the local importance of each data according to the medical interview array sequence; The global feature extraction module 14 is used to cluster all medical interview arrays, determine the mean array of each type of medical interview array, and determine the global importance of each medical interview array according to the mean array; The data deletion module 15 is used to determine the data deletion order according to the global importance of the medical interview array and the local importance of each data, and perform the deletion process regularly and quantitatively based on the data deletion order.

[0042] Furthermore, the medical interview array construction module 11 includes: The type label reading unit is used to sequentially read type labels in a preset medical record template; The data preprocessing unit is used to locate data in the medical interview data based on the type label, and perform dimensionless processing on the located data to obtain a quantization value; The quantization value statistics unit is used to statistically calculate the quantization values according to the order of the type labels to obtain a medical interview array; each element in the medical interview array corresponds to a type label; The identifier insertion unit is used to query the identity identifier and time identifier in the medical interview data, and copy and insert the medical interview data; Among them, when the type label in the medical record template has a priority, the reading order of the type label is adjusted based on the priority; the positioning results corresponding to each type label are statistically calculated in real time, the mean value of the positioning results is calculated, and when a certain positioning process fails, the mean value is used as the positioning result, and a prompt label is inserted synchronously.

[0043] Specifically, the medical interview array sorting module 12 includes: The identity identifier query unit is used to query the identity identifier of any patient; The array matching unit is used to match the medical interview array of the patient in the medical interview array library according to the identity identifier; The sorting execution unit is used to sort the medical interview array of the patient according to the time identifier to obtain a medical interview array sequence.

[0044] Even further, the local feature extraction module 13 includes: The feature set construction unit is used to randomly select a preset proportion of medical interview arrays in the medical interview array sequence as a feature set, and at the same time use all the remaining other medical interview arrays as a test set; The loop execution unit is used to loop a preset number of times to obtain a preset number of pairs of feature sets and test sets; A training unit, configured to, for each pair of a feature set and a test set, train a regression model based on the feature set and determine the average error of each data based on the test set; An error analysis unit, configured to, for each data, calculate the average error of the data in each training process and calculate the mean value of the average errors; An error application unit, configured to determine the local importance of each data of the patient based on the mean value of the average errors.

[0045] The foregoing are only preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A flexible screening method for historical medical consultation data, characterized in that: The method comprises: Based on the preset medical record template, all the consultation data are normalized to obtain the data type and its quantitative value, and a consultation array is constructed; wherein the consultation array and the consultation data have the same identity identifier and time identifier; For any patient, query his / her identity, read the patient's medical consultation array according to the identity and sort it based on the time stamp to obtain the medical consultation array sequence; Determine the local importance of each data according to the query array sequence; Clustering all the inquiry arrays, determining the mean array of each type of inquiry array, and determining the global importance of each inquiry array according to the mean array; The data deletion order is determined according to the global importance of the query array and the local importance of each data, and the deletion process is executed regularly and quantitatively based on the data deletion order.

2. The flexible screening method for historical medical consultation data according to claim 1 is characterized in that: The steps of normalizing all the medical consultation data based on the preset medical record template to obtain the data type and its quantitative value and constructing the medical consultation array include: Read the type labels in the preset medical record template in sequence; Locating data in the medical inquiry data based on the type label, performing dimensionless processing on the located data to obtain a quantized value; According to the sequential statistical quantization values ​​of the type labels, a medical consultation array is obtained; each element in the medical consultation array corresponds to a type label; Query the identity and time stamp in the medical consultation data, and copy and insert the medical consultation data; Among them, when the type tag in the medical record template contains a priority, the reading order of the type tag is adjusted based on the priority; the positioning results corresponding to each type tag are counted in real time, and the mean of the positioning results is calculated. When a positioning process fails, the mean is used as the positioning result, and the prompt tag is inserted synchronously.

3. The flexible screening method for historical medical consultation data according to claim 1 is characterized in that: The steps of querying the identity of any patient, reading the patient's medical inquiry array according to the identity and sorting the array based on the time identifier to obtain the medical inquiry array sequence include: For any patient, query their identity; Matching the patient's medical inquiry array in the medical inquiry array library according to the identity identifier; The patient's medical consultation array is sorted according to the time stamp to obtain a medical consultation array sequence.

4. The flexible screening method for historical medical consultation data according to claim 1 is characterized in that: The step of determining the local importance of each data according to the query array sequence comprises: Randomly select a preset proportion of question arrays from the question array sequence as the feature set, and use all other question arrays as the test set; The loop is executed for a preset number of times to obtain a preset number of pairs of feature sets and test sets; For each pair of feature set and test set, a regression model is trained based on the feature set, and an average error of each data is calculated based on the test set; For each data, calculate its average error in each training process and calculate the mean of the average error; The local importance of each data of the patient is determined based on the mean of the average errors.

5. The flexible screening method for historical medical consultation data according to claim 1 is characterized in that: The step of clustering all the inquiry arrays, determining the mean array of each type of inquiry array, and determining the global importance of each inquiry array according to the mean array comprises: For all the medical consultation arrays within the preset time range, the medical consultation arrays are compared pairwise and the similarity is calculated; Cluster each type of medical consultation array based on similarity. After clustering, determine the mean array of each type of medical consultation array. Inputting the mean value array into a preset evaluation model to determine an evaluation value of the mean value array; The evaluation value is adjusted according to the number of arrays of each type of consultation array to determine the global importance.

6. The flexible screening method for historical medical consultation data according to claim 1 is characterized in that: The step of determining the data deletion order according to the global importance of the query array and the local importance of each data, and executing the deletion process based on the data deletion order in a timely and quantitative manner includes: For each data of each patient, query the local importance of the data; Query the global importance of the consultation array where the data is located; Add the local importance and the global importance to get the comprehensive importance; The ascending order of comprehensive importance is used as the data deletion order, and the deletion process is executed regularly and quantitatively based on the data deletion order.

7. A flexible screening system for historical medical consultation data, characterized in that: The system comprises: A medical inquiry array construction module is used to normalize all medical inquiry data based on a preset medical record template, obtain the data type and its quantitative value, and construct a medical inquiry array; wherein the medical inquiry array and the medical inquiry data have the same identity identifier and time identifier; The medical inquiry array sorting module is used to query the identity of any patient, read the patient's medical inquiry array according to the identity, and sort it based on the time mark to obtain the medical inquiry array sequence; A local feature extraction module is used to determine the local importance of each data according to the question array sequence; A global feature extraction module, used for clustering all the inquiry arrays, determining the mean array of each type of inquiry array, and determining the global importance of each inquiry array according to the mean array; The data deletion module is used to determine the data deletion order according to the global importance of the query array and the local importance of each data, and execute the deletion process in a timely and quantitative manner based on the data deletion order.

8. The flexible screening system for historical medical consultation data according to claim 7 is characterized in that: The inquiry array building module includes: A type label reading unit, used to read the type labels in the preset medical record template in sequence; A data preprocessing unit, used for locating data in the medical inquiry data based on the type label, performing dimensionless processing on the located data, and obtaining a quantized value; A quantitative value statistical unit is used to count the quantitative values ​​according to the order of the type labels to obtain a medical inquiry array; each element in the medical inquiry array corresponds to a type label; An identification inserting unit, used to query the identity identification and time identification in the medical inquiry data, and copy and insert the medical inquiry data; Among them, when the type tag in the medical record template contains a priority, the reading order of the type tag is adjusted based on the priority; the positioning results corresponding to each type tag are counted in real time, and the mean of the positioning results is calculated. When a positioning process fails, the mean is used as the positioning result, and the prompt tag is inserted synchronously.

9. The flexible screening system for historical medical consultation data according to claim 7, characterized in that: The inquiry array sorting module comprises: An identity query unit, used to query the identity of any patient; An array matching unit, used for matching the patient's medical inquiry array in the medical inquiry array library according to the identity identifier; The sorting execution unit is used to sort the patient's medical inquiry array according to the time mark to obtain a medical inquiry array sequence.

10. The flexible screening system for historical medical consultation data according to claim 7, characterized in that: The local feature extraction module comprises: A feature set construction unit, used for randomly selecting a preset proportion of question arrays in the question array sequence as a feature set, and taking all other question arrays as a test set; A loop execution unit, used for loop execution for a preset number of times, to obtain a preset number of pairs of feature sets and test sets; A training unit, used for training a regression model based on the feature set for each pair of feature set and test set, and determining and calculating an average error of each data based on the test set; The error analysis unit is used to calculate the average error of each data in each training process and the mean of the average error; The error application unit is used for determining the local importance of each data of the patient based on the mean value of the average error.

Citation Information

Patent Citations

  • Data monitoring and processing method based on software and hardware all-in-one machine

    CN112148686A

  • Electric charge data processing method and device, terminal equipment and storage medium

    CN113554527A

  • Cache capacity reduction method and device, equipment and storage medium

    CN114546891A

  • Duplicated data deletion method for medical big data

    CN114722013A

  • Medical diagnosis data management method and platform based on block chain

    CN117524394A