Data screening method and system based on medical clinical big data analysis
By extracting and merging features from clinical manifestation data, datasets with high similarity to archived medical records are selected, solving the problem of difficulty in associating data features in traditional methods and improving diagnostic efficiency and quality.
Patent Information
- Application Number
- CN202510485333.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional medical clinical data screening methods are insufficient to fully explore data characteristics and accurately correlate similar cases, making it difficult for doctors to quickly obtain reliable reference information during diagnosis, thus affecting the efficiency and quality of diagnosis and treatment.
By extracting features from the clinical manifestation data to be analyzed, a clinical manifestation feature encoding vector is obtained. This vector is then merged and integrated with the medical record data to obtain the feature vector of the medical record to be analyzed. The distance between the feature vector and the feature vector of the archived medical records is calculated, and the archived clinical manifestation data whose distance does not exceed a preset threshold is selected as the target dataset.
It enables precise data screening, improves the efficiency of clinical data analysis, and provides more targeted diagnostic and treatment reference information.
Smart Images

Figure CN120236781B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of big data, in particular to a data screening method and system based on medical clinical big data analysis. BACKGROUND
[0002] In the medical field, with the accumulation of a large amount of clinical data, how to efficiently screen valuable information for diagnosis and treatment from massive data has become a key problem. Traditional data screening methods cannot fully mine data characteristics and accurately associate similar cases. In the face of complex and diverse clinical manifestations and medical record information of different patients, the lack of accurate and effective analysis methods makes it difficult for doctors to quickly obtain reliable reference basis when diagnosing, affecting the efficiency and quality of diagnosis and treatment. SUMMARY
[0003] The purpose of the present application is to provide a data screening method and system based on medical clinical big data analysis.
[0004] In a first aspect, the present application provides a data screening method based on medical clinical big data analysis, comprising:
[0005] performing feature extraction on the to-be-analyzed clinical manifestation data to obtain a clinical manifestation feature encoding vector;
[0006] obtaining medical record data, performing feature merging on the medical record data and the clinical manifestation feature encoding vector to obtain a multi-dimensional merged vector, performing integration processing on the multi-dimensional merged vector to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements; the medical record data is a description text used to describe the patient manifestation information corresponding to the extracted clinical manifestation data, the multi-dimensional merged vector includes expression content used to represent the patient manifestation information of the to-be-analyzed clinical manifestation data, and the preset integration number is equal to the number of feature elements of an archived medical record feature vector corresponding to an archived clinical manifestation data in a preset medical record archive;
[0007] obtaining an archived medical record feature vector corresponding to the archived clinical manifestation data; the archived medical record feature vector includes expression content used to represent the patient manifestation information of the archived clinical manifestation data;
[0008] obtaining a medical record feature distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector, determining the archived clinical manifestation data corresponding to the archived medical record feature vector with a medical record feature distance not exceeding a preset medical record feature distance threshold as a target clinical manifestation data set; the target clinical manifestation data set is used to determine a target clinical manifestation data screening result for the to-be-analyzed clinical manifestation data.
[0009] In a second aspect, an embodiment of the present application provides a server system, comprising a server configured to execute the method of the first aspect.
[0010] Compared with the prior art, the present application has the beneficial effects that: by using the data screening method and system based on medical clinical big data analysis, the feature coding vector is obtained by extracting the features of the to-be-analyzed clinical performance data, and then the medical record data is combined and integrated with the feature coding vector to obtain the to-be-analyzed medical record feature vector. Then, the archived medical record feature vector is obtained, and the distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector is calculated. The archived clinical performance data with a distance not exceeding a preset threshold is determined as the target data set, so as to realize accurate screening and improve the analysis efficiency of the clinical data. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0012] Figure 1 The step flowchart of the data screening method based on medical clinical big data analysis provided by the embodiment of the present application is shown in the figure.
[0013] Figure 2 The structural schematic block diagram of the computer device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.
[0015] The specific embodiments of the present application will be described in detail below in combination with the drawings.
[0016] In order to solve the technical problems in the foregoing background art, Figure 1 The flowchart of the data screening method based on medical clinical big data analysis provided by the embodiment of the present application is shown in the figure, and the data screening method based on medical clinical big data analysis will be described in detail.
[0017] Step S201, feature extraction is performed on the to-be-analyzed clinical performance data to obtain a clinical performance feature coding vector;
[0018] In step S202, medical record data is acquired, feature merging is performed on the medical record data and the clinical manifestation feature code vector to obtain a multi-dimensional merged vector, the multi-dimensional merged vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements; the medical record data is a description text used to describe the patient manifestation information corresponding to the extracted clinical manifestation data, the multi-dimensional merged vector includes expression content used to represent the patient manifestation information of the to-be-analyzed clinical manifestation data, and the preset integration number is equal to the number of feature elements of an already-archived medical record feature vector corresponding to already-archived clinical manifestation data in a preset medical record archive;
[0019] In step S203, an already-archived medical record feature vector corresponding to the already-archived clinical manifestation data is acquired; the already-archived medical record feature vector includes expression content used to represent the patient manifestation information of the already-archived clinical manifestation data.
[0020] In step S204, a medical record feature distance between the already-archived medical record feature vector and the to-be-analyzed medical record feature vector is obtained, and the already-archived clinical manifestation data corresponding to the already-archived medical record feature vector with a medical record feature distance not exceeding a preset medical record feature distance threshold is determined as a target clinical manifestation data set; the target clinical manifestation data set is used to determine a target clinical manifestation data screening result for the to-be-analyzed clinical manifestation data.
[0021] In an embodiment of the present application, it is assumed that there is a large general hospital which has a huge medical clinical database storing a large amount of patient medical record data and corresponding clinical manifestation data. The hospital hopes to use these data to achieve more accurate data filtering through the server, so as to provide more targeted reference information for doctors' diagnosis and treatment. In this scenario, the server undertakes the important task of executing the entire data filtering method. In the daily diagnosis and treatment process of the hospital, new patient cases will continue to be generated. For example, a patient named Mr. Li comes to the hospital for treatment due to symptoms such as persistent fever, cough, and fatigue. The doctor performs a series of examinations on Mr. Li, including body temperature monitoring, blood tests, chest X-rays, etc. The data obtained from these examinations and the doctor's detailed records of Mr. Li's symptoms constitute the to-be-analyzed clinical manifestation data. After receiving the to-be-analyzed clinical manifestation data, the server begins to perform feature extraction. The server loads the to-be-analyzed clinical manifestation data of Mr. Li into the target integrated model. This target integrated model is a model developed and trained by the hospital for a long time and is specially used for processing this kind of medical clinical data analysis. It contains multiple feature extraction components, which are further divided into input sub-components, feature extraction sub-components, and feature optimization sub-components. Data interception: In the input sub-component, the server first performs data interception according to the monitoring data time range and monitoring data update frequency of the to-be-analyzed clinical manifestation data of Mr. Li. For example, Mr. Li's body temperature monitoring data is recorded every hour, and it has lasted for two days. Therefore, the server will intercept an appropriate data segment from a large amount of body temperature monitoring data according to the time range and update frequency, and may select the data of several time periods with obvious body temperature fluctuations to obtain multiple intercepted monitoring data. Data division: Then, the server divides the multiple intercepted monitoring data. Taking blood test data as an example, the server may divide the blood test data into multiple clinical monitoring sub-data according to different blood indicators, such as white blood cell count, red blood cell count, and platelet count. Each sub-data corresponds to a specific blood indicator, and its monitoring data time range and monitoring data update frequency are also clear. Feature extraction: In the feature extraction sub-component, the server performs feature extraction on the divided multiple clinical monitoring sub-data. For example, for the clinical monitoring sub-data of white blood cell count, the server analyzes the numerical value change at different time points and the deviation degree from the normal range, thereby obtaining a clinical monitoring feature vector of white blood cell count. The same feature extraction operation is also performed on other clinical monitoring sub-data, such as red blood cell count, platelet count, and symptom-related data such as body temperature and cough, and finally multiple clinical monitoring feature vectors are obtained. Dimension conversion: Then, the server converts the feature elements indicating the spatial dimension channel in the multiple clinical monitoring feature vectors to feature dimension channels according to the mapping coefficients.The mapping coefficients here are determined through complex calculation and analysis according to the number of a large number of previous clinical monitoring feature vectors. Since the number of characteristic elements of the clinical monitoring feature vector is often greater than that of the characteristic elements of the to-be-analyzed clinical performance feature vector, through such dimension conversion, the data can be more suitable for subsequent processing. For example, a certain clinical monitoring feature vector originally has 10 characteristic elements, and after conversion, it may become 5 characteristic elements of the to-be-analyzed clinical performance feature vector, so that multiple to-be-analyzed clinical performance feature vectors are obtained. Obtain a low-rank joint compression vector: the server first obtains a low-rank joint compression vector, which is a trainable parameter obtained by the model in the previous large amount of data training process, and plays an important role in subsequent feature merging and optimization. Feature merging: the server merges the to-be-analyzed clinical performance feature vector and the low-rank joint compression vector to obtain a clinical performance feature vector set. For example, the to-be-analyzed clinical performance feature vector obtained previously about Mr. Li's various symptoms and examination indicators is merged with the low-rank joint compression vector according to certain rules to form a clinical performance feature vector set containing various feature information. Execute operation module processing: the server places the low-rank joint compression vector and the clinical performance feature vector set into multiple execution operation modules, including a target execution operation module. In the target execution operation module: first, multiply the low-rank joint compression vector and the state vector to obtain a gating signal. Assuming that the low-rank joint compression vector is [1, 2, 3] and the state vector is [4, 5, 6], the gating signal obtained after multiplication may be [4, 10, 18]. Next, multiply the clinical performance feature vector set and the input feature vector set to obtain a state representation. For example, a certain vector in the clinical performance feature vector set is [2, 3, 4], and the corresponding vector in the input feature vector set is [5, 6, 7], and the state representation obtained after multiplication may be [10, 18, 28]. Then, multiply the clinical performance feature vector set and the output feature vector set to obtain transformed features. For example, another vector in the clinical performance feature vector set is [3, 4, 5], and the corresponding vector in the output feature vector set is [6, 7, 8], and the transformed features obtained after multiplication may be [18, 28, 40]. Gating weight vector processing: the server obtains a gating weight vector according to the gating signal and the state representation. Then, the gating weight vector is processed by dimension reduction according to the correlation measure of the state representation, for example, by calculating the correlation between elements in the state representation, the elements with weak correlation in the gating weight vector are removed or merged to realize dimension reduction. After that, the dimension-reduced gating weight vector is normalized to obtain a gating control coefficient. Assuming that the dimension-reduced gating weight vector is [0.2, 0.3, 0.5], the gating control coefficient obtained after normalization may be [0.2 / 1, 0.3 / 1, 0.5 / 1], i.e. [0.2, 0.3, 0.5].Finally, the clinical manifestation feature encoding vector is obtained: the server multiplies the gating control coefficient with the transformed features to obtain the gating output vector corresponding to the target execution operation module. For example, if the gating control coefficient is [0.2, 0.3, 0.5] and the transformed features are [18, 28, 40], the gating output vector obtained after multiplication is [3.6, 8.4, 20]. According to the multiple gating output vectors and the gating control coefficients corresponding to the multiple execution operation modules, a series of complex calculations and integrations are performed to finally obtain the clinical manifestation feature encoding vector. The number of feature elements in this clinical manifestation feature encoding vector is the same as that of the low-rank joint compression vector, and it contains the concentrated feature information of Mr. Li's clinical manifestation data after a series of processing. The server loads the previously obtained clinical manifestation feature encoding vector of Mr. Li into the target integrated model. This target integrated model also contains a multi-dimensional feature extraction component and a feature integration processing component, where the multi-dimensional feature extraction component includes a feature extraction component. Obtain medical record data: the server obtains Mr. Li's medical record data from the hospital's medical record database. This medical record data is a description text used to describe the patient's performance information corresponding to the extracted clinical manifestation data, which details Mr. Li's basic information, medical history, diagnosis and treatment process, doctor's diagnosis conclusion, etc. Feature extraction: the server performs feature extraction on the obtained medical record data to obtain a description text feature vector. For example, for the description of Mr. Li's medical history in the medical record data, the server analyzes the disease types, onset time, severity of illness, etc. mentioned in the description and converts these information into corresponding feature vectors; for the description of the diagnosis and treatment process, the server analyzes the treatment methods, treatment time, treatment effect, etc. and converts them into feature vectors, and finally obtains a description text feature vector. Feature merging: the server merges the description text feature vector with the clinical manifestation feature encoding vector to obtain a multi-dimensional merged vector. Assuming that the clinical manifestation feature encoding vector is [3.6, 8.4, 20] and the description text feature vector is [5, 6, 7], the multi-dimensional merged vector obtained by some merging rule (such as adding corresponding elements) may be [8.6, 14.4, 27]. This multi-dimensional merged vector includes the expression content used to represent the patient's performance information of Mr. Li's clinical manifestation data to be analyzed. In the feature integration processing component, mode one: processing through the feature compression and reconstruction component. 1. Load into the feature integration processing component: the server loads the multi-dimensional merged vector into the feature integration processing component, which includes a feature compression component and a feature reconstruction component. The feature reconstruction component is trained according to the pre-set integration number, and the number of feature elements output is fixed at the pre-set integration number. 2. Feature compression processing: in the feature compression component, the server performs feature compression processing on the multi-dimensional merged vector to obtain a compressed feature vector.For example, a certain compression algorithm (such as principal component analysis, etc.) is performed on the multi-dimensional merged vector [8.6, 14.4, 27], and a compressed feature vector [4, 7, 12] can be obtained. Then, according to the sequence processing step, the compressed feature vector is processed to obtain a reduced dimension feature vector. Assuming that the sequence processing step is 2, the compressed feature vector [4, 7, 12] is processed by a certain dimension reduction algorithm (such as selecting one element every other element), and a reduced dimension feature vector [4, 12] is obtained. 3. Feature reconstruction processing: In the feature reconstruction component, the server obtains a feature vector of the medical record to be analyzed with a preset integration number of feature elements according to the feature correlation of the reduced dimension feature vector in each sequence processing step. For example, by analyzing the correlation of the reduced dimension feature vector [4, 12] with the original multi-dimensional merged vector at different sequence processing steps, and the correlation between the elements, after complex calculation and adjustment, a feature vector of the medical record to be analyzed with a preset integration number of feature elements is finally obtained, such as [3, 5]. Way two: processing by clinical term analysis and feature weight adjustment component: 1. Load to feature integration processing component: the server also loads the multi-dimensional merged vector to the feature integration processing component, which includes the clinical term analysis component and the feature weight adjustment component. 2. Clinical term analysis component processing: in the clinical term analysis component, the server obtains the clinical correlation score corresponding to each clinical term in the clinical term library according to the multi-dimensional merged vector. For example, for the part of the multi-dimensional merged vector related to Mr. Li's cough symptoms, the server will search for terms related to cough in the clinical term library, such as "dry cough" and "cough sputum", and determine the clinical correlation score of these terms according to the information in the multi-dimensional merged vector. Then, according to the clinical correlation score, the patient performance information of Mr. Li's performance data to be analyzed is obtained, and the patient performance information is segmented according to the preset integration number to obtain multiple clinical performance description segments. Assuming that the preset integration number is 5, by reasonably segmenting Mr. Li's patient performance information, 5 clinical performance description segments are obtained, such as "continuous fever", "dry cough", "weakness", "chest pain", and "rapid breathing". 3. Obtain relevant indicators and calculate influence factors: the server obtains the frequency of occurrence of each of the multiple clinical performance description segments in the patient performance information, and the clinical term rarity index of each of the multiple clinical performance description segments in the patient performance information. For example, "continuous fever" has a high frequency of occurrence in Mr. Li's patient performance information, while "chest pain" has a relatively low frequency of occurrence; "dry cough" is relatively common in the clinical term library, while "rapid breathing" is relatively rare. Then, according to the multiple frequency of occurrence and the multiple clinical term rarity index, multiple clinical performance influence factors are obtained. Assuming that the frequency of occurrence of "continuous fever" is 0.3 and the clinical term rarity index is 0.2, according to a certain calculation formula (such as multiplication), the clinical performance influence factor of "continuous fever" is 0.06.4. Feature weight adjustment component processing: In the feature weight adjustment component, the server multiplies each of the plurality of clinical manifestation influence factors with the multi-dimensional merged vector to obtain a plurality of clinical manifestation weight vectors. For example, the clinical manifestation influence factor 0.06 of "persistent fever" is multiplied with the multi-dimensional merged vector [8.6, 14.4, 27] to obtain a clinical manifestation weight vector [0.516, 0.864, 1.62]. According to the plurality of clinical manifestation weight vectors, a feature element number of the to-be-analyzed medical record feature vector is obtained. Through comprehensive processing (such as weighted summation) of each clinical manifestation weight vector, a to-be-analyzed medical record feature vector with a preset integration number of feature elements is finally obtained, such as [2, 4]. Method three: directly determining the last transition vector as the to-be-analyzed medical record feature vector: 1. Transition vector situation of multi-dimensional merged vector: Assuming that the multi-dimensional merged vector includes a plurality of transition vectors, the feature element number of each of the plurality of transition vectors is the preset integration number, and the last transition vector in the plurality of transition vectors includes the global expression content obtained by performing feature merging on the medical record data and the clinical manifestation feature code vector. 2. Determine the to-be-analyzed medical record feature vector: In the feature integration processing component, the server determines the last transition vector in the plurality of transition vectors as the to-be-analyzed medical record feature vector with the preset integration number of feature elements. For example, the plurality of transition vectors are [1, 2, 3], [4, 5, 6], and [7, 8, 9], and the last transition vector [7, 8, 9] is determined as the to-be-analyzed medical record feature vector. The server obtains the archived medical record feature vector corresponding to the archived clinical manifestation data from the preset medical record archive of the hospital. These archived medical record feature vectors also include expression content for representing patient manifestation information of the archived clinical manifestation data. For example, in the preset medical record archive, there is a medical record of Ms. Wang who has a similar symptom, and she also has symptoms such as fever and cough. The server will obtain the archived medical record feature vector corresponding to the medical record of Ms. Wang, which contains condensed feature information of the symptoms, diagnosis and treatment process, and other related information of Ms. Wang at that time. The server loads the archived medical record feature vector and the to-be-analyzed medical record feature vector into the target integrated model, which includes a feature retrieval component. According to the values of each feature element in the archived medical record feature vector, the server obtains a first retrieval parameter of the archived medical record feature vector; according to the values of each feature element in the to-be-analyzed medical record feature vector, the server obtains a second retrieval parameter of the to-be-analyzed medical record feature vector. For example, the archived medical record feature vector is [2, 4, 6], and the first retrieval parameter is a numerical value obtained by performing certain operations (such as summation, multiplication, etc.) on these elements; the to-be-analyzed medical record feature vector is [3, 5, 7], and the second retrieval parameter is also obtained by similar means. The server multiplies the archived medical record feature vector and the to-be-analyzed medical record feature vector to obtain a retrieval parameter coefficient.Then, according to the search parameter coefficient, the first search parameter and the second search parameter, a medical record feature distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector is obtained through a specific calculation formula (such as a certain distance calculation formula). Assuming that the archived medical record feature vector is [2, 4, 6], the to-be-analyzed medical record feature vector is [3, 5, 7], and after a series of calculations, the obtained medical record feature distance is 5. The server first obtains a preset medical record feature distance threshold. The method for determining the threshold is as follows: the server obtains a plurality of archived medical record feature vectors corresponding to the archived clinical manifestation data in the preset medical record archive, and calculates the average medical record feature distance between each set of archived medical record feature vectors. For example, 10 sets of archived medical record feature vectors are selected from the archive, the medical record feature distances between them are calculated respectively, and then the average value is obtained to obtain the average medical record feature distance. Then, the average medical record feature distance is adjusted according to the preset proportion coefficient to obtain the preset medical record feature distance threshold. Assuming that the average medical record feature distance is 8 and the preset proportion coefficient is 0.75, the preset medical record feature distance threshold is 6. The server determines the archived clinical manifestation data corresponding to the archived medical record feature vector whose medical record feature distance does not exceed the preset medical record feature distance threshold as the target clinical manifestation data set. For example, in the above calculation, the medical record feature distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector is 5, and the preset medical record feature distance threshold is 6, so the archived clinical manifestation data corresponding to the archived medical record feature vector is determined as the target clinical manifestation data set. This target clinical manifestation data set can be used to determine the target clinical manifestation data screening result for the to-be-analyzed clinical manifestation data of Mr. Li, and the doctor can refer to the diagnosis and treatment of similar cases in the past to provide more targeted suggestions for the diagnosis and treatment of Mr. Li.
[0022] In the embodiments of the present application, the feature extraction of the to-be-analyzed clinical manifestation data to obtain a clinical manifestation feature encoding vector can be implemented through the following examples.
[0023] Load the to-be-analyzed clinical manifestation data into a target integrated model; the target integrated model includes a multi-dimensional feature extraction component, and the multi-dimensional feature extraction component includes an input subcomponent, a feature extraction subcomponent and a feature optimization subcomponent;
[0024] In the input subcomponent, according to the monitoring data time range and the monitoring data update frequency of the to-be-analyzed clinical manifestation data, the to-be-analyzed clinical manifestation data is subjected to data interception to obtain a plurality of intercepted monitoring data, and the plurality of intercepted monitoring data are subjected to data division respectively to obtain a plurality of clinical monitoring sub-data monitoring data time range and monitoring data update frequency;
[0025] In the feature extraction subassembly, feature extraction is performed on the plurality of clinical monitoring sub-data to obtain a plurality of clinical monitoring feature vectors, and according to a mapping coefficient, feature elements in the plurality of clinical monitoring feature vectors indicating spatial dimension channels are re-converted to feature dimension channels to obtain a plurality of to-be-analyzed clinical performance feature vectors; the mapping coefficient is determined according to the number of the clinical monitoring feature vectors, and the number of feature elements of the clinical monitoring feature vectors is greater than the number of feature elements of the to-be-analyzed clinical performance feature vectors.
[0026] In the feature optimization subassembly, a low-rank joint compression vector is used to perform a gating mechanism process on the to-be-analyzed clinical performance feature vector to obtain a clinical performance feature encoding vector; the number of feature elements of the clinical performance feature encoding vector is the same as the number of feature elements of the low-rank joint compression vector.
[0027] In the embodiments of the present application, an example is now assumed that a patient named Ms. Zhang comes to the hospital for treatment, who has symptoms of headache, dizziness, blurred vision, and occasional palpitations. The doctor arranges a series of examinations for her, including blood pressure monitoring (measured every half hour for a day), brain CT scan, eye examination, and electrocardiogram examination, etc. The data generated by these examinations and the detailed records of Ms. Zhang's symptoms by the doctor together constitute the analyzed clinical performance data to be analyzed. After receiving the analyzed clinical performance data of Ms. Zhang, the server starts to work on feature extraction. First, the server will load these analyzed clinical performance data to the target integrated model. This target integrated model is a comprehensive analysis model developed by the hospital after a long period of research and development, a large amount of data training, and continuous optimization, which is specially used to process medical clinical related data. Its internal structure design is very fine, which contains a multi-dimensional feature extraction component, and this multi-dimensional feature extraction component is further subdivided into input subcomponent, feature extraction subcomponent and feature optimization subcomponent, each subcomponent bears a specific function, and works together to realize the conversion from raw data to feature encoding vector. In the input subcomponent, the server will intercept the data according to the monitoring data time range and monitoring data update frequency of Ms. Zhang's analyzed clinical performance data, to obtain more representative and valuable data segments. Taking Ms. Zhang's blood pressure monitoring data as an example, since it is measured every half hour for a day, the server will analyze the blood pressure fluctuation in this whole period. Considering that there may be some time periods with obvious blood pressure fluctuations or special change trends, the server will intercept the data of several key time periods. For example, the server finds that Ms. Zhang's blood pressure fluctuation amplitude is relatively large in the two time periods of 9:00 to 11:00 am and 3:00 to 5:00 pm, and shows certain regularity. Therefore, the server will intercept the blood pressure monitoring data in these two time periods as part of the multiple intercepted monitoring data. Similarly, for the brain CT scan data, although it is not like blood pressure monitoring that has frequent updates in time series, the server will also select some key layer scan image data as intercepted monitoring data according to its own characteristics and relevance with other symptoms. For example, select those scan image data that can clearly show the brain blood vessel condition and the brain region that may be related to the symptoms of headache and dizziness. After completing the data interception, the server will divide the multiple intercepted monitoring data respectively to analyze the features of each data segment in more detail. For the intercepted blood pressure monitoring data, the server will divide it according to different attributes.For example, the blood pressure values are divided into systolic and diastolic pressure, and the measurement time corresponding to each blood pressure value is recorded, thus forming multiple clinical monitoring sub-data, and the monitoring data time range (e.g., systolic pressure data from 9 am to 11 am) and monitoring data update frequency (updated every half hour) of each sub-data are clear. For the intercepted monitoring data of brain CT scan, the server will divide it according to different tissue structures and possible pathological features in the image. For example, the part of the brain CT image related to the brain blood vessels is divided into a clinical monitoring sub-data, and the corresponding scanning layer information and relationship with the whole image are recorded; the part related to the brain nerve tissue is also divided into a clinical monitoring sub-data, and the related image features and position information are also clear. Through such data division operation, the server further refines the originally complex intercepted monitoring data into multiple clinical monitoring sub-data with clear features and attributes, laying a foundation for subsequent feature extraction work. In the feature extraction sub-component, the server will perform in-depth feature extraction on the divided multiple clinical monitoring sub-data to mine the key feature information contained in each sub-data. Taking the systolic blood pressure clinical monitoring sub-data divided from the blood pressure monitoring data as an example, the server will analyze the numerical change of the systolic blood pressure at different time points. For example, observe how the systolic blood pressure changes over time in the time period from 9 am to 11 am, whether it gradually increases, gradually decreases, or shows a fluctuating state. At the same time, the server will also compare these systolic blood pressure values with the normal blood pressure range to calculate the degree of deviation from the normal range. Through these analyses, the server can extract a series of features for this systolic blood pressure clinical monitoring sub-data, such as blood pressure change trend feature, deviation from normal range feature, etc., and convert these features into a clinical monitoring feature vector. For the brain blood vessels clinical monitoring sub-data divided from the brain CT scan, the server will analyze the shape, thickness, whether there is blockage or deformation, etc. of the brain blood vessels in the image. Through quantification and coding of these features, the server can also extract the corresponding clinical monitoring feature vector for this brain blood vessels clinical monitoring sub-data. For example, the thickness of the brain blood vessels is represented by a numerical value, and whether there is blockage is represented by a binary value (0 represents no blockage, 1 represents blockage), and then these feature values are combined to form a clinical monitoring feature vector. Through similar feature extraction operations on each clinical monitoring sub-data, the server ultimately obtains multiple clinical monitoring feature vectors, each of which represents the key feature information of the corresponding clinical monitoring sub-data. After obtaining multiple clinical monitoring feature vectors, the server needs to convert the feature elements indicating the spatial dimension path in these vectors to the feature dimension channel according to the mapping coefficient to obtain multiple clinical performance feature vectors to be analyzed. The mapping coefficient is determined through complex calculation and analysis according to the number of previous clinical monitoring feature vectors.Since the number of feature elements of the clinical monitoring feature vector is often greater than that of the clinical performance feature vector to be analyzed, this dimension conversion helps to simplify the data structure and make it more suitable for subsequent processing. For example, assume that a certain clinical monitoring feature vector originally has 8 feature elements, which describe the characteristics of a certain clinical monitoring sub-data from different spatial dimension angles. After dimension conversion according to the mapping coefficients, some of the elements may be combined or converted to the feature dimension channel according to certain rules, and finally a clinical performance feature vector to be analyzed with 5 feature elements is obtained. In this way, by performing similar dimension conversion operations on all obtained clinical monitoring feature vectors, the server can obtain multiple clinical performance feature vectors to be analyzed, each of which contains optimized and converted key feature information, which is more convenient for subsequent processing in the feature optimization subcomponent. In the feature optimization subcomponent, the server will process the previously obtained clinical performance feature vector to be analyzed according to the low-rank joint compression vector to obtain a clinical performance feature encoding vector. The number of feature elements of this clinical performance feature encoding vector is the same as that of the low-rank joint compression vector, and it will serve as an important basis for subsequent data processing, condensing the key feature information of the clinical performance data to be analyzed after a series of processing. The server first needs to obtain the low-rank joint compression vector. This low-rank joint compression vector is a trainable parameter that the target integrated model gradually learns and optimizes during the previous processing of a large amount of training data. It plays a crucial role in the entire gating mechanism processing process, and through interaction with the clinical performance feature vector to be analyzed, it can effectively filter and optimize feature information. For example, during the long-term data training process in the hospital, the model analyzes and processes the clinical performance data of countless patients, continuously adjusts the values of the elements of the low-rank joint compression vector, so that it can better adapt to different types of clinical performance data, thereby achieving the purpose of accurately extracting key feature information in subsequent processing. The server multiplies the low-rank joint compression vector and the state vector to obtain a gating signal. Assuming that the low-rank joint compression vector is [2, 3, 4] and the state vector is [5, 6, 7], then by the vector multiplication operation rule, multiplying each element of the low-rank joint compression vector with the corresponding element of the state vector, the gating signal obtained is [2x5, 3x6, 4x7], i.e. [10, 18, 28]. The server multiplies the clinical performance feature vector to be analyzed and the input feature vector set to obtain a state representation. For example, the clinical performance feature vector to be analyzed is [3, 4, 5], and the corresponding vector in the input feature vector set is [6, 7, 8]. Then, by the vector multiplication operation rule, multiplying each element of the clinical performance feature vector to be analyzed with the corresponding element of the input feature vector set, the state representation obtained is [3x6, 4x7, 5x8], i.e. [18, 28, 40].The server multiplies the feature vector of the clinical manifestation to be analyzed with the set of output feature vectors to obtain the transformed feature. For example, if the feature vector of the clinical manifestation to be analyzed is [4, 5, 6] and the corresponding vector in the set of output feature vectors is [7, 8, 9], then by the operation rule of vector multiplication, each element of the feature vector of the clinical manifestation to be analyzed is multiplied by the corresponding element of the set of output feature vectors, and the transformed feature obtained is [4x7, 5x8, 6x9], i.e., [28, 40, 54]. The server obtains the gating weight vector according to the gating signal and the state representation. Through a series of mathematical operations (such as vector addition, subtraction, multiplication, etc.) on the gating signal and the state representation, a gating weight vector is calculated according to its internal logical relationship. Assuming that the gating weight vector obtained through these operations is [0.2, 0.3, 0.5]. The server performs dimensionality reduction processing on the gating weight vector according to the correlation measure of the state representation. The server analyzes the correlation between each element in the state representation, such as by calculating the covariance, correlation coefficient, etc. between the elements to determine the degree of correlation. Then, according to the degree of correlation, the elements with weak correlation in the gating weight vector are removed or combined, etc. to achieve dimensionality reduction. For example, through analysis, it is found that in the gating weight vector [0.2, 0.3, 0.5], the correlation between the first element and the other two elements is weak, so the first element can be removed, and the dimensionality-reduced gating weight vector is [0.3, 0.5]. The server normalizes the dimensionality-reduced gating weight vector to obtain the gating control coefficient. Each element of the dimensionality-reduced gating weight vector is divided by the length of the vector (i.e., the square root of the sum of the squares of all elements) to achieve normalization. Assuming that the dimensionality-reduced gating weight vector is [0.3, 0.5] and the length of the vector is √(0.3. 2+0.52) = V0.34, then the normalized gating control coefficient is [0.3 / V0.34, 0.5 / V0.34]. The server multiplies the gating control coefficient with the transformed feature to obtain the gating output vector corresponding to the target execution operation module, and obtains the clinical manifestation feature encoding vector according to the plurality of gating output vectors and the gating control coefficients respectively corresponding to the plurality of execution operation modules. For example, the gating control coefficient is [0.3 / V0.34, 0.5 / V0.34], and the transformed feature is [28, 40, 54], then the gating output vector obtained after multiplication is [28x(0.3 / V0.34), 40x(0.5 / V0.34), 54x(0.5 / V0.34)]. By performing similar operations on all execution operation modules, the plurality of gating output vectors are combined and integrated according to certain rules, and finally the clinical manifestation feature encoding vector is obtained. This clinical manifestation feature encoding vector contains the concentrated feature information of the clinical manifestation data of Ms. Zhang to be analyzed after a series of complex processing, and the number of feature elements is the same as that of the low-rank joint compression vector, which provides an important basis for subsequent merging with medical record data and further data screening and other operations.
[0028] In the embodiment of the application, in the feature optimization subassembly, the low-rank joint compression vector is used to perform gating mechanism processing on the clinical manifestation feature vector to be analyzed to obtain a clinical manifestation feature encoding vector, which can be implemented by the following example.
[0029] In the feature optimization subassembly, a low-rank joint compression vector is obtained, and the clinical manifestation feature vector to be analyzed and the low-rank joint compression vector are combined to obtain a clinical manifestation feature vector set; the low-rank joint compression vector is a trainable parameter.
[0030] The low-rank joint compression vector and the clinical manifestation feature vector set are placed into a plurality of execution operation modules, and conditional gating fusion processing is performed on the low-rank joint compression vector and the clinical manifestation feature vector set in each execution operation module to obtain a plurality of gating output vectors, and a clinical manifestation feature encoding vector is obtained according to the plurality of gating output vectors and the gating control coefficients respectively corresponding to the plurality of execution operation modules.
[0031] In the embodiments of the present application, for example, the server is processing the clinical manifestation data to be analyzed of a patient named Mr. Zhao. Mr. Zhao came to the hospital for medical treatment due to illness, and showed symptoms such as abdominal pain, diarrhea, fever, and fatigue. The doctor arranged a series of examinations for him, including blood tests, stool tests, body temperature monitoring, etc. The data generated by these examinations and the detailed records of Mr. Zhao's symptoms constitute the clinical manifestation data to be analyzed to be processed this time. After the previous steps, the clinical manifestation feature vector to be analyzed has been obtained, and now the server needs to process it in the feature optimization subcomponent according to the low-rank joint compression vector to obtain the clinical manifestation feature encoding vector. In the feature optimization subcomponent, the server first needs to obtain the low-rank joint compression vector. This low-rank joint compression vector is a trainable parameter obtained by the hospital's target integrated model through continuous training and optimization in the process of processing a large amount of medical clinical data. It is like a special "key" that can help the server better filter and integrate information in the clinical manifestation feature vector to be analyzed in subsequent processing. For example, in the previous analysis of similar symptoms and examination data of numerous patients, the model gradually adjusts the values of each element of the low-rank joint compression vector according to the characteristics and laws of the data, so that it can adapt to different clinical manifestations, thereby playing a role in this time processing Mr. Zhao's data. Taking Mr. Zhao's case as an example, the clinical manifestation feature vector to be analyzed extracted from his blood test data (such as analysis of white blood cell count, red blood cell count, etc.), stool test data (such as analysis of whether there are bacteria, parasites, etc.), and body temperature monitoring data (analysis of body temperature trend, etc.) is assumed to be [3, 5, 2] (here is only a simple indication, the actual vector dimension and elements will be determined according to the specific analysis situation). The obtained low-rank joint compression vector is assumed to be [1, 2, 4]. The server merges the two vectors according to a specific merging rule (which is pre-set by the model based on a large amount of data training and is a reasonable way). For example, the corresponding elements are added, so the vector in the clinical manifestation feature vector set obtained after merging is [3+1, 5+2, 2+4], that is, [4, 7, 6]. Through this way, all the clinical manifestation feature vectors to be analyzed are merged with the low-rank joint compression vector, and finally a clinical manifestation feature vector set containing various feature information is formed. The server places the low-rank joint compression vector and the clinical manifestation feature vector set obtained just now into multiple execution operation modules. These execution operation modules are part of the target integrated model, and they each undertake a specific operation task and work together to achieve fine processing of data. For example, assume that there are three execution operation modules, marked as module A, module B, and module C.The server sends the low-rank joint compression vector [1, 2, 4] and the set of clinical manifestation feature vectors (containing multiple vectors such as [4, 7, 6] mentioned above) to the three modules respectively, so as to perform subsequent conditional gate fusion processing in each module. In each execution operation module, the server performs conditional gate fusion processing on the low-rank joint compression vector and the set of clinical manifestation feature vectors to obtain multiple gate output vectors. Taking module A as an example, in this module: first, according to a specific algorithm and rule, the low-rank joint compression vector is multiplied with the state vector in the module (which is also a vector determined by the model during the training process and is related to the operation logic of the module) to obtain a gate signal. Assuming that the state vector in module A is [5, 3, 2], then the gate signal obtained by multiplying the low-rank joint compression vector [1, 2, 4] with it is [1x5, 2x3, 4x2], that is, [5, 6, 8]. Next, the vector in the set of clinical manifestation feature vectors (such as [4, 7, 6]) is multiplied with the input feature vector set in the module (also determined by the model and related to the operation of the module) to obtain a state representation. Assuming that the corresponding vector in the input feature vector set of module A is [2, 4, 3], then the state representation obtained by multiplying [4, 7, 6] with it is [4x2, 7x4, 6x3], that is, [8, 28, 18]. Then, the vector in the set of clinical manifestation feature vectors (such as [4, 7, 6]) is multiplied with the output feature vector set in the module (also determined by the model and related to the operation of the module) to obtain transformed features. Assuming that the corresponding vector in the output feature vector set of module A is [3, 5, 4], then the transformed features obtained by multiplying [4, 7, 6] with it are [4x3, 7x5, 6x4], that is, [12, 35, 24]. Subsequently, the gate weight vector is obtained according to the gate signal [5, 6, 8] and the state representation [8, 28, 18]. This may involve a series of complex mathematical operations, such as combination operations through addition, subtraction, multiplication, etc. of vectors, to calculate a gate weight vector according to its internal logical relationship. Assuming that the gate weight vector obtained through these operations is [0.2, 0.3, 0.5]. Then, the gate weight vector is dimensionally reduced according to the correlation measure of the state representation. The server analyzes the correlation between elements in the state representation, such as by calculating the covariance, correlation coefficient, etc. of the elements to determine the degree of correlation. Then, according to the degree of correlation, the elements with weak correlation in the gate weight vector are removed or combined, etc. to achieve dimension reduction. Assuming that through analysis it is found that the first element in the gate weight vector [0.2, 0.3, 0.5] has weak correlation with the other two elements, then the first element can be removed to obtain the dimensionally reduced gate weight vector [0.3, 0.5]. Finally, the dimensionally reduced gate weight vector is normalized to obtain a gate control coefficient.Each element of the dimension-reduced gate weight vector is divided by its vector length (i.e., the square root of the sum of squares of all elements) to achieve normalization. Assuming that the dimension-reduced gate weight vector is [0.3, 0.5] and its vector length is √(0.3. 2 +0.52) = √0.34, then the normalized gate control coefficient is [0.3 / √0.34, 0.5 / √0.34]. And the gate control coefficient is multiplied with the transformed feature to obtain the gate output vector corresponding to module A. Assuming that the transformed feature is [12, 35, 24] and the gate control coefficient is [0.3 / √0.34, 0.5 / √0.34], then the gate output vector obtained after multiplication is [12x(0.3 / √0.34), 35x(0.5 / √0.34), 24x(0.5 / √0.34)]. The same process is also performed in module B and module C and other operation modules to obtain the gate output vectors corresponding to each module. The server obtains the clinical performance feature encoding vector according to the multiple gate output vectors and the gate control coefficients corresponding to the multiple operation modules. Assuming that the gate output vector obtained by module B is [15x(0.4 / √0.41), 28x(0.6 / √0.41), 30x(0.6 / √0.41)] (here only for illustration, the actual value is obtained according to specific operation), and the gate output vector obtained by module C is [18x(0.5 / √0.50), 32x(0.7 / √0.50), 35x(0.7 / √0.50)]. The gate control coefficient corresponding to module A is [0.3 / √0.34, 0.5 / √0.34], the gate control coefficient corresponding to module B is [0.4 / √0.41, 6 / √0.41] (here the data is assumed, the actual value is obtained according to operation), and the gate control coefficient corresponding to module C is [0.5 / √0.50, 0.7 / √0.50]. The server integrates these gate output vectors according to the respective gate control coefficients according to specific integration rules (also a reasonable way based on a large amount of data training). For example, the gate output vectors are multiplied by the respective gate control coefficients and then added to obtain the clinical performance feature encoding vector. This clinical performance feature encoding vector contains the condensed feature information of Mr. Zhao's to-be-analyzed clinical performance data after a series of complex processing, and the number of feature elements is the same as that of the low-rank joint compression vector, which provides an important basis for subsequent merging with medical record data and further data screening and other operations.
[0032] In the embodiment of the present application, the plurality of operation modules includes a target operation module; and the conditional gate fusion processing of the low-rank joint compression vector and the clinical performance feature vector set in each operation module to obtain a plurality of gate output vectors can be implemented through the following examples.
[0033] In the target execution operation module, the low-rank joint compression vector is multiplied with the state vector to obtain a gating signal, the set of clinical manifestation feature vectors is multiplied with the set of input feature vectors to obtain a state representation, and the set of clinical manifestation feature vectors is multiplied with the set of output feature vectors to obtain transformed features.
[0034] A gating weight vector is obtained according to the gating signal and the state representation, the gating weight vector is dimensionally reduced according to a correlation measure of the state representation, the dimensionally reduced gating weight vector is normalized to obtain a gating regulation coefficient, and the gating regulation coefficient is multiplied with the transformed features to obtain a gating output vector corresponding to the target execution operation module.
[0035] In the embodiments of the present application, an example is provided. A patient named Mr. Chen comes to the hospital for treatment due to persistent abdominal pain, abdominal distension, loss of appetite, and occasional nausea and vomiting. The doctor arranges a series of detailed examinations for him, including abdominal ultrasound, blood biochemical examination, gastroscopy, etc. The data generated by these examinations and the doctor's detailed records of Mr. Chen's symptoms together constitute the clinical manifestation data to be analyzed. After a series of data processing steps, the stage of processing the relevant vectors in the target execution operation module to obtain the gating output vector is reached, which is crucial for subsequent accurate analysis of Mr. Chen's condition and data screening. In the target execution operation module, the server first needs to perform specific operation processing on the low-rank joint compression vector and the state vector to obtain the gating signal. Taking Mr. Chen's case as an example, the low-rank joint compression vector is a special vector trained from a large amount of clinical data accumulated by the hospital over a long period of time, which contains a lot of feature information related to various disease manifestations after integration and compression. For example, this low-rank joint compression vector may be like a "feature information aggregator" that has refined the data of many similar patients with similar symptoms. The state vector is a vector set by the target execution operation module according to the logic and requirements of the entire data processing, which interacts with the low-rank joint compression vector to help filter more targeted information. When the server processes the low-rank joint compression vector and the state vector together, it is like combining two "tools" with different functions, and through a specific operation method between them (here, the specific formula is not involved, it is a processing logic based on model setting), a gating signal that plays a key role in subsequent processing is finally obtained. This gating signal is like a "sign", which will tell the subsequent processing flow which aspects of information should be focused on. Next, the server needs to perform corresponding operation processing on the clinical manifestation feature vector set and the input feature vector set to obtain the state representation. The clinical manifestation feature vector set of Mr. Chen is composed of multiple feature vectors extracted from the in-depth analysis of his examination data and symptom description. These feature vectors cover various aspects of information from the organ morphology features seen in abdominal ultrasound, abnormal conditions of various indicators in blood biochemical examination, to gastric lesion features found in gastroscopy. The input feature vector set is another set of vectors set by the target execution operation module according to the processing flow and requirements, which interacts with the clinical manifestation feature vector set to reflect a comprehensive state of the current data processing stage. After the server performs operation processing on the clinical manifestation feature vector set and the input feature vector set, a state representation is obtained.This state representation is like a "portrait" that combines various features related to Mr. Chen's current condition in a specific way, providing a comprehensive reference for subsequent processing, allowing subsequent operations to further analyze and adjust the direction of processing based on this "portrait". Then, the server operates the clinical manifestation feature vector set and the output feature vector set to obtain the transformed features. The clinical manifestation feature vector set is the same as the previously mentioned vector set containing Mr. Chen's various condition feature information. The output feature vector set is also a set of vectors set by the target execution operation module, which interacts with the clinical manifestation feature vector set to obtain the transformed feature information. After the server operates these two sets of vectors, it obtains the transformed features. This transformed feature is like a "repackaging" of Mr. Chen's condition feature information, which presents the original condition features in a new way that is more suitable for subsequent processing, providing important basic data for subsequent operations such as generating the gating weight vector. After obtaining the gating signal and state representation, the server generates the gating weight vector based on these two quantities. The gating signal and state representation are like two "keys" that contain information that is combined through a server-internal logic and algorithm (not involving specific formulas here) to generate a gating weight vector after a series of complex processing. This gating weight vector acts as a "regulator" that determines the degree and direction of subsequent processing of various feature information based on the gating signal and state representation obtained earlier. The server then performs dimensionality reduction on the gating weight vector based on the relevance measure of the state representation. The information contained in the state representation can reflect the relationships between Mr. Chen's condition-related features. The server analyzes these relationships (not involving specific analysis methods, based on a model-internal judgment logic of relevance) to determine which elements in the gating weight vector are relatively unimportant or have weak relevance to other elements. Then, based on this relevance judgment, the server performs dimensionality reduction on the gating weight vector, i.e., removes those relatively unimportant or weakly relevant elements, making the gating weight vector more concise and accurately reflecting the truly important feature information, just like making a "big and complete" regulator more "precise and effective". After dimensionality reduction of the gating weight vector, the server normalizes the dimensionality-reduced gating weight vector to obtain the gating control coefficient. The purpose of normalization is to allow the gating control coefficient to be within a suitable range, so that when it is multiplied by the transformed features in subsequent processing, it can more reasonably adjust the weight of the feature information. The server adjusts the elements in the dimensionality-reduced gating weight vector to meet the normalization requirements, thereby obtaining the gating control coefficient.The gating control coefficient is like a "standardized regulator" that can more accurately control and adjust the transformed features. Finally, the server operates the gating control coefficient with the transformed features to obtain the gating output vector corresponding to the target execution operation module. As a "standardized regulator", the gating control coefficient and the transformed features (i.e., the previously re-packaged Mr. Chen's disease feature information) are operated to obtain the gating output vector. The gating output vector is like the "final presentation" of Mr. Chen's disease feature information after a series of fine processing, which integrates the results of all previous processing steps. It is integrated with other gating output vectors obtained by other execution operation modules, and finally obtains the clinical performance feature encoding vector, thereby providing important basic data for subsequent data screening and other operations in medical clinical big data analysis.
[0036] In the embodiment of the present application, the medical record data is obtained, the medical record data and the clinical performance feature encoding vector are merged to obtain a multi-dimensional merged vector, the multi-dimensional merged vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements, which can be implemented by the following examples.
[0037] The clinical performance feature encoding vector is loaded into a target integrated model; the target integrated model includes a multi-dimensional feature extraction component and a feature integration processing component, and the multi-dimensional feature extraction component includes a feature extraction component;
[0038] In the feature extraction component, medical record data is obtained, feature extraction is performed on the medical record data to obtain a description text feature vector, and the description text feature vector and the clinical performance feature encoding vector are merged to obtain a multi-dimensional merged vector;
[0039] In the feature integration processing component, the multi-dimensional merged vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements.
[0040] In the embodiments of the present application, it is assumed that the server is processing the data of a patient named Mr. Wang. Mr. Wang comes to the hospital for medical treatment due to physical discomfort, and shows symptoms such as chest pain, shortness of breath, and fatigue. The doctor conducts a comprehensive examination on him, including electrocardiogram examination, chest X-ray examination, blood test, etc. After a series of processing, the characteristic code vector of Mr. Wang's clinical manifestations has been obtained, and the subsequent processing steps related to the medical record data are now to be completed to obtain the final analysis feature vector for analysis. After the server completes the preliminary feature extraction and other operations on the data of Mr. Wang's clinical manifestations, obtains the clinical manifestation characteristic code vector, and loads this vector into the target integrated model. This target integrated model is a comprehensive model specially used for processing medical clinical big data analysis, which has been developed and optimized by the hospital for a long time. The internal structure design of the model is very fine, which contains multi-dimensional feature extraction components and feature integration processing components. The multi-dimensional feature extraction components have feature extraction components, and these different components bear specific functions and work together to accurately extract and process useful feature information from various data. Just like a complex and precise machine, each part has its unique role, and the model is ready to further process Mr. Wang's medical record data and the existing clinical manifestation characteristic code vector through the cooperation of various components to mine more valuable information for subsequent disease analysis. In the feature extraction component, the first thing the server needs to do is to obtain Mr. Wang's medical record data from the hospital's medical record database. Mr. Wang's medical record data is a detailed document recording the whole process of his medical treatment, which covers his basic information (such as age, gender, medical history, etc.), the specific situation of this illness (such as the time, frequency, and severity of symptoms, etc.), the doctor's diagnosis and treatment process (such as which examinations were performed, the results of the examinations, what treatment measures were taken, etc.), and the preliminary diagnosis conclusion, etc. For example, the medical record records that Mr. Wang is 70 years old and has a history of hypertension. The chest pain symptom started to appear a week ago, and it was occasional at first, but later became more frequent and the pain intensified. The doctor performed an electrocardiogram examination on him and found signs of myocardial ischemia, and then performed a chest X-ray examination and found no obvious lung abnormalities, etc. After obtaining Mr. Wang's medical record data, the server begins to perform feature extraction operations, aiming to convert the text form medical record data into a description text feature vector that can be combined with the clinical manifestation characteristic code vector for processing. For the basic information part of the medical record, the server will analyze the influence of Mr. Wang's age, gender, and other factors on the disease, such as the fact that older people are more likely to develop certain chronic disease-related complications, and men and women have different incidence rates of certain diseases, etc. These factors are converted into corresponding feature values to form part of the description text feature vector. For the description of the onset situation, the server will focus on the time sequence and trend of symptom occurrence.For example, the change from occasional to frequent chest pain in Mr. Wang, the server extracts feature values such as the rate of change in the frequency of attacks, the duration of symptoms, and adds them to the text feature vector. For the diagnosis and treatment process and diagnosis conclusion part, the server will extract feature values such as the relevance of the examination items, the effectiveness of the treatment measures, and the certainty of the diagnosis results according to the different examination items, treatment measures, and diagnosis results, and further improve the description of the text feature vector. Through detailed analysis and feature extraction of each part of the medical record data, the server finally obtains a description of the text feature vector, which contains the quantitative representation of many key information in Mr. Wang's medical record data, and prepares for the subsequent merging of the clinical performance feature coding vector. After obtaining the description of the text feature vector, the server will perform feature merging operation on it and the previously loaded clinical performance feature coding vector to obtain a multi-dimensional merged vector. For example, the clinical performance feature coding vector may contain feature information about his body function, lesion characteristics, etc. extracted from his examination data (such as electrocardiogram, chest X-ray, blood test, etc.) in a specific coding form. The description of the text feature vector is the feature information about his illness background, diagnosis and treatment process, etc. extracted from the medical record text data. When the two vectors are merged, it is like integrating the information obtained from different angles of Mr. Wang's illness, forming a more comprehensive and richer multi-dimensional merged vector. For example, the information of the two vectors can be fused by adding corresponding elements or following a certain preset merging rule, so that the multi-dimensional merged vector contains both the body condition features reflected by the examination data and the illness development process features described by the medical record text, providing more abundant materials for subsequent processing in the feature integration processing component. After obtaining the multi-dimensional merged vector, the server will send it to the feature integration processing component for integration processing, with the purpose of obtaining a to-be-analyzed medical record feature vector with a preset integration number of feature elements. This preset integration number is equal to the number of feature elements in the archived medical record feature vector corresponding to the archived clinical performance data in the hospital's preset medical record archive. This is done to facilitate more accurate selection of similar cases when comparing with the archived medical record feature vector in subsequent operations. In the feature integration processing component, several different methods may be used for integration processing, and the following are some common examples: (1) Based on feature compression and reconstruction: Assuming that the feature integration processing component uses the method of compression and reconstruction to process the multi-dimensional merged vector. First, the server will perform feature compression processing on the multi-dimensional merged vector, which is like compressing a "data packet" with a large amount of information to make it occupy less space, but trying to preserve the key information. Through a specific compression algorithm (such as principal component analysis), some secondary information in the multi-dimensional merged vector is removed or combined to obtain a compressed feature vector.Then, according to a certain sequence processing step, the compressed feature vector is processed for dimension reduction. For example, according to the way of selecting one every few elements, the dimension of the vector is gradually reduced, so that it is more concise. Finally, in the feature reconstruction component (this component is obtained by training according to the preset integration number, and the number of feature elements output by it is fixed as the preset integration number), according to the feature correlation of the dimension-reduced feature vector in each sequence processing step, a feature element number of the preset integration number of the feature vector to be analyzed is reconstructed by complex calculation and adjustment. This process is like decompressing the "data packet" after compression and dimension reduction and assembling it according to new requirements, so that it meets the preset feature element number requirement for subsequent use. Based on clinical term analysis and feature weight adjustment method: another possible processing method is to realize integration processing through clinical term analysis and feature weight adjustment. In the feature integration processing component, the server first analyzes each clinical term item in the clinical term library according to the multi-dimensional merged vector to determine their clinical correlation scores with Mr. Wang's illness. For example, for the clinical term "chest pain", combined with the information about Mr. Wang's chest pain symptoms, frequency of onset, pain degree, etc. in the multi-dimensional merged vector, the term "chest pain" is given a corresponding clinical correlation score. Then, according to these clinical correlation scores, the server can obtain the patient performance information of Mr. Wang's clinical performance data to be analyzed, and then according to the preset integration number, the patient performance information is segmented to obtain multiple clinical performance description segments. Then, the server will obtain the respective appearance frequency of each clinical performance description segment in the patient performance information, and the respective clinical term rarity indicators of them in the clinical term library. For example, the description segment "chest pain" has a high frequency of appearance in Mr. Wang's patient performance information, and "chest pain" is not particularly rare in the clinical term library. According to these appearance frequencies and clinical term rarity indicators, the server can calculate multiple clinical performance influence factors. Finally, in the feature weight adjustment component, these clinical performance influence factors are multiplied with the multi-dimensional merged vector respectively to obtain multiple clinical performance weight vectors, and then the multiple clinical performance weight vectors are processed comprehensively (such as weighted summation, etc.), and finally a feature element number of the preset integration number of the feature vector to be analyzed is obtained.
[0041] In the embodiment of the application, the integration processing of the multi-dimensional merged vector in the feature integration processing component to obtain a feature element number of the preset integration number of the feature vector to be analyzed can be implemented by the following examples.
[0042] loading the multi-dimensional merged vector to the feature integration processing component, the feature integration processing component comprising a feature compression component and a feature reconstruction component; the feature reconstruction component is obtained by training according to a preset integration number, and the number of feature elements output by the feature reconstruction component is fixed as the preset integration number;
[0043] In the feature compression component, the feature compression processing is performed on the multi-dimensional merged vector to obtain a compressed feature vector, and dimension reduction processing is performed on the compressed feature vector according to a sequence processing step length to obtain a dimension-reduced feature vector.
[0044] In the feature reconstruction component, according to the feature correlation of the dimension-reduced feature vector in each sequence processing step length, a to-be-analyzed medical record feature vector with a feature element number of the preset integration number is obtained.
[0045] In the embodiment of the application, for example, Ms. Li comes to the hospital for medical treatment due to multiple discomforts in her body, and she has symptoms such as joint pain, fatigue, blurred vision, and the like. The doctor arranges a series of detailed examinations for her, including blood tests, joint X-ray examinations, eye examinations, and the like, and records her medical record information in detail. After a series of previous data processing steps, the multidimensional merged vector about Ms. Li has been obtained, and now the server needs to further process the multidimensional merged vector in the feature integration processing component to obtain a to-be-analyzed medical record feature vector that meets specific requirements, thereby providing more accurate data support for subsequent disease analysis and data screening operations and the like. After completing the previous related steps, the server obtains the multidimensional merged vector about Ms. Li. This multidimensional merged vector is like an “information package” containing multiple aspects of information about Ms. Li's illness, which integrates various feature information extracted from the clinical manifestation feature code vector of Ms. Li and the medical record data. Next, the server will load the multidimensional merged vector into the feature integration processing component. This feature integration processing component is an important module specially designed and trained by the hospital for processing such data, which internally includes a feature compression component and a feature reconstruction component. These two components work together to reasonably process the multidimensional merged vector, so that it is converted into a to-be-analyzed medical record feature vector that meets the subsequent analysis requirements. In the feature compression component, the server first needs to perform feature compression processing on the loaded multidimensional merged vector. This step is like arranging a large box (multidimensional merged vector) full of various items (representing different feature information) and cleaning out some less important or repetitive items to make the box smaller and more compact (obtaining a compressed feature vector), while trying to retain those most critical items (feature information) for subsequent analysis. Taking Ms. Li's case as an example, the multidimensional merged vector may include various blood index-related feature information obtained from blood tests, joint condition feature information from joint X-ray examinations, vision-related feature information from eye examinations, and multiple aspects of information such as her medical history and symptom onset frequency recorded in the medical record. In the feature compression processing, the server will filter and integrate these feature information according to their importance and correlation between them and other factors through a specific compression algorithm (such as an algorithm based on principal component analysis). For example, for some indexes in blood tests, if it is found that there is a strong correlation between some indexes, and one of the indexes can represent the trend of change of other related indexes to a large extent, then the server may retain this most representative index, and appropriately compress or combine the other indexes with strong correlation. Similarly, the feature information of joint X-ray examinations and eye examinations and the like will also be processed according to similar principles. After such feature compression processing, the server obtains a compressed feature vector.This compressed feature vector, although it has reduced the amount of information compared to the original multi-dimensional merged vector, retains the most critical and reflective information of Ms. Li's condition, preparing for subsequent dimension reduction processing. After obtaining the compressed feature vector, the server will next perform dimension reduction processing on this compressed feature vector according to the sequence processing step, to obtain a reduced dimension feature vector. The sequence processing step here is a parameter determined by analyzing factors such as the distribution of feature elements in the multi-dimensional merged vector, which determines how to gradually reduce the dimension of the vector during dimension reduction. For example, for Ms. Li's compressed feature vector, assume the sequence processing step is set to every 2 elements, select 1 element to retain (this is just an example, the actual sequence processing step will be determined according to the specific situation). Then the server will follow this rule, starting from the first element of the compressed feature vector, every 2 elements select 1 element, discard the other elements, gradually reduce the dimension of the vector. In this process, the server will continuously scan and process the compressed feature vector according to the sequence processing step, like along a regular route (determined by the sequence processing step) in the "information channel" of the compressed feature vector, filtering out elements that do not meet the step requirements, and finally obtaining a reduced dimension feature vector. The dimension of this reduced dimension feature vector is lower than that of the compressed feature vector, and it is more concise, further refining the key feature information of Ms. Li's condition, making it more convenient for subsequent processing in the feature reconstruction component. In the feature reconstruction component, the server will obtain a feature element number of the pre-set integration number of the feature vector to be analyzed according to the feature correlation of the reduced dimension feature vector in each sequence processing step. This feature reconstruction component is a module specially used for reconstructing feature vectors obtained by the hospital after a large amount of training according to the pre-set integration number, and the number of feature elements output by it is fixed as the pre-set integration number, which ensures that the final obtained feature vector to be analyzed can be consistent in dimension with the archived feature vector of the archived clinical manifestation data in the hospital's pre-set medical record archive, facilitating subsequent data comparison and analysis operations. Taking Ms. Li's case as an example, in the feature reconstruction component, the server will first analyze the feature correlation of the reduced dimension feature vector in each sequence processing step. For example, the reduced dimension feature vector may be a lower dimension vector obtained after the previous dimension reduction processing, and the elements in it may represent different aspects of Ms. Li's condition, such as key changes in blood indicators, key features of joint conditions, key factors of vision problems, etc. The server will analyze the correlation of these elements with the corresponding elements in the original multi-dimensional merged vector under different sequence processing steps, as well as the relationship between these elements.For example, at a certain sequence processing step, one element in the reduced dimension feature vector represents a certain key change in blood indicators. The server will analyze the correlation of this element with the elements in the blood test related part of the original multi-dimensional merged vector, as well as the correlation of this element with other elements in the reduced dimension feature vector (such as elements related to joint conditions or vision problems). Based on these correlation analyses, the server will reconstruct a feature vector for the medical record to be analyzed with a preset number of integrated elements through a series of complex calculations and adjustments. This process is like recombining and adjusting various key feature information after dimension reduction and correlation analysis according to new rules and requirements, so that it finally forms a vector that meets the preset number of integrated requirements. This vector is the feature vector for the medical record to be analyzed. For example, assuming that the preset number of integrated elements is 5, after a series of processing in the feature reconstruction component, the server finally obtains a feature vector for the medical record to be analyzed, which may contain 5 carefully selected and adjusted feature elements that accurately reflect key information about Ms. Li's condition, such as key changes in blood indicators, key features of joint conditions, key factors of vision problems, and information about medical history and symptom frequency, thereby providing more accurate data support for subsequent condition analysis and data screening operations.
[0046] In the embodiment of the present application, the integration processing of the multi-dimensional merged vector in the feature integration processing component to obtain a feature vector for the medical record to be analyzed with a preset number of integrated elements can be implemented by the following examples.
[0047] Load the multi-dimensional merged vector into the feature integration processing component, which includes a clinical term analysis component and a feature weight adjustment component;
[0048] In the clinical term analysis component, obtain the clinical relevance scores of each clinical term in the clinical term library according to the multi-dimensional merged vector, obtain the patient performance information of the clinical performance data to be analyzed according to the clinical relevance scores, and perform word segmentation on the patient performance information according to the preset number of integrated elements to obtain a plurality of clinical performance description segments; the preset number of integrated elements is the number of clinical performance description segments corresponding to the plurality of clinical performance description segments;
[0049] Obtain the occurrence frequencies of the plurality of clinical performance description segments in the patient performance information and the clinical term rarity indicators of the plurality of clinical performance description segments in the patient performance information, and obtain a plurality of clinical performance influence factors according to the plurality of occurrence frequencies and the plurality of clinical term rarity indicators;
[0050] In the feature weight adjusting component, the plurality of clinical manifestation influence factors are multiplied with the multi-dimensional merging vector respectively to obtain a plurality of clinical manifestation weight vectors, and a feature vector to be analyzed is obtained according to the plurality of clinical manifestation weight vectors, wherein a number of feature elements of the feature vector to be analyzed is a preset integration number.
[0051] In the embodiment of the present application, it is assumed that the information management environment is in a large general hospital, the server undertakes the important task of deeply analyzing and processing the mass patient medical record data and the related clinical manifestation data, so as to provide strong support for the accurate diagnosis and effective treatment of doctors.
[0052] At this moment, the server is processing the data of a patient named Mr. Zhang. Mr. Zhang came to the hospital due to his physical discomfort, presenting symptoms such as chest pain, shortness of breath, cough, fatigue, and lower extremity edema. The doctor arranged a comprehensive examination for him, including electrocardiogram, chest X-ray, cardiac ultrasound, blood test, and lower extremity vascular ultrasound. The data generated from these examinations and the detailed records of Mr. Zhang's symptoms have been processed through a series of steps, forming a multi-dimensional merged vector for Mr. Zhang. Now, the server needs to process this multi-dimensional merged vector in the feature integration processing component according to a specific method to obtain the feature vector of the medical record that meets the requirements of subsequent analysis. The server first loads the multi-dimensional merged vector of Mr. Zhang into the feature integration processing component. This feature integration processing component is carefully designed and optimized by the hospital through long-term training, and is specifically used for targeted processing of such multi-dimensional merged vectors. It contains a clinical term analysis component and a feature weight adjustment component, which work together to complete the integration processing task of the multi-dimensional merged vector. In the clinical term analysis component, the server will determine the clinical relevance scores of each clinical term item in the clinical term library according to the loaded multi-dimensional merged vector. The clinical term library is a large vocabulary set accumulated and continuously improved by the hospital over a long period of time, which contains various terms related to medical conditions, symptoms, examinations, treatments, etc., such as "chest pain", "shortness of breath", "myocardial infarction", "heart failure", "electrocardiogram abnormalities", etc. Taking Mr. Zhang's case as an example, the multi-dimensional merged vector contains information extracted from the examination results and symptom descriptions. The server will compare and analyze these information with the terms in the clinical term library. For example, for the clinical term "chest pain", the server will check the specific description of Mr. Zhang's chest pain symptoms in the multi-dimensional merged vector, including the location of chest pain (left chest, right chest, or behind the sternum, etc.), the nature of chest pain (stabbing pain, dull pain, or squeezing pain, etc.), the frequency of chest pain (occasional, frequent, or continuous, etc.), and the association of chest pain with other symptoms (such as shortness of breath, cough, etc.). Based on these detailed information, the server will determine a clinical relevance score for the "chest pain" term according to a pre-set evaluation rule (this rule is based on a large amount of clinical data and expert experience). Assuming that after evaluation, the clinical relevance score of "chest pain" for Mr. Zhang's condition is determined to be 0.8 (the score range here can be between 0 and 1, and the higher the value, the stronger the relevance). Similarly, for other clinical term items such as "shortness of breath", "cough", "electrocardiogram abnormalities", etc., the server will also determine their respective clinical relevance scores according to the relevant information in the multi-dimensional merged vector.For example, the clinical relevance score of "shortness of breath" can be determined as 0.7, the clinical relevance score of "cough" can be 0.6, the clinical relevance score of "ECG abnormality" can be 0.9, and so on. After determining the clinical relevance scores of various clinical term items, the server will then obtain the patient performance information of the patient performance data to be analyzed according to the clinical relevance scores, and perform word segmentation on the patient performance information according to the preset integration number to obtain a plurality of clinical performance description segments. Still taking Mr. Zhang as an example, the server will comprehensively consider the clinical relevance scores of various clinical term items and the related information in the multi-dimensional merging vector to outline the complete patient performance information of Mr. Zhang. For example, according to the high relevance score of "chest pain" and the detailed description thereof in the multi-dimensional merging vector, the server will determine the specific situation of Mr. Zhang's chest pain; combined with the related information of symptoms such as "shortness of breath" and "cough", the server will depict Mr. Zhang's overall respiratory discomfort symptoms; at the same time, considering the relevance scores of clinical terms related to examination results such as "ECG abnormality", the server will also integrate the examination results into the patient performance information. Assuming that the preset integration number is 5, the server will perform word segmentation on the patient performance information of Mr. Zhang according to certain logic and semantic rules (these rules are also trained based on a large amount of clinical data) to obtain a plurality of clinical performance description segments. For example, the following five clinical performance description segments can be obtained: 1. "chest pain with a posterior sternal site, a dull pain nature, and frequent onset"; 2. "shortness of breath, aggravated after exercise"; 3. "cough with a small amount of white sputum"; 4. "fatigue, obvious four-limb weakness"; and 5. "lower extremity edema, bilateral, and mild degree". After obtaining the plurality of clinical performance description segments, the server will further obtain the respective appearance frequencies of these description segments in the patient performance information and the respective clinical term rarity indexes of the clinical terms corresponding to these description segments in the clinical term library, and then calculate a plurality of clinical performance influence factors according to the plurality of appearance frequencies and the plurality of clinical term rarity indexes. For each clinical performance description segment, the server will count the number of times of appearance of the segment in the patient performance information of Mr. Zhang to determine the appearance frequency thereof. For example, the description segment "chest pain with a posterior sternal site, a dull pain nature, and frequent onset" appears more frequently in the entire patient performance information of Mr. Zhang, and the appearance frequency thereof is assumed to be 0.4 (the appearance frequency is also within the range of 0 to 1). At the same time, the server will query the clinical term library to determine the rarity index of the clinical term in each clinical performance description segment in the entire clinical term library. For example, the clinical term "posterior sternal dull pain" is not particularly common in the clinical term library, and the clinical term rarity index thereof is assumed to be 0.3; while the clinical term "cough" in the description segment "cough with a small amount of white sputum" is relatively common, and the clinical term rarity index thereof can be 0.1, and so on.Then, the server will calculate the clinical manifestation impact factor of each description segment according to the frequency of occurrence of each clinical manifestation description segment and the clinical term rarity index, through a pre-set calculation formula (this formula is also based on a large amount of clinical data and expert experience). For example, for the description segment "chest pain and the site is behind the sternum, the nature is dull pain, and the frequency of attack is high", its clinical manifestation impact factor may be obtained by multiplying the frequency of occurrence 0.4 by the clinical term rarity index 0.3, i.e. 0.12. Similarly, for other clinical manifestation description segments, their respective clinical manifestation impact factors will also be calculated in this way. In the feature weight adjustment component, the server will multiply the previously calculated multiple clinical manifestation impact factors with the multi-dimensional merging vector respectively to obtain multiple clinical manifestation weight vectors. Taking Mr. Zhang's multi-dimensional merging vector as an example, assuming that the multi-dimensional merging vector is a vector containing multiple elements, these elements represent the information extracted from Mr. Zhang's various examinations, symptom descriptions, etc. When multiplying each clinical manifestation impact factor with the multi-dimensional merging vector, the server will multiply the clinical manifestation impact factor with each element of the multi-dimensional merging vector respectively according to the rules of vector multiplication. For example, for the clinical manifestation description segment "chest pain and the site is behind the sternum, the nature is dull pain, and the frequency of attack is high", its clinical manifestation impact factor is 0.12. Assuming that the first few elements of the multi-dimensional merging vector are [1, 2, 3] (here is only a simple illustration, the actual vector elements will be determined according to the specific circumstances), then the first few elements of the clinical manifestation weight vector obtained after multiplication may be [0.12x1, 0.12x2, 0.12x3], i.e. [0.12, 0.24, 0.36]. Similarly, for other clinical manifestation description segments, their respective clinical manifestation impact factors will also be multiplied with the multi-dimensional merging vector in this way to obtain multiple clinical manifestation weight vectors. After obtaining multiple clinical manifestation weight vectors, the server will obtain a feature vector of the to-be-analyzed medical record with a preset number of integrated elements according to these clinical manifestation weight vectors. The server will comprehensively process multiple clinical manifestation weight vectors, and the comprehensive processing method can be weighted summation, averaging, etc. (these methods are also based on a large amount of clinical data and expert experience). Assuming that the preset number of integration is 5, the server will process multiple clinical manifestation weight vectors according to a certain comprehensive processing method, so that the final to-be-analyzed medical record feature vector obtained has 5 elements. For example, through the method of weighted summation, each clinical manifestation weight vector is added according to a certain weight distribution (this weight distribution is also based on a large amount of clinical data and expert experience). Assuming that after weighted summation processing, the to-be-analyzed medical record feature vector obtained is [0.5, 0.4, 0.3, 0.2, 0.1] (here is only a simple illustration, the actual vector elements will be determined according to the specific circumstances).The feature vector of the case to be analyzed is obtained after a series of processing in the feature integration processing component, the number of elements thereof meets the requirement of the preset integration number, and it integrates the influence of the clinical performance description segments of Mr. Zhang and the related information in the multi-dimensional combined vector, thereby providing a suitable data basis for subsequent comparison and analysis with the feature vector of the archived case, and thus helping to more accurately screen out cases similar to Mr. Zhang in medical clinical big data analysis, and providing more targeted reference for the diagnosis and treatment of doctors.
[0053] In the embodiment of the application, the multi-dimensional combined vector includes a plurality of transition vectors, the number of feature elements of the plurality of transition vectors is the preset integration number, and the last transition vector in the plurality of transition vectors includes global expression content obtained by performing feature integration on the case data and the clinical performance feature code vector; the integration processing of the multi-dimensional combined vector in the feature integration processing component to obtain the feature vector of the case to be analyzed with the number of feature elements being the preset integration number can be implemented through the following examples.
[0054] In the feature integration processing component, the last transition vector in the plurality of transition vectors is determined as the feature vector of the case to be analyzed with the number of feature elements being the preset integration number.
[0055] In the embodiment of the present application, the server is processing the data of a patient named Ms. Liu. Ms. Liu comes to the hospital because of her illness, and she has symptoms such as headache, dizziness, blurred vision, nausea, and limb weakness. The doctor arranges a series of comprehensive examinations for her, including brain CT scan, eye examination, blood test, electrocardiogram examination, etc. The data generated by these examinations and the detailed records of Ms. Liu's symptoms and other information are processed through a series of processing steps to form a multi-dimensional merged vector about Ms. Liu. This multi-dimensional merged vector has a specific structure and contains multiple transition vectors. Now the server needs to process this multi-dimensional merged vector in the feature integration processing component according to its characteristics to obtain the final analysis feature vector for analysis. For Ms. Liu's multi-dimensional merged vector, it is like an "information aggregate" that integrates multiple aspects of information. This multi-dimensional merged vector contains multiple transition vectors, each of which carries relevant information obtained after a certain stage of processing, and the number of feature elements of these transition vectors is the preset integration number. For example, during the processing of Ms. Liu's data, different levels or angles of analysis and integration may be performed, and a transition vector may be generated after each step of processing. These transition vectors gradually accumulate and improve the expression of information related to Ms. Liu's illness. The last transition vector is particularly special, as it contains the global expression content obtained by integrating the medical record data and the clinical manifestation feature coding vector. This means that the last transition vector integrates all the information about Ms. Liu's basic information, medical history, diagnosis and treatment process recorded in the medical record, and the clinical manifestation feature information extracted and coded from various examinations and symptom manifestations, forming a relatively complete and comprehensive expression of Ms. Liu's illness. For example, the medical record data records Ms. Liu's age, medical history (such as having had mild hypertension), the time of this illness, and the sequence of symptom appearance, etc. The clinical manifestation feature coding vector contains information such as brain structure features from brain CT scan, changes in various indicators of the eyes from eye examination, abnormal conditions of various blood indicators from blood test, and related features of cardiac electrical activity from electrocardiogram examination. The last transition vector fuses these two aspects of information in a specific way, allowing it to fully reflect Ms. Liu's illness condition. In the feature integration processing component, the server needs to integrate the multi-dimensional merged vector to obtain the analysis feature vector with the preset integration number of feature elements. In this case, the processing method is relatively direct, which is to determine the last transition vector in the multiple transition vectors as the analysis feature vector with the preset integration number of feature elements. After receiving the multi-dimensional merged vector containing multiple transition vectors, the server identifies the last transition vector.Since the previous processing flow has ensured that the number of feature elements of each transition vector is the preset integration number, the last transition vector itself already has a suitable dimension and information content, and can be directly used as a feature vector of the case to be analyzed for subsequent analysis operations. Taking Ms. Liu as an example, assuming that the preset integration number is 10, then each transition vector, including the last transition vector, has 10 feature elements. When the server determines to use the last transition vector as the feature vector of the case to be analyzed, it directly extracts the last transition vector. For example, the 10 feature elements of the last transition vector can represent different aspects of information: the first element can represent a certain degree of correlation between the age of Ms. Liu and the severity of the disease (a quantitative relationship obtained by analyzing a large number of similar cases); the second element can reflect the potential influence of Ms. Liu's previous history of hypertension on the current symptoms (a quantitative representation based on data analysis and expert experience); the third element can represent the synergistic change relationship between the headache and dizziness symptoms in this episode (obtained from the analysis of various examination data and symptom records); the fourth element can be the correlation coefficient between a key indicator in the eye examination and the overall condition; the fifth element can be the correlation between the abnormal degree of an important blood indicator in the blood test and other symptoms; the sixth element can be the correspondence between a certain feature in the electrocardiogram and the overall condition of Ms. Liu's body; the seventh element can involve the influence of the time sequence of Ms. Liu's symptoms on the development of the disease; the eighth element can represent the effect of Ms. Liu's lifestyle (such as smoking, drinking, etc.) on the disease (obtained by comprehensive analysis of medical record data and other related information); the ninth element can be the influence of a certain synergistic effect found between different examination items on the disease judgment; and the tenth element can be the degree of correspondence between the doctor's preliminary diagnosis and the actual condition (based on further observation and analysis of the disease). By determining the last transition vector as the feature vector of the case to be analyzed, the server obtains a vector that can comprehensively reflect the various aspects of information about Ms. Liu's condition and has a dimension that meets the requirements of subsequent analysis, which can be directly used for comparison and analysis with the feature vector of the archived case, thereby providing an important data basis for screening similar cases to Ms. Liu's condition in medical clinical big data analysis, helping doctors to more accurately diagnose and treat Ms. Liu's condition, and also laying a good foundation for subsequent related data processing and analysis.
[0056] In the embodiments of the present application, the medical record feature distance between the archived medical record feature vector and the feature vector of the case to be analyzed can be implemented by the following examples.
[0057] loading the archived medical record feature vector and the feature vector of the case to be analyzed into a target integrated model; the target integrated model includes a feature retrieval component;
[0058] According to the value of each feature element in the archived medical record feature vector, a first search parameter of the archived medical record feature vector is obtained, and according to the value of each feature element in the medical record to be analyzed feature vector, a second search parameter of the medical record to be analyzed feature vector is obtained;
[0059] The archived medical record feature vector and the medical record to be analyzed feature vector are multiplied to obtain a search parameter coefficient, and according to the search parameter coefficient, the first search parameter and the second search parameter, a medical record feature distance between the archived medical record feature vector and the medical record to be analyzed feature vector is obtained.
[0060] In the embodiments of the present application, for example, the server is processing the relevant data of a patient named Mr. Zhao. Mr. Zhao comes to the hospital for medical treatment due to illness, and shows a series of symptoms such as abdominal pain, diarrhea, fever, and fatigue. The doctor conducts a comprehensive examination on him, including blood test, stool test, abdominal ultrasound, etc. After a series of data processing steps, the feature vector of the case to be analyzed about Mr. Zhao has been obtained. At the same time, the server needs to find out the cases similar to Mr. Zhao's condition from the archived medical records data in order to provide reference for the doctor, which requires calculating the medical record feature distance between the archived medical record feature vector and the case to be analyzed feature vector. The first thing the server needs to do is to load the archived medical record feature vector and the case to be analyzed feature vector into the target integrated model. This target integrated model is a comprehensive model developed and optimized by the hospital for a long time, which is specially used for processing medical clinical data related analysis tasks. It contains a feature retrieval component inside, which plays a key role in the subsequent calculation of medical record feature distance. The archived medical record feature vector is extracted from the hospital's medical record database, which corresponds to the medical record data of different patients in the past, and has undergone similar feature extraction and processing steps, condensing the key feature information of the patients' condition at that time. The case to be analyzed feature vector is obtained by processing a series of Mr. Zhao's clinical manifestation data, medical record data, etc., and also contains the key feature information of Mr. Zhao's condition. The server accurately loads the two vectors into the target integrated model, preparing for subsequent calculation and analysis, just like putting the sample to be detected and the known standard sample into a professional detection equipment for accurate comparison and analysis. In the target integrated model, the server will then obtain the first retrieval parameter of the archived medical record feature vector according to the values of each feature element in the archived medical record feature vector. Taking an archived medical record feature vector as an example, assuming that this vector corresponds to the medical record data of a Ms. Li who once had similar symptoms of abdominal pain and diarrhea. The archived medical record feature vector may contain multiple feature elements, such as the first element may represent a certain association between the patient's age and the severity of the condition at that time (a kind of quantitative relationship obtained by analyzing a large number of similar cases); the second element may reflect the influence degree of the patient's past medical history on the symptoms at that time (a kind of quantitative representation based on data analysis and expert experience); the third element may embody the synergistic change relationship between abdominal pain and diarrhea symptoms at that time (obtained from the analysis of various examination data and symptom records), etc. The server will perform specific operations on these feature element values according to the pre-set algorithm (which is based on a large amount of clinical data and expert experience), such as weighted sum, product or other complex combination operations, so as to obtain a first retrieval parameter that can represent the overall features of this archived medical record feature vector.Assuming that after the operation, the first retrieval parameter of Ms. Li's archived medical record feature vector is a specific value, such as 50 (here is just an example, the actual value will be determined according to the specific operation and data). Similarly, the server will obtain the second retrieval parameter of the feature vector of the medical record to be analyzed according to the values of the feature elements in the feature vector of the medical record to be analyzed. For Mr. Zhao's medical record feature vector to be analyzed, each feature element also has its own meaning. For example, the first element may represent the correlation between Mr. Zhao's age and the severity of the current symptoms; the second element may reflect the potential impact of Mr. Zhao's past medical history on the symptoms of this episode; the third element may reflect the synergistic change relationship between symptoms such as abdominal pain, diarrhea, fever, etc. in this episode. The server will also perform specific operations on these feature element values of Mr. Zhao's medical record feature vector to be analyzed according to the pre-set algorithm, so as to obtain a second retrieval parameter that can represent the overall characteristics of the medical record feature vector to be analyzed. Assuming that after the operation, the second retrieval parameter of Mr. Zhao's medical record feature vector to be analyzed is another specific value, such as 40 (also just an example, the actual value will be determined according to the specific operation and data). After obtaining the first retrieval parameter of the archived medical record feature vector and the second retrieval parameter of the medical record to be analyzed, the server will multiply the archived medical record feature vector and the medical record to be analyzed to obtain the retrieval parameter coefficient. Continue to take Ms. Li's archived medical record feature vector and Mr. Zhao's medical record feature vector to be analyzed as an example, assuming that Ms. Li's archived medical record feature vector is [1, 2, 3] (here is just a simple example, the actual vector dimension and element will be determined according to the specific situation), and Mr. Zhao's medical record feature vector to be analyzed is [4, 5, 6]. According to the rules of vector multiplication, multiply the corresponding elements of the two vectors and sum them up, that is: (1x4) + (2x5) + (3x6) = 4 + 10 + 18 = 32, so the retrieval parameter coefficient obtained here is 32. Finally, the server will obtain the medical record feature distance between the archived medical record feature vector and the medical record feature vector to be analyzed according to the retrieval parameter coefficient, the first retrieval parameter and the second retrieval parameter, through a specific calculation formula (this formula is also based on a large amount of clinical data and expert experience). Assuming that the first retrieval parameter of Ms. Li's archived medical record feature vector obtained earlier is 50, the second retrieval parameter of Mr. Zhao's medical record feature vector to be analyzed is 40, and the retrieval parameter coefficient is 32.According to the calculation formula (here it is assumed that the calculation formula is: medical record feature distance = √((first search parameter-search parameter coefficient)2+ (second search parameter-search parameter coefficient)2)), the corresponding numerical values are substituted to obtain: medical record feature distance = √((50-32)2+ (40-32)2) = √(182+82) = √(324+64) = √388≈19.7 (here one decimal place is retained) Through such calculation, the server obtains the medical record feature distance between the archived medical record feature vector (corresponding to the medical record of Ms. Li) and the to-be-analyzed medical record feature vector (corresponding to the medical record of Mr. Zhao). The value of the medical record feature distance can be used to judge the similarity degree of the medical record of Ms. Li and the medical record of Mr. Zhao in terms of disease characteristics. The smaller the distance, the higher the similarity degree, which means that the medical record of Ms. Li is more likely to have reference value for analyzing the disease of Mr. Zhao.
[0061] In addition, for the determination of the preset medical record feature distance threshold, the server first obtains a plurality of groups of archived medical record feature vectors corresponding to the archived clinical manifestation data from the preset medical record archive. For example, there are a large number of medical record data of patients with different diseases in the archive, and the server selects 10 groups of feature vectors corresponding to the medical records of patients who have similar symptoms of fever and cough but have different disease details. Then, the average medical record feature distance between each group of archived medical record feature vectors is calculated. Assuming that the first two archived medical record feature vectors are vector A and vector B, the medical record feature distance between them is calculated to be 5 by a specific distance calculation method (such as the method mentioned above based on vector element operation); the distance between the two vectors of the second group is calculated to be 6, and so on. Add the distance values of the 10 groups and divide by 10 to get the average medical record feature distance, which is assumed to be 7. Then, the average medical record feature distance is adjusted according to the preset proportion coefficient. If the preset proportion coefficient is 0.8, then 7 multiplied by 0.8 gives 5.6, which is the preset medical record feature distance threshold. It is used to judge the similarity between the analyzed medical record feature vector and the archived medical record feature vector in the subsequent process. The archived medical record with a distance not exceeding the threshold has more reference value. When processing the clinical manifestation data of a new patient, the server determines the mapping coefficient in the feature extraction subcomponent. For example, for a patient with abdominal pain and diarrhea, the server will count the number of clinical monitoring feature vectors N. Assuming that feature vectors are extracted from blood test, stool test and other test data, a total of 20 clinical monitoring feature vectors are obtained, i.e. N = 20. Then, according to the preset mapping function f(N), the mapping coefficient is calculated by substituting N. Since the preset mapping function satisfies the rule that the value of the mapping coefficient gradually decreases as N increases, when N = 20 is substituted, a specific mapping coefficient value is obtained through function calculation, such as 0.3. This mapping coefficient is used for subsequent operations of converting the feature elements indicating the spatial dimension channel in the clinical monitoring feature vector to the feature dimension channel. After obtaining the patient's medical record data, the server extracts the feature to obtain a description text feature vector. Taking a patient with heart disease as an example, a word vector embedding model trained based on a large amount of medical record text data in advance is used. The training goal of this model is to minimize the difference between the predicted text and the real text. The server inputs the text content in the medical record, such as the patient's medical history description, symptom onset, diagnosis and treatment process, etc. into the word vector embedding model. The model will convert the words in the medical record text into corresponding vector representations according to the knowledge learned before. For example, the word "palpitation" will be converted into a specific vector, and the word "chest tightness" will also have a corresponding vector. Through the processing of the entire medical record text, a description text feature vector is finally obtained, which integrates the vector representation of the key text information in the medical record, and is used for subsequent merging operation with the clinical manifestation feature encoding vector. When integrating the multi-dimensional merging vector, the server determines the sequence processing step.Take a multi-dimensional combined vector of a patient with multiple symptoms as an example.
[0062] The server first analyzes the distribution of the characteristic elements of the multi-dimensional combined vector, and calculates the standard deviation of the characteristic element distribution. Assuming that the multi-dimensional combined vector contains multiple characteristic elements extracted from the patient's various examinations and medical record information, the standard deviation is calculated to be 3 after analysis. Then, according to the mapping relationship table of the preset standard deviation and the sequence processing step length, the corresponding sequence processing step length value of the standard deviation 3 is found in the table, which is assumed to be 2. This sequence processing step length value is used for subsequent dimensionality reduction processing of the compressed feature vector and other operations. When calculating the medical record feature distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector, the feature retrieval component of the target integrated model of the server adopts a weighted calculation method. For example, for an archived medical record feature vector and a to-be-analyzed medical record feature vector, the server assigns a first weight value to each characteristic element in the archived medical record feature vector, and assigns a weight of 0.2 to the age-related characteristic element, a weight of 0.3 to the disease severity-related element, and so on. Assign a second weight value to each characteristic element in the to-be-analyzed medical record feature vector, and similarly assign different weights. When calculating the retrieval parameter coefficient, the first retrieval parameter and the second retrieval parameter, the corresponding characteristic element values and the corresponding weight values are multiplied and then subjected to subsequent operations. For example, when calculating the retrieval parameter coefficient, the element value of the archived medical record feature vector is 5, the weight is 0.2, the corresponding element value of the to-be-analyzed medical record feature vector is 4, and the weight is 0.3. Then, the product of 1 and 1.2 is obtained, and the original calculation method of the retrieval parameter coefficient is continued to operate, and finally the weighted medical record feature distance is obtained. This can more accurately measure the similarity between the two vectors and provide a more accurate basis for data screening.
[0063] The embodiment of the present application provides a computer device 100, which comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the data screening method based on medical clinical big data analysis as described above. Figure 2 As shown in the figure, Figure 2 The computer device 100 provided by the embodiment of the present application is a structural block diagram. The computer device 100 comprises a memory 111, a processor 112 and a communication unit 113. In order to realize the transmission or interaction of data, the memory 111, the processor 112 and the communication unit 113 are directly or indirectly electrically connected with each other. For example, the electrical connection between these elements can be realized by one or more communication buses or signal lines.
[0064] The foregoing description, for purposes of explanation, is provided as to specific embodiments and implementations. However, the foregoing description is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. Many modifications and variations are possible in light of the foregoing teaching. These embodiments and implementations were chosen and described in order to best explain the principles of the disclosure and its practical application, to thereby enable others skilled in the art to best utilize the disclosure, and to best enable others skilled in the art to best utilize the disclosure in various embodiments and with various modifications as are suited to the particular situation for each individual application.
Claims
1. A data screening method based on medical clinical big data analysis, characterized in that, The method comprises the following steps: feature extraction is performed on the clinical manifestation data to be analyzed to obtain a clinical manifestation feature encoding vector; medical record data is obtained, feature merging is performed on the medical record data and the clinical manifestation feature encoding vector to obtain a multi-dimensional merged vector, integration processing is performed on the multi-dimensional merged vector to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements; the medical record data is a description text used to describe the patient manifestation information corresponding to the extracted clinical manifestation data, the multi-dimensional merged vector includes expression content used to represent the patient manifestation information of the to-be-analyzed clinical manifestation data, and the preset integration number is equal to the number of feature elements of an archived medical record feature vector corresponding to an archived clinical manifestation data in a preset medical record archive; an archived medical record feature vector corresponding to the archived clinical manifestation data is obtained; the archived medical record feature vector includes expression content used to represent the patient manifestation information of the archived clinical manifestation data; a medical record feature distance between the archived medical record feature vector and the to-be-analyzed medical record feature vector is obtained, and the archived clinical manifestation data corresponding to the archived medical record feature vector whose medical record feature distance does not exceed a preset medical record feature distance threshold is determined as a target clinical manifestation data set; the target clinical manifestation data set is used to determine a target clinical manifestation data screening result for the to-be-analyzed clinical manifestation data.
2. The method of claim 1, wherein, The method of performing feature extraction on the to-be-analyzed clinical manifestation data to obtain a clinical manifestation feature encoding vector comprises the following steps: loading the to-be-analyzed clinical manifestation data into a target integrated model; the target integrated model comprises a multi-dimensional feature extraction component, and the multi-dimensional feature extraction component comprises an input subcomponent, a feature extraction subcomponent, and a feature optimization subcomponent; in the input subcomponent, data interception is performed on the to-be-analyzed clinical manifestation data according to a monitoring data time range and a monitoring data update frequency of the to-be-analyzed clinical manifestation data to obtain a plurality of intercepted monitoring data, and the plurality of intercepted monitoring data are respectively divided to obtain a plurality of clinical monitoring sub-data monitoring data time range and monitoring data update frequency; in the feature extraction subcomponent, feature extraction is performed on the plurality of clinical monitoring sub-data to obtain a plurality of clinical monitoring feature vectors, and feature elements in the plurality of clinical monitoring feature vectors indicating spatial dimension channels are re-converted to feature dimension channels according to a mapping coefficient to obtain a plurality of to-be-analyzed clinical manifestation feature vectors; the mapping coefficient is determined according to the number of the clinical monitoring feature vectors, and the number of feature elements of the clinical monitoring feature vectors is greater than the number of feature elements of the to-be-analyzed clinical manifestation feature vectors; in the feature optimization subcomponent, a low-rank joint compression vector is obtained, feature merging is performed on the to-be-analyzed clinical manifestation feature vectors and the low-rank joint compression vector to obtain a clinical manifestation feature vector set; the low-rank joint compression vector is a trainable parameter; the low-rank joint compression vector and the clinical manifestation feature vector set are placed into a plurality of execution operation modules, and the plurality of execution operation modules comprise a target execution operation module; In the target execution operation module, the low-rank joint compression vector is multiplied with the state vector to obtain a gating signal, the set of clinical manifestation feature vectors is multiplied with the set of input feature vectors to obtain a state representation, and the set of clinical manifestation feature vectors is multiplied with the set of output feature vectors to obtain transformed features; According to the gating signal and the state representation, a gating weight vector is obtained, the gating weight vector is dimensionally reduced according to the correlation measure of the state representation, the dimensionally reduced gating weight vector is normalized to obtain a gating control coefficient, the gating control coefficient is multiplied with the transformed features to obtain a gating output vector corresponding to the target execution operation module, and a clinical manifestation feature encoding vector is obtained according to a plurality of gating output vectors and gating control coefficients corresponding to the plurality of execution operation modules respectively. The number of feature elements of the clinical manifestation feature encoding vector is the same as the number of feature elements of the low-rank joint compression vector.
3. The method of claim 1, wherein, The medical record data is obtained, the medical record data and the clinical manifestation feature encoding vector are combined to obtain a multi-dimensional combined vector, and the multi-dimensional combined vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements, including: The clinical manifestation feature encoding vector is loaded into a target integrated model; the target integrated model includes a multi-dimensional feature extraction component and a feature integration processing component, and the multi-dimensional feature extraction component includes a feature extraction component; In the feature extraction component, medical record data is obtained, feature extraction is performed on the medical record data to obtain a description text feature vector, and the description text feature vector and the clinical manifestation feature encoding vector are combined to obtain a multi-dimensional combined vector; In the feature integration processing component, the multi-dimensional combined vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements.
4. The method of claim 3, wherein, In the feature integration processing component, the multi-dimensional combined vector is integrated to obtain a to-be-analyzed medical record feature vector with a preset integration number of feature elements, including: The multi-dimensional combined vector is loaded into the feature integration processing component, and the feature integration processing component includes a feature compression component and a feature reconstruction component; the feature reconstruction component is obtained by training according to a preset integration number, and the feature reconstruction component outputs a fixed number of feature elements, which is the preset integration number; In the feature compression component, the multi-dimensional combined vector is compressed to obtain a compressed feature vector, and the compressed feature vector is dimensionally reduced according to a sequence processing step to obtain a dimensionally reduced feature vector; In the feature reconstruction component, according to the feature correlation of the dimensionally reduced feature vector in each sequence processing step, a to-be-analyzed medical record feature vector with a preset integration number of feature elements is obtained.
5. The method of claim 4, wherein, The sequence processing step is determined in the following manner: Analyze the feature element distribution of the multi-dimensional combined vector to determine the standard deviation of the feature element distribution; According to a preset mapping relationship table of standard deviation and sequence processing step length, a sequence processing step length value corresponding to the standard deviation is searched.
6. The method of claim 3, wherein, The multi-dimensional merged vector is loaded into the feature integration processing component, and the feature integration processing component includes a clinical term analysis component and a feature weight adjustment component. In the clinical term analysis component, according to the multi-dimensional merged vector, a clinical relevance score corresponding to each clinical term in a clinical term library is obtained, patient performance information of the analyzed clinical performance data is obtained according to the clinical relevance score, and a plurality of clinical performance description segments are obtained by performing word segmentation on the patient performance information according to the preset integration number, wherein the preset integration number is a number corresponding to the plurality of clinical performance description segments. The plurality of clinical performance description segments correspond to respective appearance frequencies in the patient performance information, and the plurality of clinical performance description segments correspond to respective clinical term rarity indicators in the patient performance information, a plurality of clinical performance influence factors are obtained according to the plurality of appearance frequencies and the plurality of clinical term rarity indicators. In the feature weight adjustment component, the plurality of clinical performance influence factors are multiplied by the multi-dimensional merged vector respectively to obtain a plurality of clinical performance weight vectors, and a feature vector with a feature element number of the preset integration number is obtained according to the plurality of clinical performance weight vectors. The multi-dimensional merged vector includes a plurality of transition vectors, the feature element number of each of the plurality of transition vectors is the preset integration number, and the last transition vector in the plurality of transition vectors includes global expression content obtained by performing feature merging on the medical record data and the clinical performance feature encoding vector.
7. The method of claim 3, wherein, In the feature integration processing component, the last transition vector in the plurality of transition vectors is determined as the feature vector with the feature element number of the preset integration number. The medical record feature distance between the archived medical record feature vector and the analyzed medical record feature vector is obtained, including:
8. The method of claim 1, wherein, The archived medical record feature vector and the analyzed medical record feature vector are loaded into a target integrated model; the target integrated model includes a feature retrieval component; According to each feature element value in the archived medical record feature vector, a first retrieval parameter of the archived medical record feature vector is obtained, and according to each feature element value in the analyzed medical record feature vector, a second retrieval parameter of the analyzed medical record feature vector is obtained. The archived medical record feature vector and the analyzed medical record feature vector are multiplied to obtain a retrieval parameter coefficient, and a medical record feature distance between the archived medical record feature vector and the analyzed medical record feature vector is obtained according to the retrieval parameter coefficient, the first retrieval parameter and the second retrieval parameter. 9. The method of claim 1, wherein, The method for determining the preset medical record feature distance threshold comprises the following steps: Obtaining a plurality of groups of archived medical record feature vectors corresponding to the archived clinical manifestation data in a preset medical record archive, and calculating the average medical record feature distance between each group of archived medical record feature vectors; Adjusting the average medical record feature distance according to a preset proportion coefficient to obtain the preset medical record feature distance threshold.
10. A server system, characterized by The server is configured to perform the method of any one of claims 1-9.
Citation Information
Patent Citations
Comparative learning-based diagnosis coding method and system
CN114974602A
KR20220005674A