A method and system for constructing a medical examination database
By segmenting and identifying correlation information in the hospital database's test data, a database framework encompassing all laboratories is generated, solving the problem of weak data correlation in existing technologies and achieving more efficient and accurate data analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KUNSHAN TRADITIONAL CHINESE MEDICINE HOSPITAL
- Filing Date
- 2025-05-15
- Publication Date
- 2026-07-24
AI Technical Summary
The existing hospital database has weak correlations between data points, making it impossible to combine data from different laboratories for data analysis, resulting in poor data analysis efficiency and accuracy.
By acquiring test data from various laboratories, performing data segmentation and processing, identifying data characteristics and correlation information of test data groups, generating a database containing all laboratories, using clustering and similarity algorithms to identify explicit and implicit correlation information, and combining user type laboratory correlation information, a database framework containing all laboratories is constructed.
It enables comprehensive analysis of various detection data from multiple angles and in all aspects, improving the accuracy and efficiency of database data analysis.
Smart Images

Figure CN120763131B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data and database construction technology, and in particular to a method and system for constructing a medical examination database. Background Technology
[0002] Traditional hospital databases are built by integrating data from various laboratories and summarizing the data from each laboratory to create a comprehensive database. However, this method results in weak correlations between the data in the database, making it difficult to combine data from different laboratories for analysis. This significantly impacts the efficiency of data analysis for users and leads to poor accuracy in the data analysis of the constructed database. Summary of the Invention
[0003] The main objective of this invention is to provide a method and system for constructing a medical examination database, aiming to solve the problem that existing technologies suffer from poor data analysis accuracy due to weak correlation between data in the constructed database, making it impossible to combine data from various laboratories for data analysis, which greatly affects the efficiency of data analysis for users in various specialties.
[0004] To achieve the above objectives, the present invention provides a method for constructing a medical examination database, the method comprising:
[0005] Multiple test data from each laboratory are acquired, and the test data from each laboratory are divided and processed to obtain test data groups for each laboratory.
[0006] For each laboratory, data features of each test data group in the laboratory are extracted, and based on the data features of each test data group, data association information between the test data groups is identified;
[0007] Based on the data characteristics of each user type in the test data groups of each laboratory, laboratory association information between each laboratory is identified, and a database containing all laboratories is generated based on the data association information between the test data groups and the laboratory association information between each laboratory.
[0008] Optionally, the step of dividing the test data from each laboratory into data groups for each laboratory includes:
[0009] For each laboratory, identify the data source information and data identification information of each test data of the laboratory, and identify the test user corresponding to each test data based on the data source information of each test data;
[0010] In the user database, query the user type to which each detected user belongs, and based on the data identifier information of each detected data, query the data type of the detected data in the detection database;
[0011] Based on the user type and data type of each test data, the test data are divided into groups to obtain the test data groups of the laboratory.
[0012] Optionally, the extraction of data features from each set of test data in the laboratory includes:
[0013] For each detection data group, a clustering algorithm is used to identify the central detection data of the detection data group;
[0014] Based on the data type of the central detection data and the user type corresponding to the central detection data, data feature extraction processing is performed on the central detection data to obtain the data features of the central detection data.
[0015] The data characteristics of the central detection data are used as the data characteristics of the detection data group.
[0016] Optionally, identifying the data association information between the detection data groups based on the data characteristics of each detection data group includes:
[0017] Each of the data features is vectorized to obtain a feature vector corresponding to each data feature. Then, a similarity algorithm is used to calculate the similarity between each feature vector to obtain the data similarity between each detection data group.
[0018] The data similarity between each detection data group is used as the basis for determining whether there is an explicit correlation between the detection data groups.
[0019] The correlation between each data feature is calculated using a feature matching algorithm, and the detection data group to which each data feature has a correlation is considered as a detection data group with implicit correlation information.
[0020] The explicit correlation information between each of the detection data groups and the implicit correlation information between each of the detection data groups are used as the data correlation information between each of the detection data groups.
[0021] Optionally, identifying laboratory association information between the laboratories based on the data characteristics of each user type's test data sets in each of the laboratories includes:
[0022] In each laboratory's test data set, the target test data sets for each user type in each laboratory are filtered, and for each user type, based on the target data characteristics of the user type in the target test data sets of each laboratory, the explicit association information of the user type between the laboratories and the implicit association information of the user type between the laboratories are identified.
[0023] The explicit association information of each user type among the laboratories and the implicit association information of each user type among the laboratories are used as the laboratory association information among the laboratories.
[0024] Optionally, the step of generating a database containing all laboratories based on the data association information between the various test data groups and the laboratory association information between the various laboratories includes:
[0025] Based on the data association information between the various test data groups of each laboratory, a data association map of the laboratory is generated.
[0026] Based on the laboratory association information between the laboratories, the data association map of each laboratory is processed by map linking to obtain an association map containing all laboratories, and a database framework is generated based on the association map.
[0027] The test data of each user in each laboratory are stored in the database framework according to the user type and the data type corresponding to the test data of each user, so as to obtain a database containing all laboratories.
[0028] Furthermore, to achieve the above objectives, the present invention also provides a system for constructing a medical examination database, the system comprising:
[0029] The acquisition module is used to acquire multiple test data from each laboratory and perform data segmentation processing on the test data from each laboratory to obtain test data groups for each laboratory.
[0030] The identification module is used to extract data features of each test data group of each laboratory, and to identify data association information between each test data group based on the data features of each test data group;
[0031] The generation module is used to identify laboratory association information between the laboratories based on the data characteristics of the test data groups of each user type in each of the laboratories, and to generate a database containing all laboratories based on the data association information between the test data groups and the laboratory association information between the laboratories.
[0032] Optionally, the acquisition module is specifically used for:
[0033] For each laboratory, identify the data source information and data identification information of each test data of the laboratory, and identify the test user corresponding to each test data based on the data source information of each test data;
[0034] In the user database, query the user type to which each detected user belongs, and based on the data identifier information of each detected data, query the data type of the detected data in the detection database;
[0035] Based on the user type and data type of each test data, the test data are divided into groups to obtain the test data groups of the laboratory.
[0036] Optionally, the identification module is specifically used for:
[0037] For each detection data group, a clustering algorithm is used to identify the central detection data of the detection data group;
[0038] Based on the data type of the central detection data and the user type corresponding to the central detection data, data feature extraction processing is performed on the central detection data to obtain the data features of the central detection data.
[0039] The data characteristics of the central detection data are used as the data characteristics of the detection data group.
[0040] Optionally, the identification module is specifically used for:
[0041] Each of the data features is vectorized to obtain a feature vector corresponding to each data feature. Then, a similarity algorithm is used to calculate the similarity between each feature vector to obtain the data similarity between each detection data group.
[0042] The data similarity between each detection data group is used as the basis for determining whether there is an explicit correlation between the detection data groups.
[0043] The correlation between each data feature is calculated using a feature matching algorithm, and the detection data group to which each data feature has a correlation is considered as a detection data group with implicit correlation information.
[0044] The explicit correlation information between each of the detection data groups and the implicit correlation information between each of the detection data groups are used as the data correlation information between each of the detection data groups.
[0045] Optionally, the generation module is specifically used for:
[0046] In each laboratory's test data set, the target test data sets for each user type in each laboratory are filtered, and for each user type, based on the target data characteristics of the user type in the target test data sets of each laboratory, the explicit association information of the user type between the laboratories and the implicit association information of the user type between the laboratories are identified.
[0047] The explicit association information of each user type among the laboratories and the implicit association information of each user type among the laboratories are used as the laboratory association information among the laboratories.
[0048] Optionally, the generation module is specifically used for:
[0049] Based on the data association information between the various test data groups of each laboratory, a data association map of the laboratory is generated.
[0050] Based on the laboratory association information between the laboratories, the data association map of each laboratory is processed by map linking to obtain an association map containing all laboratories, and a database framework is generated based on the association map.
[0051] The test data of each user in each laboratory are stored in the database framework according to the user type and the data type corresponding to the test data of each user, so as to obtain a database containing all laboratories.
[0052] Thirdly, this application provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in any one of the first aspects.
[0053] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0054] Fifthly, this application provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0055] This invention provides a method and system for constructing a medical examination database. The method includes: acquiring multiple test data from various laboratories, and dividing the test data from each laboratory into data groups for each laboratory; for each laboratory, extracting data features from each test data group, and identifying data association information between the test data groups based on the data features of each test data group; identifying laboratory association information between the laboratories based on the data features of each user type in the test data groups of each laboratory, and generating a database containing all laboratories based on the data association information between the test data groups and the laboratory association information between the laboratories. This solution identifies the data association information between the test data groups of each laboratory by grouping the actual test data of each laboratory and then identifying the data features of each test data group. Then, by identifying the data characteristics of each user type across the test data groups in each laboratory, the laboratory association information between each laboratory is identified, thereby constructing a database containing the laboratories. This not only enables comprehensive analysis of test data from the same laboratory based on the data association information between test data groups, but also enables comprehensive analysis of test data from different laboratories by combining the laboratory association information between laboratories. As a result, the database constructed through this solution can comprehensively analyze test data from multiple angles and in all aspects, thereby effectively improving the accuracy of data analysis in the constructed database. Attached Figure Description
[0056] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart of a database construction method provided in an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram of the structure of the database construction system provided in this embodiment of the invention;
[0059] Figure 3 An internal structural diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0060] The database construction method provided in this invention is applied to a database construction system. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application. The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or accompanying drawings of this application are used to distinguish different objects, not to describe a particular order.
[0061] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0062] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0063] The database construction method provided in this application embodiment can be applied to database construction environments. This method can be applied to terminals, servers, and systems including both terminals and servers, and is implemented through interaction between the terminal and server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. The terminal groups the actual test data from each laboratory, identifies the data characteristics of each test data group, and thus identifies the data association information between the test data groups of each laboratory. Then, based on the data characteristics of each user type among the test data groups of each laboratory, it identifies the laboratory association information between each laboratory, thereby constructing a database containing laboratories. This not only enables comprehensive analysis of test data from the same laboratory based on the data association information between test data groups, but also enables comprehensive analysis of test data from different laboratories by combining the laboratory association information between different laboratories. Therefore, the database constructed through this solution can comprehensively analyze test data from multiple angles and in all aspects, thereby effectively improving the data analysis accuracy of the constructed database.
[0064] In one embodiment, such as Figure 1As shown, a method for constructing a medical examination database is provided. Taking the application of this method to a terminal as an example, the method includes the following steps:
[0065] Step S101: Obtain multiple test data from each laboratory, and perform data segmentation processing on the test data from each laboratory to obtain each test data group for each laboratory.
[0066] In this embodiment, the terminal receives test data from each laboratory's testing equipment, thus obtaining test data from each laboratory. Each test data point represents test data from a different user. These laboratories include, but are not limited to, dynamic electrocardiogram (ECG) rooms, transcranial Doppler (TCD) rooms, ECG rooms, electroencephalogram (EEG) rooms, dynamic blood pressure monitoring rooms, electromyography (EMG) rooms, pulmonary function testing rooms, visual function testing rooms, hearing testing rooms, and skin examination rooms. The terminal then groups the test data according to the data type and the user type of the user to whom the data belongs, resulting in separate test data groups. User types include, but are not limited to, different age groups (age ranges, e.g., 0-15 years, 15-30 years, 30-45 years, 45-60 years, 60-75 years, 75-90 years, over 90 years old), different genders, and different physical states (e.g., healthy, slightly ill, sick, seriously ill, etc., where each physical state is a range determined by staff based on the user's specific illness). Each data point corresponds to a data type, which in turn corresponds to the instrument type of the testing instrument. This laboratory includes, but is not limited to, a Holter monitor room, a transcranial Doppler room, an electrocardiogram (ECG) room, an electroencephalogram (EEG) room, a Holter monitor room, an electromyography (EMG) room, a pulmonary function testing room, a visual function testing room, a hearing testing room, and a skin testing room. The specific grouping process will be explained in detail later.
[0067] Step S102: For each laboratory, extract the data features of each test data group in the laboratory, and based on the data features of each test data group, identify the data association information between each test data group.
[0068] In this embodiment, the terminal extracts data features from each test data group for each laboratory and identifies data association information between the test data groups based on these features. This data association information includes explicit and implicit association information between the test data groups. Explicit association information represents the similarity between the data features of each test data group, while implicit association information represents the relationship between the data features of each test data group. The specific processes for extracting data features and identifying association information will be described in detail later.
[0069] Step S103: Based on the data characteristics of the test data groups of each user type in each laboratory, identify the laboratory association information between each laboratory, and generate a database containing all laboratories based on the data association information between each test data group and the laboratory association information between each laboratory.
[0070] In this embodiment, the terminal identifies laboratory association information between laboratories based on the data characteristics of each user type's test data sets in each laboratory. Based on this data association information, a database containing all laboratories is generated. This laboratory association information refers to the correlation between the test data characteristics of different laboratories for different user types, characterizing the feature data associations between different user types in each laboratory. For example, based on the data characteristics of test data sets detected in the ECG testing room and the data characteristics of test data sets detected in the electromyography (EMG) testing room for each user type aged 15-30, the terminal analyzes the association information between the ECG and EMG testing rooms for the 15-30 year old user type. This association information includes, for example, the association between heart rate change information and muscle state detection results. For the 30-45 year old female user type, the data characteristics of test data sets detected in the EEG testing room and the data characteristics of test data sets detected in the transcranial Doppler (TCD) room include, for example, the association between EEG signal discharge detection results and cerebral blood flow state detection. The analysis of laboratory association information is used to improve the database association information between multiple laboratories when constructing each database, so as to ensure that when analyzing and identifying relevant test data of user types, the test results associated with and corresponding to different laboratories can be detected simultaneously.
[0071] The laboratory association information includes the similarity and correlation between data features of each user type in the test data groups of each laboratory. The process for identifying the similarity and correlation between data features is the same as in step S102. Each user type can have one or more test data groups in each laboratory. When there are multiple test data groups for each user type in each laboratory, it is necessary to separately determine the similarity and correlation between each test data group in each laboratory to obtain the laboratory association information. The process of constructing a database containing all laboratories will be described in detail later.
[0072] Based on the above scheme, after grouping the actual test data from each laboratory, the data characteristics of each test data group are identified, thereby identifying the data correlation information between the test data groups of each laboratory. Then, by identifying the data characteristics of each user type among the test data groups of each laboratory, the laboratory correlation information between each laboratory is identified, thus constructing a database containing laboratories. This not only enables comprehensive analysis of test data from the same laboratory based on the data correlation information between test data groups, but also enables comprehensive analysis of test data from different laboratories by combining the laboratory correlation information between different laboratories. Therefore, the database constructed through this scheme can comprehensively analyze test data from multiple angles and in all aspects, thereby effectively improving the accuracy of data analysis in the constructed database.
[0073] Optionally, the test data from each laboratory is segmented to obtain test data groups for each laboratory. This includes: for each laboratory, identifying the data source information and data identification information for each test data, and identifying the test user corresponding to each test data based on the data source information; querying the user type of each test user in the user database, and querying the data type of the test data in the test database based on the data identification information of each test data; and segmenting the test data based on the user type and data type of each test data to obtain test data groups for each laboratory.
[0074] In this embodiment, the terminal identifies the data source information and data identifier information of each test data for each laboratory, and identifies the test user corresponding to each test data based on the data source information of each test data. The data source information of each test data includes the user information of the test user for that test data, which includes the user's identity information (age, gender) and the user's physical condition.
[0075] In the user database, the terminal queries the user type of each user based on their age group, gender, and physical condition, and then queries the data type of each data point in the detection database based on its data identifier. This data identifier is the device identifier of the detection device corresponding to the data; each device identifier corresponds to one detection device.
[0076] Finally, the terminal performs data segmentation based on the user type and data type of each test data point, resulting in test data groups for the laboratory. Specifically, the terminal divides each test data point into initial test data groups based on the user type, then further segments each initial test data group based on its data type, obtaining test data groups for each data type. Finally, the terminal combines all test data groups into the laboratory's test data groups.
[0077] Based on the above scheme, the data is divided from multiple perspectives, such as the user type to which each data belongs and the data type of each data, thereby improving the comprehensiveness and precision of the data division.
[0078] Optionally, extract data features from each test data set in the laboratory, including: for each test data set, using a clustering algorithm to identify the central test data of the test data set; based on the data type of the central test data and the user type corresponding to the central test data, perform data feature extraction processing on the central test data to obtain the data features of the central test data; and use the data features of the central test data as the data features of the test data set.
[0079] In this embodiment, for each detection data group, the terminal identifies the central detection data of the group using a clustering algorithm. Based on the data type and user type of the central detection data, the terminal performs data feature extraction processing to obtain the data features of the central detection data. Based on the data type of the central detection data, the terminal identifies the corresponding data modality in the database, where data modalities include text modalities, image modalities, etc. Then, the terminal uses the feature extraction network corresponding to the data modality of the central detection data to perform data extraction processing on the central detection data to obtain the data features of each central detection data. For example, the feature extraction network corresponding to each data modality is a deep learning neural network based on natural language processing technology for text modalities, and a convolutional neural network based on a self-attention mechanism for image modalities.
[0080] Finally, the terminal uses the data characteristics of the central detection data of each detection data group as the data characteristics of each detection data group.
[0081] Based on the above scheme, by clustering to identify the central detection data of each detection data group and then extracting data features, not only is the amount of data features optimized and the extraction efficiency improved, but also the interference problem of discrete data affecting data features can be avoided for the central detection data, thereby improving the accuracy of the extracted data features.
[0082] Optionally, based on the data features of each detection data group, the data association information between each detection data group is identified, including: vectorizing each data feature to obtain the feature vector corresponding to each data feature, and calculating the similarity between each feature vector using a similarity algorithm to obtain the data similarity between each detection data group; using the data similarity between each detection data group as each detection data group with explicit association information; calculating the association relationship between each data feature using a feature matching algorithm, and using the detection data group to which each data feature with an association relationship belongs as each detection data group with implicit association information; and using both the explicit association information and the implicit association information between each detection data group as the data association information between each detection data group.
[0083] In this embodiment, the terminal vectorizes each data feature to obtain a feature vector corresponding to each data feature. This vectorization method is based on the TF-IDF (Term Frequency-Inverse Document Frequency) bag-of-words model, which vectorizes each feature vector.
[0084] Then, the terminal calculates the similarity between each feature vector using a similarity algorithm to obtain the data similarity between each detection data group. This similarity algorithm is a cosine similarity algorithm. Finally, the terminal uses the data similarity between each detection data group as the basis for determining which detection data groups exhibit explicit correlation information.
[0085] The terminal first establishes a feature matrix containing feature vectors of all data features. Then, based on this feature matrix, the terminal calculates the feature matching degree between each data feature using a feature matching algorithm, and uses the feature matching degree as the association relationship between the data features. Next, the terminal identifies the detection data groups to which the data features with association relationships belong as detection data groups with implicit association information. The feature matching algorithm can be, but is not limited to, the SURF (Speeded Up Robust Features) feature matching algorithm.
[0086] Finally, the terminal uses both the explicit and implicit correlation information between the detection data groups as the data correlation information between the detection data groups.
[0087] Based on the above scheme, after feature vectorization, the data association information between each detection data group is identified based on the similarity between each feature vector and the correlation between each feature vector, thereby improving the comprehensiveness and accuracy of identifying the data association information between each detection data group.
[0088] Optionally, based on the data characteristics of the test data sets of each user type in each laboratory, laboratory association information is identified, including: in the test data sets of each laboratory, filtering the target test data sets of each user type in each laboratory, and for each user type, based on the target data characteristics of the target test data sets of the user type in each laboratory, identifying explicit association information and implicit association information between user types in each laboratory; and using the explicit association information and implicit association information between user types in each laboratory as laboratory association information between laboratories.
[0089] In this embodiment, the terminal filters the target detection data groups for each user type in each laboratory's detection data group. For each user type, based on the target data features of the target detection data groups in each laboratory, it identifies the explicit and implicit correlation information between user types across laboratories. The identification method for the explicit and implicit correlation information between the target data features of each target detection data group is the same as described above. The terminal uses the explicit and implicit correlation information between the target data features of each user type in each target detection data group in each laboratory as the explicit and implicit correlation information between that user type and each laboratory.
[0090] Finally, the terminal uses the explicit association information of each user type between each laboratory, as well as the implicit association information of each user type between each laboratory, as the laboratory association information between each laboratory.
[0091] Based on the above scheme, by separating each user type to identify explicit and implicit correlations between laboratories, the comprehensiveness and accuracy of the identification are improved.
[0092] Optionally, based on the data association information between each test data group and the laboratory association information between each laboratory, a database containing all laboratories is generated, including: generating a data association map of the laboratory based on the data association information between each test data group of the laboratory for each laboratory; performing map linking processing on the data association map of each laboratory based on the laboratory association information between laboratories to obtain an association map containing all laboratories, and generating a database framework based on the association map; storing the test data of each user in each laboratory according to the user type and the data type corresponding to the test data of each user in the database framework to obtain a database containing all laboratories.
[0093] In this embodiment, the terminal generates a data association graph for each laboratory based on the data association information between the various test data groups within that laboratory. This data association graph is a knowledge graph constructed from the data association information between different user types and data types.
[0094] Then, based on the laboratory association information between laboratories, the terminal performs graph linking processing on the data association graph of each laboratory to obtain an association graph containing all laboratories. Each data association graph, after graph linking processing, yields its association chains, where each chain represents the explicit and implicit association information between each user type and the various laboratories.
[0095] Next, the terminal generates a database framework based on the association graph. This database framework includes multiple storage areas. Each user type in each laboratory corresponds to a storage area under different data type conditions. There are data retrieval links between the storage areas. These links are the association chains of the association graph. When analyzing data information of different user types or different data types, these links allow the retrieval of detection data from only a single storage area to obtain all the related detection data.
[0096] Finally, the terminal stores the test data of each user in each laboratory in the database framework according to the user type and the data type of each user's test data, resulting in a database containing all laboratories.
[0097] Based on the above scheme, by identifying the data characteristics of each user type among the test data groups in each laboratory, the laboratory association information between each laboratory is identified, thereby constructing a database containing laboratories. This not only enables comprehensive analysis of test data from the same laboratory based on the data association information between test data groups, but also enables comprehensive analysis of test data from different laboratories by combining the laboratory association information between laboratories. As a result, the database constructed through this scheme can comprehensively analyze test data from multiple angles and in all aspects, thereby effectively improving the accuracy of data analysis in the constructed database.
[0098] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0099] Based on the same inventive concept, this application also provides a database construction system for implementing the database construction method described above. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations of one or more database construction system embodiments provided below can be found in the limitations of the database construction method described above, and will not be repeated here.
[0100] Further reference Figure 2 As a response to the above Figure 1 The present application provides an embodiment of a medical examination database construction system 200, which includes an acquisition module 210, an identification module 220, and a generation module 230, wherein:
[0101] The acquisition module 210 is used to acquire multiple test data from each laboratory and perform data segmentation processing on the test data from each laboratory to obtain each test data group from each laboratory.
[0102] The identification module 220 is used to extract data features of each test data group of each laboratory, and identify data association information between each test data group based on the data features of each test data group;
[0103] The generation module 230 is used to identify laboratory association information between the laboratories based on the data characteristics of the test data groups of each user type in each of the laboratories, and to generate a database containing all laboratories based on the data association information between the test data groups and the laboratory association information between the laboratories.
[0104] Optionally, the acquisition module 210 is specifically used for:
[0105] For each laboratory, identify the data source information and data identification information of each test data of the laboratory, and identify the test user corresponding to each test data based on the data source information of each test data;
[0106] In the user database, query the user type to which each detected user belongs, and based on the data identifier information of each detected data, query the data type of the detected data in the detection database;
[0107] Based on the user type and data type of each test data, the test data are divided into groups to obtain the test data groups of the laboratory.
[0108] Optionally, the identification module 220 is specifically used for:
[0109] For each detection data group, a clustering algorithm is used to identify the central detection data of the detection data group;
[0110] Based on the data type of the central detection data and the user type corresponding to the central detection data, data feature extraction processing is performed on the central detection data to obtain the data features of the central detection data.
[0111] The data characteristics of the central detection data are used as the data characteristics of the detection data group.
[0112] Optionally, the identification module 220 is specifically used for:
[0113] Each of the data features is vectorized to obtain a feature vector corresponding to each data feature. Then, a similarity algorithm is used to calculate the similarity between each feature vector to obtain the data similarity between each detection data group.
[0114] The data similarity between each detection data group is used as the basis for determining whether there is an explicit correlation between the detection data groups.
[0115] The correlation between each data feature is calculated using a feature matching algorithm, and the detection data group to which each data feature has a correlation is considered as a detection data group with implicit correlation information.
[0116] The explicit correlation information between each of the detection data groups and the implicit correlation information between each of the detection data groups are used as the data correlation information between each of the detection data groups.
[0117] Optionally, the generation module 230 is specifically used for:
[0118] In each laboratory's test data set, the target test data sets for each user type in each laboratory are filtered, and for each user type, based on the target data characteristics of the user type in the target test data sets of each laboratory, the explicit association information of the user type between the laboratories and the implicit association information of the user type between the laboratories are identified.
[0119] The explicit association information of each user type among the laboratories and the implicit association information of each user type among the laboratories are used as the laboratory association information among the laboratories.
[0120] Optionally, the generation module 230 is specifically used for:
[0121] Based on the data association information between the various test data groups of each laboratory, a data association map of the laboratory is generated.
[0122] Based on the laboratory association information between the laboratories, the data association map of each laboratory is processed by map linking to obtain an association map containing all laboratories, and a database framework is generated based on the association map.
[0123] The test data of each user in each laboratory are stored in the database framework according to the user type and the data type corresponding to the test data of each user, so as to obtain a database containing all laboratories.
[0124] The modules in the aforementioned database construction system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0125] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, communication interface, display screen, and input system connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for constructing a medical examination database. The display screen can be an LCD screen or an e-ink screen. The input system can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0126] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0127] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any one of the first aspects.
[0128] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0129] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0131] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0132] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0133] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for constructing a medical examination database, characterized in that, The method includes: Multiple test data from each laboratory are acquired, and the test data from each laboratory are divided and processed to obtain test data groups for each laboratory. For each laboratory, data features of each test data group in the laboratory are extracted, and based on the data features of each test data group, data association information between the test data groups is identified; Based on the data characteristics of each user type in the test data groups of each laboratory, the laboratory association information between each laboratory is identified, and a database containing all laboratories is generated based on the data association information between the test data groups and the laboratory association information between each laboratory. The step of identifying data association information between the detection data groups based on the data characteristics of each detection data group includes: Each of the data features is vectorized to obtain a feature vector corresponding to each data feature. Then, a similarity algorithm is used to calculate the similarity between each feature vector to obtain the data similarity between each detection data group. The data similarity between each detection data group is used as the basis for determining whether there is an explicit correlation between the detection data groups. The correlation between each data feature is calculated using a feature matching algorithm, and the detection data group to which each data feature has a correlation is considered as a detection data group with implicit correlation information. The explicit correlation information between each of the detection data groups and the implicit correlation information between each of the detection data groups are used as the data correlation information between each of the detection data groups. The method of identifying laboratory association information between the laboratories based on the data characteristics of each user type's test data groups in each of the laboratories includes: In each laboratory's test data set, the target test data sets for each user type in each laboratory are filtered, and for each user type, based on the target data characteristics of the user type in the target test data sets of each laboratory, the explicit association information of the user type between the laboratories and the implicit association information of the user type between the laboratories are identified. The explicit association information of each user type among the laboratories and the implicit association information of each user type among the laboratories are used as the laboratory association information among the laboratories.
2. The method according to claim 1, characterized in that, The step of dividing the test data from each laboratory into data groups for each laboratory includes: For each laboratory, identify the data source information and data identification information of each test data of the laboratory, and identify the test user corresponding to each test data based on the data source information of each test data; In the user database, query the user type to which each detected user belongs, and based on the data identifier information of each detected data, query the data type of the detected data in the detection database; Based on the user type and data type of each test data, the test data are divided into groups to obtain the test data groups of the laboratory.
3. The method according to claim 2, characterized in that, The extraction of data features from each set of test data in the laboratory includes: For each detection data group, a clustering algorithm is used to identify the central detection data of the detection data group; Based on the data type of the central detection data and the user type corresponding to the central detection data, data feature extraction processing is performed on the central detection data to obtain the data features of the central detection data. The data characteristics of the central detection data are used as the data characteristics of the detection data group.
4. The method according to claim 1, characterized in that, The database containing all laboratories is generated based on the data association information between the various test data groups and the laboratory association information between the various laboratories, including: For each laboratory, a data association map is generated based on the data association information between the various test data groups of the laboratory; Based on the laboratory association information between the laboratories, the data association map of each laboratory is processed by map linking to obtain an association map containing all laboratories, and a database framework is generated based on the association map. The test data of each user in each laboratory are stored in the database framework according to the user type and the data type corresponding to the test data of each user, so as to obtain a database containing all laboratories.
5. A system for constructing a medical examination database, characterized in that, The system includes: The acquisition module is used to acquire multiple test data from each laboratory and perform data segmentation processing on the test data from each laboratory to obtain test data groups for each laboratory. The identification module is used to extract data features of each test data group of each laboratory, and to identify data association information between each test data group based on the data features of each test data group; The generation module is used to identify laboratory association information between the laboratories based on the data characteristics of the test data groups of each user type in each of the laboratories, and to generate a database containing all laboratories based on the data association information between the test data groups and the laboratory association information between the laboratories. The identification module identifies data association information between the detection data groups based on the data characteristics of each detection data group, including: Each of the data features is vectorized to obtain a feature vector corresponding to each data feature. Then, a similarity algorithm is used to calculate the similarity between each feature vector to obtain the data similarity between each detection data group. The data similarity between each detection data group is used as the basis for determining whether there is an explicit correlation between the detection data groups. The correlation between each data feature is calculated using a feature matching algorithm, and the detection data group to which each data feature has a correlation is considered as a detection data group with implicit correlation information. The explicit correlation information between each of the detection data groups and the implicit correlation information between each of the detection data groups are used as the data correlation information between each of the detection data groups. The generation module identifies laboratory association information between the laboratories based on the data characteristics of each user type's test data groups in each of the laboratories, including: In each laboratory's test data set, the target test data sets for each user type in each laboratory are filtered, and for each user type, based on the target data characteristics of the user type in the target test data sets of each laboratory, the explicit association information of the user type between the laboratories and the implicit association information of the user type between the laboratories are identified. The explicit association information of each user type among the laboratories and the implicit association information of each user type among the laboratories are used as the laboratory association information among the laboratories.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.