A sample matching method, device, apparatus and storage medium
By identifying a first group and a second group in the experimental and control groups, and using neural networks and clustering algorithms for sample matching, the problem of inaccurate experimental results caused by differences in covariates between the experimental and control groups was solved, thus improving the accuracy of sample matching and the experimental results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST
- Filing Date
- 2022-10-31
- Publication Date
- 2026-08-04
AI Technical Summary
When studying the effect of any independent variable on experimental results, existing techniques lead to inaccurate experimental results due to differences in various covariates between the experimental and control groups.
By determining the first and second groups for the experimental and control groups, and using neural networks and clustering algorithms, candidate sample sets for each first sample are determined from the second group based on the classification probability and similarity of the samples, thus achieving accurate sample matching and eliminating covariate differences.
Without reducing the sample size, we can ensure that the samples in different groups remain consistent across all covariates after matching, thereby improving the accuracy of the experimental results.
Smart Images

Figure CN115691821B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a sample matching method, apparatus, device, and storage medium. Background Technology
[0002] In any research field, when studying the influence of any independent variable on the experimental results, it is usually affected by some uncontrollable covariates. For example, in the medical field, medical institutions often use Diagnosis Related Groups (DRG) management tools to group patients. Based on whether patients use DRG tools for payment, corresponding DRG groups and non-DRG groups can be obtained. Then, by comparing the payment situations of different patient cases within the two groups, the cost control effect of DRG in the medical field can be studied. However, using basic information such as gender, age, length of hospital stay, and surgical level of patients within the DRG and non-DRG groups as covariates will introduce certain covariate differences, making it impossible to ensure the accuracy of the cost control effect of DRG.
[0003] Therefore, in order to accurately analyze the experimental effect of any independent variable, for the two groups selected using that independent variable, it is first necessary to eliminate the differences in covariates between the two groups. Typically, by limiting the values of the covariates referenced when selecting the groups, the two groups specified by that independent variable can maintain basic consistency in each covariate.
[0004] However, the above method requires that the number of covariates not be too large to avoid increasing the complexity of the preliminary work. Moreover, it will correspondingly reduce the number of samples in the two groups specified by any independent variable, resulting in missing experimental samples for that independent variable, which will affect the accuracy of the experimental results for that independent variable. Summary of the Invention
[0005] This application provides a sample matching method, apparatus, device, and storage medium to achieve accurate matching of samples within different groups specified by any independent variable. This ensures that the matched samples within different groups maintain a basic consistency across various covariates, thereby eliminating covariate differences among samples within different groups and ensuring the accuracy of experimental results under any independent variable.
[0006] In a first aspect, embodiments of this application provide a sample matching method, the method comprising:
[0007] Determine the first and second groups in the experimental and control groups specified by any independent variable;
[0008] Based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group, determine the first candidate sample set corresponding to each first sample from the second group;
[0009] Based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, a second candidate sample set corresponding to each first sample is determined from the second group;
[0010] Based on the first candidate sample set and the second candidate sample set corresponding to each first sample, a matching sample for each first sample is determined to form a matching group for the first group.
[0011] Secondly, embodiments of this application provide a sample matching device, the device comprising:
[0012] The grouping determination module is used to determine the first and second groups in the experimental and control groups for any given independent variable;
[0013] The first candidate sample determination module is used to determine the first candidate sample set corresponding to each first sample from the second group based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group;
[0014] The second candidate sample determination module is used to determine the second candidate sample set corresponding to each first sample from the second group based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering.
[0015] The sample matching module is used to determine the matching samples of each first sample based on the first candidate sample set and the second candidate sample set corresponding to each first sample, so as to form the matching group of the first group.
[0016] Thirdly, embodiments of this application provide an electronic device, which includes:
[0017] A processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the sample matching method provided in the first aspect of this application.
[0018] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that causes a computer to perform the sample matching method provided in the first aspect of this application.
[0019] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the sample matching method as provided in the first aspect of this application.
[0020] This application provides a sample matching method, apparatus, device, and storage medium. First, a first group and a second group are determined within the experimental and control groups specified by any independent variable. Then, based on the first classification probability of a first sample within the first group and the second classification probability of a second sample within the second group, a first candidate sample set corresponding to each first sample is determined from the second group. Furthermore, based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, a second candidate sample set corresponding to each first sample is determined from the second group. Subsequently, matching samples for each first sample are determined from the first and second candidate sample sets corresponding to each first sample, thus forming a matching group for the first group. This achieves accurate matching of samples within different groups specified by any independent variable, ensuring that the matched samples within different groups maintain basic consistency across various covariates. Therefore, without reducing covariates and while ensuring the comprehensiveness of samples within different groups, the differences in covariates among samples within different groups can be eliminated, ensuring the accuracy of the experimental results under any independent variable. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a sample matching method according to an embodiment of this application;
[0023] Figure 2 This is a schematic diagram illustrating the principle of determining any one of the first and second subclusters in an embodiment of this application.
[0024] Figure 3 A flowchart illustrating another sample matching method as shown in an embodiment of this application;
[0025] Figure 4 This is a schematic diagram illustrating the principle of predicting the classification probability of either the first sample or the second sample in an embodiment of this application.
[0026] Figure 5 This is a schematic block diagram of a sample matching device according to an embodiment of this application;
[0027] Figure 6 This is a schematic block diagram of an electronic device shown in an embodiment of this application. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0030] To address the issue of inaccurate experimental results caused by differences in various covariates between two groups under a given independent variable when studying its impact on experimental outcomes, this application proposes a scheme to eliminate covariate differences between samples within different groups. The experimental group and control group, specified by any independent variable, are designated as the first group and the second group, respectively. Then, using two different methods, the first candidate sample set and the second candidate sample set corresponding to each first sample in the first group are determined from each second sample within the second group. This approach considers both single-dimensional and multi-dimensional factors to comprehensively determine the matching sample for each first sample, ensuring the accuracy of sample matching within different groups specified by any independent variable. This ensures that the matched samples within different groups maintain a high degree of consistency in various covariates, thereby eliminating covariate differences between samples within different groups.
[0031] Figure 1 This is a flowchart illustrating a sample matching method according to an embodiment of this application. (Refer to...) Figure 1 The method may include the following steps:
[0032] S110, determine the first and second groups in the experimental and control groups specified by any independent variable.
[0033] When analyzing the impact of using or not using any independent variable in an experiment, the corresponding experimental group and control group are usually obtained based on whether each experimental subject uses the independent variable.
[0034] For any independent variable, the experimental group is a group consisting of experimental subjects using the independent variable as samples, and the control group is another group consisting of experimental subjects not using the independent variable as samples.
[0035] Taking the patients' use of DRG tools for payment as an example, resulting in DRG groups and non-DRG groups, when analyzing the cost control effect of DRG, the DRG group can be used as the experimental group, while the non-DRG group can be used as the control group.
[0036] In this application, when conducting experimental analysis on the effect of any independent variable in its respective field, the corresponding experimental group and control group are usually selected from a large number of experimental subjects based on whether each experimental subject in the field uses the independent variable.
[0037] However, when selecting experimental and control groups, each sample in the experimental and control groups has some unique and uncontrollable basic attributes that cannot be modified by the experimenter. These sample characteristics exist as covariates in the experimental analysis, resulting in certain differences between the experimental and control groups in each covariate, which affects the accuracy of the experimental effect of the independent variable.
[0038] Therefore, in order to eliminate the covariate differences between the experimental and control groups, this application can perform feature matching on each sample in the experimental group and each sample in the control group to divide the matched samples in the experimental and control groups into two new groups. This ensures that the sample characteristics in the two matched groups remain basically consistent, thereby eliminating the covariate differences between the experimental and control groups. Accordingly, this application can designate one of the experimental and control groups as the first group and the other as the second group.
[0039] In order to facilitate accurate differentiation of samples in different groups, this application can use each sample in the first group as the first sample and each sample in the second group as the second sample, so that the matching sample of each first sample in the first group can be found sequentially from each second sample in the second group.
[0040] It should be understood that the characteristics of the first sample can be represented by the basic attributes of the first sample that are uncontrollable on each covariate. Similarly, the characteristics of the second sample can be represented by the basic attributes of the second sample that are uncontrollable on each covariate. Therefore, by subsequently performing sample feature matching on each first sample within the first group and each second sample within the second group, the matched first and second samples can eliminate the differences in their corresponding covariates.
[0041] Taking the first group as the DRG group and the second group as the non-DRG group as an example, the first sample in the DRG group can be each patient who paid using the DRG tool. The basic attributes of the first sample are the gender, age, length of hospital stay, and surgical level of each patient who paid using the DRG tool, used as covariates. Similarly, the second sample in the non-DRG group can be each patient who did not pay using the DRG tool. The basic attributes of the second sample are the gender, age, length of hospital stay, and surgical level of each patient who did not pay using the DRG tool, used as covariates.
[0042] As an optional implementation scheme in this application, considering that the sample sizes in the experimental group and the control group may be different, and that the first sample in the first group is used to actively match the second sample in the second group, and that the matching properties of the samples in the two groups are also different, this application first determines the sample sizes of the experimental group and the control group for each independent variable to ensure the comprehensiveness of the samples participating in the experimental analysis of any independent variable; the group with the larger sample size in the experimental group and the control group is designated as the first group, and the other group in the experimental group and the control group is designated as the second group.
[0043] In other words, the first group is the one with the larger sample size in both the experimental and control groups, while the second group is the one with the smaller sample size in both groups. Therefore, each first sample in the first group can obtain at least one matching sample from each second sample in the second group, ensuring that the number of matched samples remains consistent with the larger sample size in both the experimental and control groups, thus preventing sample omissions in subsequent experiments for both groups.
[0044] S120, based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group, determine the first candidate sample set corresponding to each first sample in the second group.
[0045] In this application, the first group consists of multiple first samples, and the second group consists of multiple second samples. The sample features of the first and second samples are multidimensional features composed of basic attribute information on each covariate.
[0046] Therefore, in order to reduce the complexity of matching the first sample and the second sample on each covariate, this application first performs a comprehensive analysis of the sample features of the first sample and the second sample on each covariate, so as to summarize the sample features under each covariate into a feature value, which helps to obtain the inherent implicit relationship between the first sample and the second sample on each covariate, and facilitates the subsequent matching of the first sample and the second sample.
[0047] In order to summarize the multidimensional feature values of the sample features of the first sample and the second sample on each covariate, considering that one of the first sample and the second sample will use the corresponding independent variable, while the other sample will not use the corresponding independent variable, this application can pre-train a neural network for classifying whether any sample uses the independent variable.
[0048] By sequentially inputting the sample features of each first sample within the first group into a trained neural network, the neural network sequentially performs corresponding classification processing on the sample features of each first sample, thereby sequentially outputting the first classification probability of each first sample.
[0049] Similarly, by sequentially inputting the sample features of each second sample within the second group into the trained neural network, the neural network sequentially performs corresponding classification processing on the sample features of each second sample, thereby sequentially outputting the second classification probability of each second sample.
[0050] Then, considering that the first group and the second group belong to different categories under whether or not the independent variable is used, the first category probability of the first sample and the second category probability of the second sample will tend to the standard value set under different categories. This standard value can include "1" when the independent variable is used and "0" when the independent variable is not used.
[0051] Therefore, by analyzing the classification deviation between the first classification probability of each first sample in the first group and the standard value of the category to which the first group belongs, and the classification deviation between the second classification probability of each second sample in the second group and the standard value of the category to which the second group belongs, we can more accurately uncover the impact of the sample features of each first sample on each covariate and the sample features of each second sample on each covariate on classification accuracy. Then, for each first sample, we can find multiple second samples with a high degree of similarity to the classification deviation of the first sample from the various second samples in the second group, forming the first candidate sample set for that first sample.
[0052] It should be understood that, following the same method described above, the first candidate sample set for each first sample can be determined from each second sample in the second group. Furthermore, the covariate differences between each second sample in the first candidate sample set of each first sample and that first sample are relatively low.
[0053] S130, based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, determine the second candidate sample set corresponding to each first sample from the second group.
[0054] To ensure the accuracy of sample matching between the experimental and control groups, in addition to using the classification method in S120 to perform matching analysis on each first sample in the first group and each second sample in the second group, this application will further analyze the degree of matching between each first sample in the first group and each second sample in the second group through feature similarity.
[0055] In this application, a corresponding clustering algorithm can be used to perform similarity analysis on the sample features of every two first samples within the first group, thereby clustering each first sample within the first group to obtain multiple first sub-clusters. Each first sub-cluster can consist of at least one first sample. Furthermore, the number of first sub-clusters can be the number of clustering categories in the first group. In this application, the number of clustering categories can be 5% of the number of samples in the first group, or it can be other values; this application does not impose any limitations on this.
[0056] Similarly, using a corresponding clustering algorithm, similarity analysis is performed on the sample features of every two second samples within the second group to cluster each second sample within the second group, thereby obtaining multiple second sub-clusters. Each second sub-cluster can consist of at least one second sample. Moreover, the number of second sub-clusters can be the number of clustering categories in the second group. In this application, the number of clustering categories can be 5% of the number of samples in the second group, or it can be other values, which are not limited in this application.
[0057] It should be understood that, since the number of samples in the first group is higher than the number of samples in the second group, the number of samples in the first sub-cluster is greater than the number of samples in the second sub-cluster.
[0058] Taking the K-means clustering algorithm as an example, when clustering samples in either the first or second group, the feature similarity (i.e., Euclidean distance) between any two samples within that group can be used as the distance between those two samples. Then, if, based on the distance between two samples, one sample is determined to be one of the k most similar samples of another, then these two samples can be connected using the distance between them as weights, thus obtaining the clustering topology for that group. Then, as... Figure 2 As shown, a corresponding graph partitioning algorithm can be used to partition the original clustering topology graph into independent subclusters.
[0059] After obtaining multiple first subclusters corresponding to the first group and multiple second subclusters corresponding to the second group, for each first subcluster, the sample features of the multiple first samples contained in that first subcluster can be comprehensively analyzed to obtain the features of that first subcluster. Similarly, the features of each second subcluster can be determined. Then, based on the feature similarity between each first subcluster and each second subcluster, the degree of matching between each first subcluster and each second subcluster can be determined.
[0060] Furthermore, for each first sample, by determining the second sub-cluster that best matches the first sub-cluster containing the first sample, a second candidate sample set corresponding to the first sample can be determined from the multiple second samples contained in the second sub-cluster. The second candidate sample set corresponding to each first sample can be all the second samples within the second sub-cluster that best matches the first sub-cluster containing the first sample, or it can be a subset of the second samples within the second sub-cluster.
[0061] It should be noted that there is no specific order of execution between S120 and S130 in this application, and they can be executed simultaneously.
[0062] S140, based on the first candidate sample set and the second candidate sample set corresponding to each first sample, determine the matching sample of each first sample to form the matching group of the first group.
[0063] After obtaining the first candidate sample set and the second candidate sample set corresponding to each first sample, it is possible to determine whether there is an overlapping second sample in the first candidate sample set and the second candidate sample set corresponding to each first sample. The overlapping second sample can be used as the matching sample of the first sample.
[0064] If there are multiple overlapping second samples in the first candidate sample set and the second candidate sample set corresponding to a certain first sample, then one of the overlapping second samples can be randomly selected as the matching sample of the first sample.
[0065] As an optional implementation scheme in this application, when determining the matching sample of each first sample based on the first candidate sample set and the second candidate sample set corresponding to each first sample, the statistical frequency of each second sample in the first candidate sample set and the second candidate sample set corresponding to the first sample can be determined for each first sample, and the second sample with the highest statistical frequency can be used as the matching sample of the first sample; the matching samples of each first sample in the first group are combined to obtain the matching group of the first group.
[0066] In other words, the second samples in the first and second candidate sample sets corresponding to each first sample can be merged, and the frequency of each second sample can be determined, which can be 1 or 2. Then, the second sample with the highest frequency, that is, the second sample with a frequency of 2, is found and used as the matching sample for the first sample. If there are multiple second samples with a frequency of 2, one of them is randomly selected as the matching sample for the first sample. If there is no second sample with a frequency of 2, one of the second samples is randomly selected from all the second samples as the matching sample for the first sample.
[0067] Furthermore, by combining the matching samples of each first sample, a matching group for the first group can be obtained. Then, the first group and this matching group are used as the new experimental group and control group, respectively. This ensures that the basic attribute information of each sample in the new experimental group and control group is more matched on each covariate, thereby eliminating the covariate differences between the experimental group and the control group. Using the new experimental group and control group, the effect of this independent variable in its respective field is experimentally analyzed to ensure the accuracy of the experimental effect under any independent variable.
[0068] The technical solution provided in this application first determines a first group and a second group in the experimental and control groups specified by any independent variable. Then, based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group, a first candidate sample set corresponding to each first sample is determined from the second group. Moreover, based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, a second candidate sample set corresponding to each first sample is determined from the second group. Furthermore, from the first candidate sample set and the second candidate sample set corresponding to each first sample, matching samples for each first sample are determined, thereby forming a matching group for the first group. This achieves accurate matching of samples within different groups specified by any independent variable, ensuring that the matched samples within different groups maintain basic consistency in various covariates. Thus, without reducing covariates and while ensuring the comprehensiveness of samples within different groups, the differences in covariates among samples within different groups can be eliminated, ensuring the accuracy of the experimental results under any independent variable.
[0069] As an optional implementation scheme in this application, in order to ensure the accuracy of sample matching within different groups, this application will provide a detailed description of the specific process of determining the first candidate sample set and the second candidate sample set corresponding to each first sample from each second sample within the second group.
[0070] Figure 3 A flowchart illustrating another sample matching method shown in an embodiment of this application is as follows: Figure 3 As shown, the method may include the following steps:
[0071] S310, determine the first and second groups in the experimental and control groups specified by any independent variable.
[0072] S320, each of the first and second samples is input into multiple pre-built first-class classification models to obtain multiple first-class prediction probabilities of the first sample and multiple second-class prediction probabilities of the second sample.
[0073] Considering that different classification models have different network parameters, the classification probabilities output for the same sample may also differ. Therefore, to ensure the accuracy of the classification probabilities of the first and second samples, this application pre-constructs multiple classification models, each capable of predicting the classification probability of any sample. For example, the classification models in this application can be Support Vector Machine (SVM) models, logistic regression models, decision tree models, random forest models, etc.
[0074] Furthermore, considering that multiple classification models will produce multiple predicted values when predicting the classification probability of the same sample, a third classification model is needed to combine these multiple predicted values into a final classification probability.
[0075] Therefore, this application may randomly select one of the pre-built classification models as the second classification model in this application, and use the remaining multiple classification models as the first classification model in this application.
[0076] Furthermore, such as Figure 4 As shown, for each first sample within the first group and each second sample within the second group, the sample features of that sample can be input into multiple first-class classification models. Each first-class classification model then predicts the classification probability of that sample, resulting in multiple predicted classification probabilities. In other words, through multiple first-class classification models, multiple first-class predicted probabilities can be obtained for each first sample, and multiple second-class predicted probabilities can be obtained for each second sample.
[0077] S330, input the multiple first classification prediction probabilities of each first sample into the pre-built second classification model to obtain the first classification probability of the first sample.
[0078] For each first sample, the multiple first-class predicted probabilities of the first sample can be input into the constructed second-class classification model. The second-class classification model can then summarize the multiple first-class predicted probabilities of the same first sample to output the first-class probability of the first sample.
[0079] S340, input the multiple second-class prediction probabilities of each second sample into the second-class classification model to obtain the second-class probability of the second sample.
[0080] Similarly, for each second sample, multiple second-class prediction probabilities of the second sample can be input into the constructed second-class classification model. The second-class classification model can then aggregate the multiple second-class prediction probabilities of the same second sample to output the second-class probability of the second sample.
[0081] S350, for each first sample, based on the predetermined upper limit of classification probability and the absolute deviation between the first classification probability of the first sample and the sum of the second classification probabilities of each second sample in the second group, determine the first candidate sample set corresponding to the first sample from the second group.
[0082] Considering that the first group and the second group belong to different categories under whether or not the independent variable is used, the first category probability of the first sample and the second category probability of the second sample will tend to the standard values set under different categories. If the sum of the first category probability of a certain first sample and the second category probability of a certain second sample is closest to 1, it means that the sample feature differences between the first sample and the second sample on each covariate are minimal, which can eliminate the covariate differences between the first sample and the second sample.
[0083] Therefore, this application can set a classification probability upper limit of 1. Then, for each first sample, the first classification probability of the first sample and the second classification probability of each second sample are added together to obtain the probability sum between the first sample and each second sample. Then, the absolute deviation between each probability sum under the first sample and the classification probability upper limit "1" is determined. At this point, the smaller the absolute deviation, the smaller the covariate difference between the first sample and the second sample represented by the absolute deviation. Therefore, this application can select from the second samples in the second group those whose calculated absolute deviation is less than or equal to a certain preset threshold (e.g., 0.02) to form the first candidate sample set corresponding to the first sample.
[0084] Following the same method described above, the first candidate sample set corresponding to each first sample can be obtained.
[0085] S360, cluster the first sample in the first group and the second sample in the second group respectively to obtain multiple first sub-clusters corresponding to the first group and multiple second sub-clusters corresponding to the second group.
[0086] This application can use any clustering algorithm to cluster each first sample in the first group and each second sample in the second group, thereby obtaining multiple first sub-clusters corresponding to the first group and multiple second sub-clusters corresponding to the second group.
[0087] The first sub-cluster consists of at least one first sample, and the second sub-cluster consists of at least one second sample.
[0088] S370, for each first sub-cluster, determine the matching sub-cluster of the first sub-cluster from the second sub-cluster based on the similarity between the first sub-cluster and each second sub-cluster.
[0089] After obtaining the first and second subclusters, in order to ensure a proper match between the first and second samples, it is necessary to first match the first and second subclusters. Therefore, for each subcluster in the first and second subclusters, the samples within that subcluster are first identified, and the sample features of each sample within that subcluster are comprehensively processed to obtain the features of that subcluster.
[0090] Then, by analyzing the feature similarity between each first sub-cluster and each second sub-cluster, the matching sub-cluster of each first sub-cluster can be determined from each second sub-cluster.
[0091] As an optional implementation of this application, the matching between the first sub-cluster and the second sub-cluster can be determined by the following steps:
[0092] The first step is to determine the first tuple of the first sub-cluster based on the sample characteristics of the first sample within each first sub-cluster.
[0093] For each first sub-cluster, there will be multiple first samples. By comprehensively processing the sample features of each first sample in the first sub-cluster, we can obtain the sample number, linear sum of multidimensional feature values, sample feature mean, variance, standard deviation, etc., of the first sample in the first sub-cluster. Based on this, the first tuple of the first sub-cluster can be generated.
[0094] Taking the first tuple as a quadruple as an example, the first tuple of a certain first sub-cluster can be represented as CL = (n, w1, w2, w3). Where, n is the number of samples of the first sample in the first sub-cluster, w1 is the linear sum of the multidimensional feature values of the n first samples in the first sub-cluster, w2 is the mean of the sample features of the n first samples in the first sub-cluster, and w3 is the standard deviation of the sample features of the n first samples in the first sub-cluster.
[0095] The second step is to determine the second tuple of each second sub-cluster based on the sample characteristics of the second sample within each second sub-cluster.
[0096] Similarly, following the method for determining the first tuple of the first sub-cluster, each second sub-cluster contains multiple second samples. By comprehensively processing the sample features of each second sample within the second sub-cluster, the sample number, linear sum of multidimensional feature values, sample feature mean, variance, standard deviation, etc., of the second sample within the second sub-cluster can be obtained. Based on this, the second tuple of the second sub-cluster can be generated.
[0097] The third step is to determine the matching sub-cluster of the first sub-cluster from the second sub-cluster based on the similarity between the first tuple of the first sub-cluster and the second tuple of each second sub-cluster.
[0098] For any subcluster within the first tuple of each first subcluster and the second tuple of each second subcluster, the sum of the terms of the tuple for that subcluster can be calculated. Then, for each first subcluster, the similarity between the first tuple of the first subcluster and the second tuple of each second subcluster can be analyzed based on the ratio of the difference between the sum of the terms of the first tuple of the first subcluster and the sum of the terms of the second tuple of each second subcluster. Finally, the second subcluster whose sum of terms of the first tuple of the first subcluster is closest to that of the first subcluster is taken as the matching subcluster of the first subcluster.
[0099] For example, the matching degree between each first sub-cluster and each second sub-cluster can be determined using the following formula:
[0100] Where r is the ratio of the difference between the sum of the terms of the first tuple of each first subcluster and the sum of the terms of the second tuple of each second subcluster, and m is the number of first subclusters. CL 1i CL is the sum of the terms of the first tuple of the i-th first subcluster within the first group. 1a CL is the sum of the terms of the first tuple of the a-th first subcluster within the first group. 2b It is the sum of the terms of the second tuple of the b-th second subcluster within the second group.
[0101] According to the above formula, for each first sub-cluster, the sum CL of the terms of the second tuple of each second sub-cluster is changed. 2b This allows us to determine the ratio of differences between the first tuple of the first subcluster and the second tuples of each second subcluster. Then, the second subcluster with the r value closest to 0 is determined as the matching subcluster of the first subcluster.
[0102] S380, for each first sample, determine the second candidate sample set corresponding to the first sample from the matching sub-cluster of the first sub-cluster where the first sample is located.
[0103] For each first sample, the first sub-cluster to which the first sample belongs can be determined, and then the matching sub-cluster of the first sub-cluster can be determined. Then, the similarity between the sample features of the first sample and the sample features of each second sample in the matching sub-cluster is calculated, so as to determine the k most similar second samples from each second sample in the matching sub-cluster, as the second candidate sample set corresponding to the first sample.
[0104] For example, this application can employ the reverse approach of a corresponding graph segmentation algorithm to merge each first sub-cluster and its matching sub-clusters, such that the merged sub-cluster contains both first and second samples. Then, the merged sub-cluster containing each first sample is determined, and the k most similar second samples to the first sample are identified from each second sample in the merged sub-cluster, serving as the second candidate sample set corresponding to the first sample.
[0105] S390, based on the first candidate sample set and the second candidate sample set corresponding to each first sample, determine the matching sample of each first sample to form the matching group of the first group.
[0106] The technical solution provided in this application, through multiple classification models, ensures the accuracy of the first classification probability of the first sample and the second classification probability of the second sample. Furthermore, by representing the features of the first and second subclusters using tuples, it ensures convenient and efficient sample matching.
[0107] Figure 5 This is a schematic block diagram illustrating a sample matching device according to an embodiment of this application. Figure 5 As shown, the device 500 may include:
[0108] Grouping determination module 510 is used to determine the first and second groups in the experimental and control groups specified by any independent variable;
[0109] The first candidate sample determination module 520 is used to determine the first candidate sample set corresponding to each first sample from the second group based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group;
[0110] The second candidate sample determination module 530 is used to determine the second candidate sample set corresponding to each first sample from the second group based on the similarity between the first sub-cluster after sample clustering of the first group and the second sub-cluster after sample clustering of the second group.
[0111] The sample matching module 540 is used to determine the matching sample of each first sample based on the first candidate sample set and the second candidate sample set corresponding to each first sample, so as to form the matching group of the first group.
[0112] In some implementations, the first candidate sample determination module 520 can be specifically used for:
[0113] Each sample in the first sample and the second sample is input into multiple pre-built first-class classification models to obtain multiple first-class prediction probabilities of the first sample and multiple second-class prediction probabilities of the second sample.
[0114] The multiple first-class prediction probabilities of each first sample are input into the pre-built second-class classification model to obtain the first-class probability of the first sample.
[0115] The multiple second-class prediction probabilities of each second sample are input into the second-class classification model to obtain the second-class probability of the second sample.
[0116] For each first sample, the first candidate sample set corresponding to the first sample is determined from the second group based on the predetermined upper limit of classification probability and the absolute deviation between the first classification probability of the first sample and the sum of the second classification probabilities of each second sample in the second group.
[0117] In some implementations, the second candidate sample determination module 530 may include:
[0118] A clustering unit is used to cluster the first sample in the first group and the second sample in the second group respectively to obtain a plurality of first sub-clusters corresponding to the first group and a plurality of second sub-clusters corresponding to the second group. The first sub-cluster consists of at least one first sample, and the second sub-cluster consists of at least one second sample.
[0119] A sub-cluster matching unit is used to determine, for each first sub-cluster, a matching sub-cluster from the second sub-cluster based on the similarity between the first sub-cluster and each second sub-cluster;
[0120] The candidate sample determination unit is used to determine, for each first sample, the second candidate sample set corresponding to the first sample from the matching sub-cluster of the first sub-cluster where the first sample is located.
[0121] In some implementations, the sub-cluster matching unit can be specifically used for:
[0122] Based on the sample characteristics of the first sample within each first sub-cluster, determine the first tuple of that first sub-cluster;
[0123] Based on the sample characteristics of the second sample within each second sub-cluster, determine the second tuple of that second sub-cluster;
[0124] For each first sub-cluster, a matching sub-cluster of the first sub-cluster is determined from the second sub-cluster based on the similarity between the first tuple of the first sub-cluster and the second tuple of each second sub-cluster.
[0125] In some implementations, the sample matching module 540 can be specifically used for:
[0126] For each first sample, determine the statistical frequency of each second sample in the first candidate sample set and the second candidate sample set corresponding to the first sample, and take the second sample with the highest statistical frequency as the matching sample of the first sample;
[0127] The matching samples of each first sample in the first group are combined to obtain the matching group of the first group.
[0128] In some implementations, the grouping determination module 510 can be specifically used for:
[0129] Determine the sample size for the experimental and control groups specified for each independent variable;
[0130] The group with the larger sample size in the experimental group and the control group is designated as the first group, and the other group in the experimental group and the control group is designated as the second group.
[0131] In this embodiment, firstly, a first group and a second group are determined within the experimental and control groups specified by any independent variable. Then, based on the first classification probability of the first sample within the first group and the second classification probability of the second sample within the second group, a first candidate sample set corresponding to each first sample is determined from the second group. Furthermore, based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, a second candidate sample set corresponding to each first sample is determined from the second group. Subsequently, matching samples for each first sample are determined from the first and second candidate sample sets corresponding to each first sample, thus forming a matching group for the first group. This achieves accurate matching of samples within different groups specified by any independent variable, ensuring that the matched samples within different groups maintain basic consistency across various covariates. Therefore, without reducing covariates and while ensuring the comprehensiveness of samples within different groups, the differences in covariates among samples within different groups can be eliminated, ensuring the accuracy of the experimental results under any independent variable.
[0132] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 5The apparatus 500 shown can execute any of the method embodiments in this application, and the foregoing and other operations and / or functions of each module in the apparatus 500 are respectively for implementing the corresponding processes in the various methods in the embodiments of this application. For the sake of brevity, they will not be described in detail here.
[0133] The apparatus 500 of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0134] Figure 6 This is a schematic block diagram of an electronic device shown in an embodiment of this application.
[0135] like Figure 6 As shown, the electronic device 600 may include:
[0136] The system includes a memory 610 and a processor 620. The memory 610 stores computer programs and transfers the program code to the processor 620. In other words, the processor 620 can retrieve and run the computer program from the memory 610 to implement the methods described in the embodiments of this application.
[0137] For example, the processor 620 can be used to execute the above-described method embodiments according to instructions in the computer program.
[0138] In some embodiments of this application, the processor 620 may include, but is not limited to:
[0139] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0140] In some embodiments of this application, the memory 610 includes, but is not limited to:
[0141] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0142] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 610 and executed by the processor 620 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0143] like Figure 6 As shown, the electronic device may also include:
[0144] Transceiver 630, which can be connected to processor 620 or memory 610.
[0145] The processor 620 can control the transceiver 630 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 630 may include a transmitter and a receiver. The transceiver 630 may further include antennas, and the number of antennas may be one or more.
[0146] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0147] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0148] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0149] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0151] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0152] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A sample matching method for analyzing the cost control effect of disease diagnosis-related grouping in the medical field, characterized in that, include: Determine the first and second groups in the experimental and control groups specified by any independent variable; The first group is the DRG group, which consists of patients who use the DRG tool for payment, and the second group is the non-DRG group, which consists of patients who do not use the DRG tool for payment. The first group and the second group respectively contain basic attribute information as covariates for each patient, including gender, age, length of hospital stay, and surgical level; Based on the first classification probability of the first sample within the first group and the second classification probability of the second sample within the second group, a first candidate sample set corresponding to each first sample is determined from the second group, including: Each sample in the first sample and the second sample is input into multiple pre-built first-class classification models to obtain multiple first-class prediction probabilities of the first sample and multiple second-class prediction probabilities of the second sample. The multiple first-class prediction probabilities of each first sample are input into the pre-built second-class classification model to obtain the first-class probability of the first sample. The multiple second-class prediction probabilities of each second sample are input into the second-class classification model to obtain the second-class probability of the second sample. For each first sample, based on the predetermined upper limit of classification probability and the absolute deviation between the first classification probability of the first sample and the sum of the second classification probabilities of each second sample in the second group, the first candidate sample set corresponding to the first sample is determined from the second group; Based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering, a second candidate sample set corresponding to each first sample is determined from the second group; Based on the first candidate sample set and the second candidate sample set corresponding to each first sample, a matching sample for each first sample is determined to form a matching group for the first group, including: For each first sample, determine the statistical frequency of each second sample in the first candidate sample set and the second candidate sample set corresponding to the first sample, and take the second sample with the highest statistical frequency as the matching sample of the first sample; The matching samples of each first sample in the first group are combined to obtain the matching group of the first group.
2. The method according to claim 1, characterized in that, The step of determining the second candidate sample set corresponding to each first sample from the second group based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering includes: Clustering is performed on the first sample in the first group and the second sample in the second group to obtain multiple first sub-clusters corresponding to the first group and multiple second sub-clusters corresponding to the second group. The first sub-cluster consists of at least one first sample, and the second sub-cluster consists of at least one second sample. For each first sub-cluster, based on the similarity between the first sub-cluster and each second sub-cluster, a matching sub-cluster of the first sub-cluster is determined from the second sub-cluster; For each first sample, determine the second candidate sample set corresponding to the first sample from the matching sub-cluster of the first sub-cluster where the first sample is located.
3. The method according to claim 2, characterized in that, For each first sub-cluster, based on the similarity between the first sub-cluster and each second sub-cluster, a matching sub-cluster of the first sub-cluster is determined from the second sub-cluster, including: Based on the sample characteristics of the first sample within each first sub-cluster, determine the first tuple of that first sub-cluster; Based on the sample characteristics of the second sample within each second sub-cluster, determine the second tuple of that second sub-cluster; For each first sub-cluster, a matching sub-cluster of the first sub-cluster is determined from the second sub-cluster based on the similarity between the first tuple of the first sub-cluster and the second tuple of each second sub-cluster.
4. The method according to claim 1, characterized in that, The determination of the first and second groups in the experimental and control groups specified by any independent variable includes: Determine the sample size for the experimental and control groups specified for each independent variable; The group with the larger sample size in the experimental group and the control group is designated as the first group, and the other group in the experimental group and the control group is designated as the second group.
5. A sample matching device for analyzing the cost control effect of disease diagnosis-related grouping in the medical field, characterized in that, include: The grouping determination module is used to determine the first and second groups in the experimental and control groups for any given independent variable; The first group is the DRG group, which consists of patients who use the DRG tool for payment, and the second group is the non-DRG group, which consists of patients who do not use the DRG tool for payment. The first group and the second group respectively contain basic attribute information as covariates for each patient, including gender, age, length of hospital stay, and surgical level; The first candidate sample determination module is used to determine a first candidate sample set corresponding to each first sample from the second group based on the first classification probability of the first sample in the first group and the second classification probability of the second sample in the second group, including: Each sample in the first sample and the second sample is input into multiple pre-built first-class classification models to obtain multiple first-class prediction probabilities of the first sample and multiple second-class prediction probabilities of the second sample. The multiple first-class prediction probabilities of each first sample are input into the pre-built second-class classification model to obtain the first-class probability of the first sample. The multiple second-class prediction probabilities of each second sample are input into the second-class classification model to obtain the second-class probability of the second sample. For each first sample, based on the predetermined upper limit of classification probability and the absolute deviation between the first classification probability of the first sample and the sum of the second classification probabilities of each second sample in the second group, the first candidate sample set corresponding to the first sample is determined from the second group; The second candidate sample determination module is used to determine the second candidate sample set corresponding to each first sample from the second group based on the similarity between the first sub-cluster of the first group after sample clustering and the second sub-cluster of the second group after sample clustering. The sample matching module is used to determine the matching samples for each first sample based on the first candidate sample set and the second candidate sample set corresponding to each first sample, so as to form a matching group for the first group, including: For each first sample, determine the statistical frequency of each second sample in the first candidate sample set and the second candidate sample set corresponding to the first sample, and take the second sample with the highest statistical frequency as the matching sample of the first sample; The matching samples of each first sample in the first group are combined to obtain the matching group of the first group.
6. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the sample matching method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the sample matching method as described in any one of claims 1-4.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the sample matching method as described in any one of claims 1-4.