Method, electronic device and computer program product for evaluating an operation result

By evaluating operational results using a hierarchical hybrid expert network model (HME), the applicability issues of multivariate and continuous processing in existing technologies are resolved, achieving automated and accurate operational result evaluation while reducing hyperparameter dependence and subjectivity.

CN113469203BActive Publication Date: 2026-02-03NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010245570.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-31
Publication Date
2026-02-03
Estimated Expiration
2040-03-31

AI Technical Summary

Technical Problem

Existing subgroup analysis methods are not applicable to multivariate and continuous processing scenarios, and are subject to subjectivity and hyperparameter dependence, leading to inaccurate analysis results.

Method used

A hierarchical hybrid expert network (HME) model is used to evaluate operational results. By establishing an initial prediction model and optimizing the parameters of the gate nodes and expert nodes using observation data, a final prediction model is generated. This model is suitable for multivariate and continuous processing and automates the evaluation of operational results.

Benefits of technology

It enables automated and accurate evaluation of operation results without the need for hyperparameter tuning, and is suitable for diverse and continuous processing scenarios, reducing subjectivity and uncertainty of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113469203B_ABST
    Figure CN113469203B_ABST
Patent Text Reader

Abstract

A method, an electronic device and a computer readable program medium for evaluating operation results are disclosed. In one method for evaluating operation results, an initial prediction model is established for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes a gating node for grouping individuals and a plurality of different expert nodes for performing prediction based on different prediction methods, the observation data including individual features of individuals, corresponding operations performed on the individuals and corresponding operation results; parameters of the gating node and the expert nodes are determined using the observation data to obtain a final prediction model; and operation results of a predetermined operation on a sub-group of individuals in the observation data matching the predetermined operation are evaluated using each expert node in the final prediction model. With the embodiments of the present disclosure, precise operation result evaluation can be achieved in an automated manner without adjusting hyperparameters, and it is suitable for a variety of data types, with a more widely applicable scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data mining technology, and more particularly to a method, electronic device, and computer program product for evaluating operational results. Background Technology

[0002] With the rapid development of information technology, the scale of data has grown exponentially. This data contains a vast amount of useful information, thus data mining has received increasing attention. For example, causal inference is widely used in various fields, such as healthcare, education, and ecology, to extract valuable information from the data. In these fields, it is often difficult to find solutions that are effective for all individuals. For instance, for cancer patients, different treatment plans have different effects on different patients; for students receiving education and training, different training programs also have different effects on different students.

[0003] One solution to the above problem is to identify subgroups with better effects defined by individual characteristics for each treatment. This problem of finding subgroups with different treatment effects is called subgroup analysis. Subgroup analysis helps to explore, for example, the heterogeneity of treatment effects.

[0004] Subgroup analysis methods can be broadly categorized into two types: confirmatory subgroup analysis and exploratory subgroup analysis. Confirmatory subgroup analysis is primarily used to process a small number of predefined subgroups, while exploratory subgroup analysis uses a data-driven approach to identify subgroups with different therapeutic effects. In confirmatory subgroup analysis, the subgroups are predefined by professionals, which introduces significant subjectivity. This subjectivity can directly lead to questionable results and the possibility of deliberate manipulation of the analysis results. Exploratory subgroup analysis employs a tree-based approach, a widely recognized technique for identifying heterogeneity. It can automatically identify subgroups without prior knowledge and is suitable for large datasets.

[0005] However, both confirmatory and exploratory subgroup analyses may only handle binary treatments and suffer from inaccurate results. Therefore, an improved technique for evaluating operational outcomes is needed. Summary of the Invention

[0006] In view of this, this disclosure provides a method, electronic device, and computer program product for evaluating operational results.

[0007] According to a first aspect of this disclosure, a method for evaluating the results of an operation is provided. The method may include establishing an initial prediction model for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods, wherein the observation data includes multiple individual features of multiple individuals, corresponding operations in the predetermined operations performed on the multiple individuals, and corresponding operation results. The method further includes determining the parameters of the gating nodes and expert nodes of the initial prediction model using the observation data to obtain a final prediction model. The method further includes: using each expert node in the final prediction model to predict individual subgroups in the observation data that match it, to determine the operation results of the predetermined operations for each individual subgroup.

[0008] According to a second aspect of this disclosure, another method for evaluating operational outcomes is provided. The method may include receiving individual characteristics of one or more individuals and a corresponding operational mode in a predetermined operation performed on the one or more individuals. The method further includes determining one or more subgroups to which the one or more individuals belong, based on a prediction model and the individual characteristics of the one or more individuals. The prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods. The method further includes predicting corresponding operational outcomes for the one or more individuals using the expert nodes in the prediction model associated with the determined one or more subgroups, according to the corresponding operational mode.

[0009] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor. The actions include: establishing an initial prediction model for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods, wherein the observation data includes multiple individual features of multiple individuals, corresponding operations in the predetermined operations performed on the multiple individuals, and corresponding operation results; determining parameters of the gating nodes and expert nodes of the initial prediction model using the observation data to obtain a final prediction model; and using each expert node in the final prediction model to predict individual subgroups in the observation data that match them, to determine the operation result of the predetermined operations for each individual subgroup.

[0010] According to a fourth aspect of this disclosure, another electronic device is also provided. This electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the device to perform an action. The action includes: receiving individual characteristics of one or more individuals and a corresponding operation mode in a predetermined operation to be performed on the one or more individuals; determining one or more subgroups to which the one or more individuals belong, based on a prediction model and the individual characteristics of the one or more individuals, the prediction model having a hierarchical structure and including gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods; and predicting corresponding operation results for the one or more individuals using the expert nodes in the prediction model associated with the determined one or more subgroups, according to the corresponding operation mode.

[0011] In a fifth aspect of this disclosure, a computer-readable medium is provided that stores machine-executable instructions thereon, which, when executed, cause a machine to perform the method according to the first aspect.

[0012] In a sixth aspect of this disclosure, a computer-readable medium is provided having machine-executable instructions stored thereon, which, when executed, cause a machine to perform the method according to the second aspect.

[0013] In a seventh aspect of this disclosure, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform the method according to the first aspect.

[0014] In an eighth aspect of this disclosure, a computer program product is provided, which is tangibly stored on a computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform the method according to the second aspect.

[0015] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of this disclosure, nor is it intended to limit the scope of this disclosure. Attached Figure Description

[0016] The above and other features of this disclosure will become more apparent from the detailed description of the embodiments shown in conjunction with the accompanying drawings, in which the same reference numerals denote the same or similar parts. In the drawings:

[0017] Figure 1 A flowchart illustrating a method for evaluating operational results according to one embodiment of the present disclosure is shown schematically;

[0018] Figure 2 A schematic diagram illustrating the structure of an initial prediction model according to one embodiment of the present disclosure is shown.

[0019] Figure 3 A flowchart illustrating a method for optimizing an initial prediction model based on an optimization objective function according to one embodiment of the present disclosure is shown schematically.

[0020] Figure 4 A flowchart illustrating a method for evaluating the operational results of a predetermined operation for one or more individuals, according to one embodiment of the present disclosure, is shown schematically.

[0021] Figure 5 A block diagram of a system for evaluating operational results according to one embodiment of the present disclosure is shown schematically.

[0022] Figure 6 An example application of evaluating the results of a predetermined operation on an individual subgroup according to a specific implementation of this disclosure is illustrated schematically;

[0023] Figure 7 A schematic diagram of an electronic device in which embodiments of the present disclosure can be implemented is shown. Detailed Implementation

[0024] In the following, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that these drawings and descriptions relate only to preferred embodiments as examples. It should be pointed out that, based on the following description, alternative embodiments of the structures and methods disclosed herein can be readily conceived, and these alternative embodiments can be used without departing from the principles of the disclosure claimed herein.

[0025] It should be understood that these exemplary embodiments are provided merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Furthermore, in the accompanying drawings, optional steps, modules, etc., are shown with dashed boxes for illustrative purposes.

[0026] The terms “comprising,” “including,” and similar terms used herein should be understood as open-ended terms, meaning “including / including but not limited to.” The term “based on” means “at least partially based on.” The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment.” Definitions of other terms will be given in the description below.

[0027] In embodiments of this disclosure, the term "disposal variable / mode" refers to the manner in which a predetermined action is performed on each individual being studied in an observational experiment. This can be represented by a disposal level, such as taking a medication or a placebo, various dosages of a repeated medication, or potential strategies like a course program a student attends. Here, T is used to represent the disposal mode, which can also be called a "disposal variable." T takes the value t, which can be discrete or continuous data. For example, for a binary treatment with two disposal levels, t = 0, 1; for a ternary treatment with three disposal levels, t = 0, 1, 2; and so on for treatments with more disposal levels. When the disposal levels are continuous, t can also be any value within a range.

[0028] In embodiments of this disclosure, the term "individual characteristic" refers to the characteristics of an individual to be observed, also known as an individual attribute. For example, in the case of a patient, individual characteristics may include the patient's weight, age, gender, duration of illness, relevant medical examination data, etc. In the field of education, for example, it may be the student's age, gender, current level (e.g., English proficiency level), whether they have participated in similar courses, family economic situation, etc. In this document, X represents an individual characteristic; furthermore, an individual characteristic may also be referred to as a "covariate".

[0029] In embodiments of this disclosure, the term "operational outcome" refers to the result obtained or predicted / evaluated when a predetermined operation is performed on an individual; it can also be called a "treatment outcome." For example, in the case of a patient, the operational outcome may refer to the progression of the individual's condition after the administration of a medication or placebo, such as the reduction of symptoms, or a test value measuring the effect. In the field of education and healthcare, it could be the student's learning outcome, such as an improvement in relevant levels (e.g., an improvement in an English proficiency test level). In this document, Y is used to represent the treatment outcome, which can also be referred to as an "outcome variable."

[0030] In embodiments of this disclosure, "observational data" refers to data obtained from specific observational experiments or data accumulated in practical applications. For example, it could be operational result data based on pre-performed predetermined operations on individuals, or, in the medical field, data obtained from observational experiments on different patients. N can indicate the number of observational samples, i.e., how many data points are related to the individual; D is the dimension of the observational variables, or the number of observational variables, i.e., each data point contains multiple related parameter values, which is the sum of the number of covariates X, treatment variables T, and outcome variables.

[0031] In embodiments of this disclosure, the term "latent variable" refers to a variable that is not observed in causal inference but needs to be understood in causal inference. It is specifically introduced to solve the causal inference problem in this embodiment by optimizing the model.

[0032] As mentioned earlier, in many practical applications, there is a desire to predict the effect of an action on one or more subgroups, or to predict the appropriate course of action for an individual, so that computing devices can make decisions automatically or assist people in making decisions. This allows for the automatic determination of whether to perform a certain action on an individual, or which of several actions to perform on an individual. For example, it might be desirable to predict the potential impact of a drug or treatment on a patient's condition, thereby automatically or assisting doctors in developing treatment plans. It might also be desirable to predict the extent to which a training course will improve a student's grades, or to predict the impact of advertising on consumers' final purchasing behavior. Furthermore, in environmental protection and governance applications, it's possible to predict the effect of an environmental protection or governance plan on a particular region, or to automatically determine the most suitable governance plan for that region.

[0033] As mentioned earlier, existing subgroup analyses can be broadly categorized into confirmatory subgroup analyses and exploratory subgroup analyses. However, in confirmatory subgroup analyses, the number of subgroups to be analyzed and the method of subgroup division are predefined by investigators based on their experience, which is highly subjective. This subjectivity can directly lead to questionable results and raises the possibility of deliberate manipulation of the results.

[0034] In contrast, exploratory subgroup analysis is an automated process based on a tree structure. It is data-driven and does not require pre-defining the number of subgroups or the subgrouping method, thus possessing a high degree of objectivity. For illustrative purposes, an example of a prior art solution will be presented below.

[0035] In one existing subgroup analysis scheme, an initial tree structure is first grown based on the observed data. For each node, all possible partitioning methods are traversed, and the node is partitioned according to the partitioning statistics that maximize the partitioning. The same operation is performed on the left and right child nodes of this node. This generates an initial tree structure with the subgroups already partitioned. Subsequently, a pruning operation is performed on this initial tree structure, removing the weakest links to simplify the tree structure. Finally, the size of the tree structure is further reduced by calculating the partitioning statistics between each pair of terminal nodes. Node pairs with low heterogeneity are merged. This partitioning statistics calculation operation is repeated until superior heterogeneity is achieved in all remaining subgroups. The resulting tree structure is the result analysis model.

[0036] However, exploratory subgroup analysis is highly sensitive to hyperparameters, requiring manual adjustment to achieve optimal results; therefore, its performance is heavily dependent on hyperparameter settings. Furthermore, current techniques, both confirmatory and exploratory subgroup analyses, only support binary processing, making them unsuitable for scenarios requiring multivariate processing or for parameter-based continuous processing predictions. Therefore, an improved operational outcome evaluation technique is needed to at least partially address the aforementioned problems in existing technologies.

[0037] Therefore, this disclosure provides a novel scheme for evaluating operational structures. According to this scheme, a hierarchical multi-expert initial prediction model is first established for a set of observational data. Then, the parameters of the gating nodes and expert nodes of the initial prediction model are determined using the observational data to obtain a final prediction model. Subsequently, each expert node in the final prediction model can be used to predict the individual subgroups in the observational data that match it, thereby determining the operational results of a predetermined operation for each individual subgroup. Unlike existing subgroup analysis techniques, the implementation method of this disclosure is insensitive to hyperparameters, thus enabling fully automated operational result evaluation without the need for hyperparameter adjustment. Furthermore, in the implementation method of this disclosure, each expert model is suitable not only for binary processing applications but also for multivariate processing applications, and is also suitable for situations where the treatment level is continuous, thus having a wider range of application scenarios.

[0038] The technical solutions disclosed in this disclosure will be described in detail below with reference to the accompanying drawings and specific examples. However, it should be noted that the following description is for illustrative purposes only, and this patent is not limited thereto.

[0039] Figure 1 A schematic diagram of a flowchart illustrating a method for estimating operational effects according to one embodiment of the present disclosure is shown. The various steps in this method can be executed centrally by a single processing unit in an electronic device, or separately by multiple processing units in an electronic device, or executed in multiple processing units of multiple electronic devices, provided that data transmission between them is possible.

[0040] like Figure 1As shown, in block 110, an initial prediction model is first established for a set of observation data. This initial prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods. For example, the initial prediction model could be a hierarchical hybrid expert (HME) network model. The observation data includes result data that specifies corresponding operations for individuals, specifically including multiple individual characteristics of multiple individuals, corresponding operations in the predetermined operations performed on the multiple individuals, and corresponding operation results.

[0041] Observational data can be pre-stored in an observational database or imported into the system when operational results need to be evaluated. The observational data itself can originate from third parties or be collected through other means. For example, in the healthcare field, it could be data obtained from observational trials conducted by medical institutions and pharmaceutical research companies on multiple groups of patients / volunteers (e.g., drug reuse, placebo reuse, etc.). In the education and training sector, it could be data on the effectiveness of student education and training, accumulated over a long period of teaching experience.

[0042] The observation data can be an N*D matrix, where N is the number of observation samples (i.e., how many data points there are), and D is the dimension of the observed variables, or the number of observed variables, i.e., the number of related parameters in each data point. For example, in the healthcare field, N indicates the number of patients in the observation data, and D indicates the amount of data related to that patient, including the number of individual patient characteristics, corresponding operations, and corresponding results. This data can be preprocessed, such as through integration, reduction, and noise reduction of the raw data. These preprocessing operations are known in the field and will not be elaborated upon here.

[0043] In this paper, T represents the disposal variable in the observed data, i.e., the operation performed on an individual; X represents the observed covariate, i.e., the individual characteristics of the individual; and Y represents the corresponding operation result. The dimensions of T, X, and Y together define the dimension D of the observed data. For example, when X=5, T=1, and Y=1, the dimension D of the observed data is 7.

[0044] Based on observational data, a conditional distribution of the outcome Y can be established given the disposal variable T and covariate X. In this paper, a hierarchical network model with multiple experts is proposed for representation. For example, a hierarchical mixed expert (HME) model can be used. However, it should be noted that this model is given for illustrative purposes only, and this disclosure is not limited thereto.

[0045] The HME network is an extension of the hybrid expert model, represented by a tree structure that combines multiple different expert models according to a gating function. Specifically, it is a multi-level tree structure where the leaf nodes are expert nodes corresponding to the multiple expert models used for prediction based on different expert models, and the other nodes are gating nodes that indicate the grouping decision conditions used to group the observed data.

[0046] For the observed data, an initial HME network is first established. This initial HME network consists only of a basic multi-level tree structure, with leaf nodes being expert nodes. The number of levels and expert nodes in the HME network can be a predetermined value that meets the needs of the observed data. On the one hand, this value should be adapted to the number of covariates X, while on the other hand, it should provide a certain degree of design flexibility. For example, for observed data with 5 covariates, a 6-level HME network with 32 expert nodes can be generated. For fewer covariates (e.g., 3), a 5-level HME network with 16 expert nodes can be used, while for more covariates (e.g., 10), a 7-level HME network with 64 expert nodes can be used.

[0047] For illustrative purposes, in Figure 2 The diagram shows an exemplary HME model, which is a multi-layered model based on a tree structure, consisting of a 6-layer network structure and 32 expert nodes.

[0048] The initial HME network can be represented by a corresponding mathematical formula. However, in practical applications, the model is not necessarily represented mathematically in electronic devices. Instead, the HME network may be constructed by determining the parameters associated with the model. For example, the number of expert nodes and the number of gate nodes can indicate the structure of the initial HME network.

[0049] The construction and mathematical representation of the initial prediction model will be described in detail below with reference to specific examples of the HME model, and will not be repeated here.

[0050] Continue to refer to Figure 1 Next, in block 120, the parameters of the gating nodes and expert nodes of the initial prediction model are determined using the observation data to obtain the final prediction model. For example, the initial prediction model can be reduced, and the parameter values ​​of the remaining gating nodes and expert nodes can be determined. For example, it can be determined which individual feature in the individual feature X(1, ..., N) each gating node corresponds to, and what its split point value is; which expert nodes are used in the model, and what the parameters of the expert model corresponding to each expert node are.

[0051] The initial network model determined in Block 110 is merely an initial architecture; it lacks information about the individual characteristics corresponding to the gate nodes, as well as the splitting points (grouping points) for those individual characteristics. The parameters of the expert nodes in the model are also yet to be determined. In Block 120, observational data will be used to reduce the structure of the initial model and determine the parameters of the nodes within the model.

[0052] In one embodiment of this disclosure, an optimization objective function can be determined for the initial prediction model, and the initial prediction model can be optimized based on the observation data and the optimization objective function to determine the final prediction model.

[0053] For example, the optimization objective function for the initial prediction model can be determined based on the Decomposition Asymptotic Bayes (FAB) method. Optionally, latent variables can be introduced into the optimization objective function determined by the FAB method, indicating whether each observation in the observation data T matches any expert node in the prediction model. In this way, during the optimization process, not only can the parameter values ​​of each gate node and expert node be determined, but also the observation data that matches each expert can be identified.

[0054] For illustrative purposes, Figure 3 A flowchart is shown illustrating a method for optimizing an initial prediction model based on an optimization objective function, according to one embodiment of the present disclosure. Figure 3 In block 310, the latent variables are first optimized based on the objective function to determine the optimization probability of matching the observed data with each expert node. Next, in block 320, the existing prediction model is reduced based on the determined optimization probabilities to obtain a simplified prediction model. Then, in block 330, the gate nodes and expert nodes in the simplified prediction model can be optimized based on the objective function to determine the optimization parameter values ​​for each gate node and expert node, thereby generating an optimized prediction model.

[0055] It should be noted that the above steps can be repeated until the objective function converges, i.e., the difference between the results obtained from the two optimizations is within a predetermined threshold range. Alternatively, the above operations can be stopped after being repeated a predetermined number of times. This will yield the final prediction model that can be used for prediction.

[0056] The determination of the objective function and the optimization process will be described in detail below with examples, and will not be repeated here.

[0057] Next, in block 103, the individual subgroups that match the observation data T can be predicted using the expert nodes in the final prediction model to determine the operation results of the predetermined operation for each individual subgroup.

[0058] As mentioned earlier, the model structure reduction process is based on the probability of matching expert nodes with each observation data point. In other words, the final prediction model implicitly includes the matching relationships between the observation data and each expert node. In the final prediction network, the portion of observation data T that matches each expert node constitutes an individual subgroup. For each individual subgroup, the corresponding expert node can be used to predict the individual data points within it. Then, the evaluation values ​​of the operation results of each individual within the subgroup can be aggregated to serve as the evaluation value of the predetermined operation structure for that individual subgroup. For example, the average of the evaluation values ​​of the operation results of each individual can be used as the final evaluation value of the operation result for that subgroup. However, it should be noted that the average value is merely an example; other forms of values ​​can also be used as the final evaluation value of the operation result.

[0059] It is understood that, in the embodiments of this disclosure, the generated predictive model includes various expert models adapted to the corresponding groupings. These expert models are suitable not only for binary treatments but also for ternary treatments, and can also evaluate results for continuous treatments. For example, for ternary treatments, three operational outcome evaluations corresponding to three different treatment levels can be obtained, while for continuous treatment levels, a single operational evaluation result curve can be obtained.

[0060] The resulting expert model and / or the evaluation of the above-mentioned operational results can be transmitted to other processing modules in electronic devices or other processing modules in other electronic devices for use in other processes. Examples include adjusting drug formulations, automatically recommending suitable drug reuse suggestions, adjusting curriculum structure in educational settings, and automatically recommending suitable drug regimens.

[0061] The implementation method of this disclosure enables automatic grouping of individuals in observed data and allows for the evaluation of the operational results for each subgroup of individuals under predetermined operations. The entire process is data-driven and does not rely on any subjective perception. Furthermore, although hyperparameters exist in the parameter determination process of the prediction model of this disclosure, multiple observations have shown that the evaluation of operational results according to this disclosure is not sensitive to the values ​​of these hyperparameters. In other words, the evaluation of operational results does not depend on the hyperparameters, thus eliminating the need for hyperparameter adjustment. This ensures the results are accurate even when not affected by the values ​​of the hyperparameters. Simultaneously, in the case of an FAB-based optimization model, the L0 paradigm can be employed during the optimization process, which can further effectively alleviate the overfitting problem.

[0062] The following sections will describe the process of determining the objective function and optimizing the network using specific examples. However, it should be noted that this is merely an illustrative method for purposes of explanation, and this disclosure is not limited thereto.

[0063] Furthermore, it should be noted that the following description, for illustrative purposes, presents each step of the predictive model construction and objective function determination process. However, this is merely to illustrate the specific principles of this disclosure and does not imply that each step in the above process will be performed in the scheme provided in this disclosure. Rather, it is more likely that the expressions characterizing the HME network and / or the expressions for the optimization objective function have been pre-stored in an electronic device. When it is actually necessary to evaluate the observed data, it may only be necessary to determine the specific model and expressions to be used based on the pre-stored predictive model and optimization objective function according to the observed data. For example, the hierarchy of the predictive model to be used, the number of expert nodes, the number of gate nodes, and the various parameters to be used in the optimization objective function can be determined, and then these parameter values ​​can be assigned to the expressions below.

[0064] HME network model

[0065] The following section first defines some parameters that need to be used in the HME network to characterize it.

[0066] For an HME network, the number of gated nodes can be indicated by G, and the number of expert nodes can be indicated by E. For the g-th (g = 1, ..., G) gated node, a binary latent variable Ug ∈ {0, 1} can be defined. This latent variable Ug indicates whether an observation is associated with / matched to an expert node on the left branch of that gated node, i.e., whether the evaluation result of the operation on that data was generated by an expert node on the left branch of that gated node. If Ug = 1, the observation is associated with an expert node on the left branch of that gated node; and Ug = 0 indicates that the observation is not associated with an expert node on the left branch of that gated node, i.e., it is associated with an expert node on its right branch.

[0067] Let x represent the probability that the operation result is generated by an expert node on the left branch of the gate node for each individual feature in the observed data.

[0068]

[0069] Where θ g =(α g ,β g γ gThe gate node g is instructed to perform operations on the segmentation dimension γ. g With the predetermined dividing point β g The probability of performing a segmentation; γ g Indicates the segmentation dimension, that is, which covariate in covariate X is used for segmentation; β g Indicates the split point, i.e., at which value the split occurs; α g Indicates the probability of being segmented in this manner.

[0070] Based on this, the conditional distribution of the e-th expert can be represented as:

[0071]

[0072] in Wherein, the parameter τe indicates the average causal effect of the outcome Y of the action T = t in the expert e (e = 0, 1…E), and we and σe2 indicate the parameters used by the expert model. The expert model in this example is a linear prediction model, but this is merely exemplary and this disclosure is not intended to be limited thereto.

[0073] Furthermore, additional parameters can be defined to facilitate the subsequent representation of the HME model. Parameter G can be further defined. g and ε e G g Indicates the indices of all expert nodes located only in the subtree of the g-th gated node, where g = 1, ..., G; ε e This indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node, where e = 1; ...; E. Based on this, the following function can be further defined:

[0074]

[0075] Where h(ξ, g, e) is a function defined based on whether the e-th expert node is in the left / right subtree of the gate node g; ξ is a quantity with broad meaning, which can be a latent variable, other variables or functions, or a specific numerical value.

[0076] ξ is used to indicate the latent variables Ug and b(x, θ) respectively. g (Equation 1) gives us the following two functions:

[0077] h U (g, e) := h(U g , g, e)∈{0,1}, which are binary implicit variables;

[0078] h b (x, g, e) := h{b(x, θ) g ), g, e}, which is a probability function.

[0079] Where h U (g, e) indicates whether an observation is related to the e-th expert node under the g-th gated node branch;

[0080] h b (x, g, e) indicates the probability that an observation is associated with the e-th expert node under the branch of the g-th gated node.

[0081] Thus, the initial hybrid expert model can be expressed as:

[0082]

[0083] in

[0084] y indicates the result variable T;

[0085] x indicates the individual dimensions in the covariate X;

[0086] t = indicates the disposal variable;

[0087] θ = (θ1, ..., θ) G It also indicates the relevant parameters of each gate node in the HME model.

[0088] φ=(φ1,...,φ E ), and indicates the relevant parameters of each expert node in the HME model.

[0089] Introduction of latent variables

[0090] For the HME model described above, in order to perform subgroup analysis on the observed data, in addition to the observed variables mentioned above, it is also necessary to understand whether the observed data matches each expert node. Therefore, a binary hidden variable Z is further defined for the HME network. e The hidden variable Z e This indicates whether each observation in the observed data matches any of the expert nodes in the initial prediction model. Specifically, whether the operation of expert node e matches each observation can be represented as:

[0091]

[0092] in

[0093] e indicates the index of the transition node, e = 0, 1…E;

[0094] g indicates the index of the gated node, g = 0, 1, ... G;

[0095] ε e Indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node.

[0096] h U (g, e) indicates whether an observation is associated with the e-th expert node under the g-th gated node branch, where a value of 1 indicates that it is associated with the e-th expert node under the g-th gated node branch, and a value of 0 indicates that it is not associated with the e-th expert node.

[0097] If the result of an operation on an observation can match the expert node e, then Ze = 1; otherwise, Ze = 0. Z = Ze can be used to represent the information about the correlation and matching between the observation data and E expert nodes, where the latent variable is a component assignment vector. Thus, the observation data and the latent variable can be represented as follows:

[0098] Observational data:

[0099] Latent variables: Where z n =(z n1 , ..., z nE ),and

[0100] Based on the above observation data and latent variables, the HME network in Equation 4 above can be further represented as:

[0101]

[0102]

[0103] Next, an optimization objective function for the initial prediction model can be established based on the HME function that incorporates latent variables. In the following text, the FAB method will be used as an example to construct the optimization objective function.

[0104] Construction of the objective function

[0105] First, the parameters θ and φ required by the HME model in Equation 4 above can be represented by M. Then, based on Bayesian methods, the model that maximizes the following model posterior can be selected:

[0106] p(M|y N x N , t N )∝p(M|x N , t N )p(y N |x N , t N M)(Formula 7)

[0107] Using the uniform model prior p(M|x) N , t N Special attention needs to be paid to p(y).N |x N , t N M). If q(z) is used N Instruction z N The distribution of variables can then be obtained.

[0108]

[0109] And when q(z) N )=p(z N |y N x N , t N When M), the quality can be maintained. p(y) can be estimated by estimating the lower bound on the right-hand side of the above equation. N |x N , t N M).

[0110] According to the Bayesian method, p(y) N |x N , t N M) can be further expressed as

[0111]

[0112] in Other parameters are the same as above; and therefore the number of valid samples for this quantity is

[0113] Further use is possible D g D e To indicate Λ and θ respectively g and φ e The maximum likelihood estimation is then performed. The Laplace approximation method is then applied to the distributions of each decomposition. so An approximation can be made in the following way:

[0114]

[0115] Where [A, a] indicates the quadratic terms of matrix A and vector a. Aa, and

[0116]

[0117] In the above formula, Indicator Fisher Information Matrix The decomposition uses an approximation, where It can be represented as:

[0118]

[0119] In a similar manner It can be approximated as:

[0120]

[0121] in

[0122]

[0123] in Indicator Fisher Information Matrix The decomposition uses an approximation, where It can be represented as

[0124]

[0125] Therefore, logp(y) N , z N |x N , t N (θ, φ) can be approximated as:

[0126]

[0127] By substituting Equation 13 into Equation 9, we can obtain:

[0128]

[0129] Furthermore, priorsp(θ) can be further... g |M), p(φ e |M) is considered a constant, and note that:

[0130]

[0131]

[0132] Thus, p(y) N , z N |x N , t N M) can be further represented as:

[0133]

[0134] By further ignoring the asymptotic minimization term, we obtain the following optimization objective function FIC:

[0135]

[0136] in, as follows:

[0137]

[0138] In the above formula, In practice, this is difficult to obtain, making direct estimation of the FIC difficult. FAB inference, however, focuses on the asymptotically consistent lower bound of the objective function's FIC. For the Laplace function, Therefore, it is obvious Thus, according to From the definition, we can obtain the following formula.

[0139]

[0140] In this way, we can obtain FIC(y) N x N , t N The lower limit of M):

[0141]

[0142] in It is an arbitrary scalar. Thus, the initial prediction network can be optimized through the following maximization problem:

[0143]

[0144] If q, θ, and φ are fixed, then

[0145]

[0146] The optimization objective function can then be further simplified to:

[0147]

[0148] Where S = (S1, ..., S2) E ) is a vector of component function terms.

[0149] In this way, we can obtain an optimization objective function suitable for optimizing the HME model described above. Solving this objective function yields the optimized prediction network.

[0150] The process of solving the objective function optimization can include two steps. The first is to reduce the structure of the initial prediction network by optimizing the latent variables. The second is to optimize the parameters of the gate nodes and expert nodes based on the reduced prediction network structure. These two processes will be described below with examples.

[0151] Network structure reduction

[0152] In one embodiment of this disclosure, the network structure is simplified by optimizing the probability distribution of latent variables.

[0153] First, based on the above objective function, the optimization probability q, which characterizes the matching between expert nodes and observed data, can be expressed as:

[0154]

[0155] Here,

[0156]

[0157] By analyzing the above Relative to q ne Taking the derivative and setting the input = 0, we obtain the following expression:

[0158]

[0159] in

[0160]

[0161] This allows us to obtain the optimized probability q of matching expert nodes with observed data. Further, based on the optimized matching probability of expert nodes and observed data, experts and their corresponding branches with low matching probabilities (i.e., poor effectiveness) in the initial prediction model are removed. For example, the method used is shown in the following formula.

[0162]

[0163] δ indicates the threshold used to determine whether to remove an expert node; and Indicates the normalization constant, where

[0164] Therefore, by reducing or shrinking the model, unnecessary expert nodes and related branches can be removed from the network model. In this way, irrelevant expert nodes can be removed from the initial prediction model with a symmetric tree structure containing a large number of expert nodes.

[0165] Model node parameter determination

[0166] After removing irrelevant expert nodes, the parameters of the nodes in the resulting model will be optimized. In the proposed implementation, the component function term vector S and parameters θ and φ will be further optimized. Based on the above FIC objective, the optimization objective can be expressed as follows:

[0167]

[0168] In the above In this case, there are no overlapping terms between (S, φ) and θ, meaning there is no correlation between them. Therefore, we can optimize each of them separately, which will give us the following two optimization objective expressions.

[0169]

[0170]

[0171] Equation 24a can be further expressed as:

[0172]

[0173] That It is the γth g The value of the dimension is greater than or equal to β. g A collection of observational data; It is the γth g The value of the dimension is less than β g A collection of observational data; This indicates the indices of all expert nodes located only in the subtree to the left of the g-th gate node. Indicates the indexes of all expert nodes located only in the subtree to the right of the g-th gate node.

[0174] Thus, given The solution to the above formula is relative to α. g The analytical solution, i.e.

[0175]

[0176] In Equation 24b, given S e In this case, De instructed The dimension, by using φ e In the L0 paradigm, De can be represented as D e =||ω e ||0+2. In this way, the equation can be transformed into a feature selection problem with discrete constraints, which can be expressed as follows.

[0177]

[0178]

[0179]

[0180] Where C is the regularization term and The corresponding predetermined constant. In Equation 25 above, the objective function is based on φ. eThe smooth concave shape allows us to solve the L0-regularized feature selection problem by using a forward-backward (FoBa) greedy algorithm.

[0181] It can be seen that the L0 paradigm was used in the process of optimizing the parameters of the gated nodes and expert nodes, which can reduce the overfitting problem.

[0182] The steps of model structure reduction and model node parameter optimization described above can be performed iteratively until the function converges. That is, each iteration will further optimize based on the current prediction network. For example, the first optimization is based on the initial prediction network, while subsequent optimizations will use the previously optimized prediction network as a basis for further optimization. This iteration can be repeated until the difference between the prediction networks in two consecutive iterations is below a predetermined threshold. Alternatively, it can be specified that the iteration operation stops after a predetermined number of iterations, or a combination of both methods can be used.

[0183] Through several iterations, a final prediction model can be obtained, in which the individual features associated with each gate node and their segmentation points, as well as the expert model and its parameters to be used in the prediction model, have been determined.

[0184] It should be noted that while the entire process of establishing the HME network has been described in detail above, in practical applications, the objective function is not established step by step as described above. Instead, parameters M related to the HME network expression in Equation 4 may be determined, and then the prediction model is reduced based on Equations 21 and 22, for example, using these parameters and observation data. The optimized parameter values ​​of the gate nodes and expert nodes are then solved based on Equations 26 and 27, for example.

[0185] The foregoing described a technical solution for obtaining an optimized prediction network based on observational data and utilizing expert nodes within this network to evaluate operational outcomes for each subgroup. However, the technical solution disclosed herein can also be implemented in other ways, utilizing, for example, the final prediction model determined in step 120 to predict operational effects for individuals. Reference will be made below. Figure 4 Let me explain.

[0186] Figure 4 A flowchart illustrating a method for evaluating the operational results of a predetermined operation for one or more individuals, according to one embodiment of the present disclosure, is shown schematically. The various steps in this method can be performed centrally by a single processing unit in an electronic device, or separately by multiple processing units in an electronic device, or performed in multiple processing units of multiple electronic devices, provided that data transmission between them is possible.

[0187] like Figure 4As shown, in block 410, the individual characteristics of one or more individuals and the corresponding operation mode in a predetermined operation to be performed on the one or more individuals are received. The individuals here are, for example, patients, students, or other individuals, and the individual characteristics are individual attributes associated with that individual that may be relevant to the operation result. In one example, these might include, for example, a patient's weight, age, gender, and duration of illness; in another example, they might include, for example, a student's age, gender, current level (e.g., English proficiency level), whether they have attended similar courses, and their family's economic situation.

[0188] Then, in block 420, based on the prediction model and the individual characteristics of the one or more individuals, one or more subgroups to which the one or more individuals belong are determined. The prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods. The prediction model includes gating nodes that indicate corresponding individual characteristics and their segmentation points, based on which individuals can be divided into corresponding subgroups.

[0189] In one implementation, the prediction model can be a pre-established optimized model, which can be constructed through the following operations. For example, firstly, an initial prediction model is established for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes gating nodes for grouping individuals and multiple different expert nodes for performing predictions based on different prediction methods. The observation data includes multiple individual features of multiple individuals, corresponding operations in the predetermined operations performed on the multiple individuals, and corresponding operation results. Then, the parameters of the gating nodes and expert nodes of the initial prediction model are determined using the observation data to obtain the final prediction model. This process of pre-determining the prediction model is combined with the above. Figure 1 Blocks 110 and 120 are similar. For more details, please refer to [link / reference needed]. Figures 1-3 The descriptions made will not be repeated here.

[0190] Next, in block 430, according to the corresponding operation method, expert nodes associated with the identified one or more subgroups in the prediction model are used to predict the corresponding operation results for the one or more individuals. The predicted operation results may be, for example, values ​​indicating the possible effects of a patient taking the drug, or multiple effect values ​​for multiple different doses of the drug taken by the patient, or an effect curve for continuous doses of the drug taken by the patient.

[0191] Figure 5 A block diagram of a system for evaluating operational results according to one embodiment of the present disclosure is shown schematically. Figure 1As shown, the system includes an observation database 501, a prediction model construction unit 510, a model optimization unit 520, and an operation result estimation unit 530.

[0192] The observation database 501 stores observational data, such as individual patient characteristics, corresponding operations, and their effects. This data can be preprocessed, for example, through integration, reduction, and noise reduction of the raw data.

[0193] The prediction model building unit 510 receives the observation data and builds an initial prediction model based on the observation data. The prediction model can be the HME model mentioned earlier, which has a tree-like hierarchical structure and mixes multiple expert models based on gated empty nodes.

[0194] The model optimization unit 520 optimizes the initial network based on the observed variables using the objective optimization function. Optionally, the model optimization estimation unit may include an objective function determination unit 522, a latent variable optimization unit 524, a model reduction unit 526, and a node optimization unit 528.

[0195] The objective function determination unit 522 can determine the optimization objective function for the initial prediction model based on, for example, FAB. For instance, the asymptotic approximation of the marginal log-likelihood, i.e., the FIC function mentioned above, can be obtained using the variable distribution q mentioned above. Then, the lower bound of the FIC of the HME model is used as the optimization objective function of the FAB algorithm.

[0196] The latent variable optimization unit 524 can optimize the variable distribution q based on the above-mentioned objective function. The model reduction unit 526 can remove expert nodes and their related branches that have poor matching with the observed data based on the results of the above variable distribution optimization.

[0197] The node optimization unit 528 optimizes the parameters of the gated nodes and expert nodes in the reduced HME model based on the fir tree optimization objective function.

[0198] The model optimization unit will iterate the aforementioned latent variable optimization unit 524, model reduction unit 526, and node optimization unit 528 until FIC converges, thereby obtaining the final prediction model.

[0199] The operation result estimation unit 530 will obtain the prediction results of the expert nodes in the final prediction model for the observation data of each of its associated subgroups.

[0200] Furthermore, the operation result estimation unit 530 can also be an operationally separate unit from the aforementioned units. It can reside in the same electronic device as the aforementioned units, or on another electronic device storing the final prediction model. It can receive individual-related data, such as individual characteristics and the operation to be evaluated. In this way, based on the final prediction model, the subgroup to which the individual belongs and its corresponding expert model can be determined, and the operation result can be predicted for the individual based on the expert model, such as by combining... Figure 5 As described. Furthermore, it should be noted that the operational result data can also receive batches of individual data to be evaluated, and provide corresponding operational result evaluations for each individual.

[0201] Furthermore, it should be noted that the system can be implemented by a single electronic device or by multiple electronic devices; each functional block in the diagram can be implemented by a single processing unit or by multiple processing units in one or more electronic devices.

[0202] The implementation of the present disclosure will be described below with reference to some specific embodiments. However, it should be noted that these are merely exemplary and the present disclosure is not limited thereto.

[0203] In the medical field, understanding the effectiveness of a drug for different patient types is crucial for improving drug formulations and helping doctors or patients determine its suitability or appropriate dosage. Current methods typically rely on professionals' experience for analysis and selection. Even with confirmatory subgroup analysis methods, their sensitivity to hyperparameters means the accuracy of results still depends on the subjective judgment of professionals. The method disclosed herein, however, can automatically determine the final prediction network based on observational data and analyze the data to assess the drug's effectiveness for different patient types.

[0204] Figure 6 A schematic diagram illustrating the operational results of an evaluation of a predetermined operation on an individual subgroup according to a specific implementation of this disclosure is shown. Figure 6The data presented as input includes observational data from 198 children. Each data point indicates the effect of a particular calcium supplement or placebo on bone mineral density in each child. As shown in the figure, each observational data point includes seven variables: a treatment variable T, an outcome variable Y, and five covariates X1-X5. Treatment T is a binary treatment variable indicating whether the patient is taking a calcium supplement or a placebo. The outcome variable is a continuous output of whole-body bone mineral density (TBBMD). The five covariates indicate five patient characteristics, such as X1 being age (6-18 years), X2 being sex (male or female), X3 being race (white, African American, other), X4 being Donald stage (TS; values ​​1-6 for each variable), and X5 being duration of illness (2-9 years).

[0205] Data from 198 children is input into the processing unit of an electronic device. For these 198 data points, a 6-layer HME model with 32 expert nodes can be built, for example. Then, based on the FAB model, the model is reduced according to equations 19 and 20, for example, and the optimization parameters of the expert nodes and gating nodes in the model are determined using equations 24 and 25. Finally, the output can be as follows: Figure 6 The final network shown in the lower right corner displays the number of observational data points (i.e., individual subgroups) associated with each expert node, indicated to the left or right of the node. The values ​​representing the indicative treatment effects calculated using each expert network for the corresponding observational data are shown below the node. The evaluation values ​​shown are the average of all individuals within the subgroup, predicted by the expert model.

[0206] The results above show that the calcium supplement is effective for children under 15 years old with a disease duration of less than 6.5 years, but it is more effective in boys than girls. However, the calcium supplement is ineffective for patients with a disease duration of more than 6.5 years or those over 15 years of age.

[0207] Using embodiments of this disclosure, an efficacy assessment of a calcium compound for each subgroup can be output, and the results can be used to adjust subsequent drug formulations, for example. Additionally, doctors or patients can use this information to determine whether the calcium compound is suitable, assisting them in selecting appropriate medications. Alternatively or additionally, patient data can be directly input into an electronic device, which automatically determines the subgroup and corresponding expert model based on the patient's individual characteristics, and uses the expert model to determine the efficacy of the calcium compound. In cases where multiple drug-related models exist for a single disease, a more suitable and effective drug can be automatically selected for the patient.

[0208] Furthermore, in the field of education, students may need to choose from multiple similar courses with different arrangements (e.g., English courses with varying proportions of listening, speaking, reading, and writing), or educational institutions may need to recommend more suitable courses to students. Educational institutions typically possess observational data in this regard. This data may include characteristics of students who previously participated in courses D, E, F, etc., such as age, gender, whether they had participated in similar courses before, and their family's economic situation, as well as individual students' performance after participating in the corresponding courses, such as exam scores and awards received.

[0209] Based on such observational data, the method of this invention can be used to obtain different learning programs suitable for different types of students, thereby helping students currently choosing courses to make decisions or recommending more suitable courses. Similar to what was mentioned earlier, traditional methods cannot accurately and objectively recommend courses to students. However, the scheme provided in this disclosure, being insensitive to hyperparameters, can yield more accurate evaluation results.

[0210] Furthermore, it should be noted that the multiple expert models used in the prediction model disclosed herein are applicable to various types of data, such as discrete binary processing, ternary processing, or more complex processing, and are also suitable for continuous processing.

[0211] Figure 7 A schematic block diagram of an example device 700 that can be used to implement embodiments of the present disclosure is shown. As shown, device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. Various programs and data required for the operation of device 700 may also be stored in RAM 703. CPU 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0212] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0213] Processing unit 701 performs the various methods and processes described above, such as any one of processes 100, 300, and 400. For example, in some embodiments, processes 100, 300, and 400 may be implemented as computer software programs or computer program products tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more steps of any one of processes 100, 300, and 400 described above may be performed. Alternatively, in other embodiments, CPU 701 may be configured to perform any one of processes 100, 300, and 400 by any other suitable means (e.g., by means of firmware).

[0214] According to some embodiments of the present disclosure, a computer-readable medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to the present disclosure.

[0215] Those skilled in the art will understand that the various steps of the methods disclosed above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using device-executable program code, which can then be stored in a storage device for execution by the computing device. Alternatively, they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this disclosure is not limited to any particular combination of hardware and software.

[0216] It should be understood that although several devices or sub-devices of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more devices described above can be embodied in one device. Conversely, the features and functions of one device described above can be further divided and embodied by multiple devices.

[0217] The above description is merely an optional embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for evaluating drug efficacy, comprising: An initial prediction model is established for a set of observational data. This initial prediction model has a hierarchical structure and includes gating nodes for grouping individual patients and multiple expert nodes for performing predictions based on different prediction methods. The observational data includes multiple individual characteristics of the multiple patients, corresponding operations in predetermined operations performed on the multiple patients, and corresponding operation results. The multiple individual characteristics include the clinical characteristics of the multiple patients. The predetermined operations include drug intervention operations performed on the multiple patients, and the operation results include the treatment effects after performing the drug intervention operations on the multiple patients. Determine the optimization objective function for the initial prediction model, and optimize the initial prediction model based on the observation data and the optimization objective function to determine the final prediction model; as well as By utilizing the expert nodes in the final prediction model, predictions are made for patient subgroups in the observed data that match these nodes, in order to determine the outcome of the predetermined operation for each patient subgroup. The determination of the optimization objective function for the initial prediction model includes: determining the optimization objective function based on the decomposition progressive Bayesian method, and introducing latent variables into the optimization objective function, wherein the latent variables indicate whether each observation in the observation data matches each expert node in the prediction model. The optimization of the initial prediction model based on the observed data and the optimization objective function includes: The latent variables are optimized using the observation data based on the optimization objective function to determine the optimization probability of matching the observation data with each expert node; Based on the determined optimization probability and the judgment threshold used to determine whether to remove expert nodes, the existing prediction model is reduced to obtain a simplified prediction model. Based on the aforementioned objective function, the gate nodes and expert nodes in the simplified prediction model are optimized to determine the optimized parameter values ​​for each gate node and expert node, thereby generating an optimized prediction model. The process of determining the optimized probability distribution, reducing the current prediction model, and optimizing the gate nodes and expert nodes is repeated until the optimization objective function converges, thereby determining the final prediction model that can be used for prediction. Whether each observation in the observation data matches each expert node in the prediction model is expressed by the following formula: In this formula, e indicates the index of the expert node, g indicates the index of the gate node, and ε e h indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node. U (g, e) indicates whether a given observation is associated with the e-th expert node under the g-th gated node branch, where a value of 1 indicates that it is associated with the e-th expert node under the g-th gated node branch, and a value of 0 indicates that it is not associated with the e-th expert node. The initial prediction model mentioned therein is a hierarchical hybrid expert network model.

2. The method of claim 1, wherein the predetermined operation comprises at least three different treatment levels.

3. The method of claim 1, wherein the predetermined operation includes operations based on a continuous operation handling level.

4. A method for evaluating drug efficacy, comprising: Receive individual characteristics of one or more patient individuals and corresponding operational methods in a predetermined operation to be performed on the one or more patient individuals, wherein the individual characteristics include clinical characteristics of the one or more patient individuals, and the predetermined operation includes a drug intervention operation to be performed on the one or more patient individuals; Based on the prediction model and the individual characteristics of the one or more patients, one or more subgroups to which the one or more patients belong are determined. The prediction model has a hierarchical structure and includes gating nodes for grouping patients and multiple different expert nodes for performing predictions based on different prediction methods. as well as According to the corresponding operation method, expert nodes associated with the determined one or more subgroups in the prediction model are used to predict the corresponding operation results for the one or more individual patients, wherein the operation results include the treatment effect after performing the drug intervention operation for the one or more individual patients. The prediction model is predetermined through the following operations: An initial prediction model is established for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes a gating node for grouping individual patients and multiple different expert nodes for performing predictions based on different prediction methods. The observation data includes multiple individual characteristics of multiple patient individuals, corresponding operations in the predetermined operations performed on the multiple patient individuals, and corresponding operation results. A target function for optimizing the initial prediction model is determined. Based on the observed data and the target function, the initial prediction model is optimized to determine the final prediction model. The determination of the optimization objective function for the initial prediction model includes: determining the optimization objective function based on the decomposition progressive Bayesian method, and introducing latent variables into the optimization objective function, wherein the latent variables indicate whether each observation in the observation data matches each expert node in the prediction model. The optimization of the initial prediction model based on the observed data and the optimization objective function includes: The latent variables are optimized using the observation data based on the optimization objective function to determine the optimization probability of matching the observation data with each expert node; Based on the determined optimization probability and the judgment threshold used to determine whether to remove expert nodes, the existing prediction model is reduced to obtain a simplified prediction model. Based on the aforementioned objective function, the gate nodes and expert nodes in the simplified prediction model are optimized to determine the optimized parameter values ​​for each gate node and expert node, thereby generating an optimized prediction model. The process of determining the optimized probability distribution, reducing the current prediction model, and optimizing the gate nodes and expert nodes is repeated until the optimization objective function converges, thereby determining the final prediction model that can be used for prediction. Whether each observation in the observation data matches each expert node in the prediction model is expressed by the following formula: In this formula, e indicates the index of the expert node, g indicates the index of the gate node, and ε e h indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node. U (g, e) indicates whether a given observation is associated with the e-th expert node under the g-th gated node branch, where a value of 1 indicates that it is associated with the e-th expert node under the g-th gated node branch, and a value of 0 indicates that it is not associated with the e-th expert node. The initial prediction model mentioned therein is a hierarchical hybrid expert network model.

5. An electronic device for evaluating drug efficacy, comprising: processor; as well as A memory coupled to a processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions including: An initial prediction model is established for a set of observational data. This initial prediction model has a hierarchical structure and includes gating nodes for grouping individual patients and multiple expert nodes for performing predictions based on different prediction methods. The observational data includes multiple individual characteristics of the multiple patients, corresponding operations in predetermined operations performed on the multiple patients, and corresponding operation results. The multiple individual characteristics include the clinical characteristics of the multiple patients. The predetermined operations include drug intervention operations performed on the multiple patients, and the operation results include the treatment effects after performing the drug intervention operations on the multiple patients. Determine an optimization objective function for the initial prediction model, and optimize the initial prediction model based on the observed data and the optimization objective function to determine the final prediction model; and By utilizing the expert nodes in the final prediction model, predictions are made for patient subgroups in the observed data that match these nodes, in order to determine the outcome of the predetermined operation for each patient subgroup. The determination of the optimization objective function for the initial prediction model includes: determining the optimization objective function based on the decomposition progressive Bayesian method, and introducing latent variables into the optimization objective function, wherein the latent variables indicate whether each observation in the observation data matches each expert node in the prediction model. The optimization of the initial prediction model based on the observed data and the optimization objective function includes: The latent variables are optimized using the observation data based on the optimization objective function to determine the optimization probability of matching the observation data with each expert node; Based on the determined optimization probability and the judgment threshold used to determine whether to remove expert nodes, the existing prediction model is reduced to obtain a simplified prediction model. Based on the aforementioned objective function, the gate nodes and expert nodes in the simplified prediction model are optimized to determine the optimized parameter values ​​for each gate node and expert node, thereby generating an optimized prediction model. The process of determining the optimized probability distribution, reducing the current prediction model, and optimizing the gate nodes and expert nodes is repeated until the optimization objective function converges, thereby determining the final prediction model that can be used for prediction. Whether each observation in the observation data matches each expert node in the prediction model is expressed by the following formula: In this formula, e indicates the index of the expert node, g indicates the index of the gate node, and ε e h indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node. U (g, e) indicates whether a given observation is associated with the e-th expert node under the g-th gated node branch, where a value of 1 indicates that it is associated with the e-th expert node under the g-th gated node branch, and a value of 0 indicates that it is not associated with the e-th expert node. The initial prediction model mentioned therein is a hierarchical hybrid expert network model.

6. The electronic device of claim 5, wherein the predetermined operation comprises at least three different levels of operation.

7. The electronic device of claim 5, wherein the predetermined operation includes operation based on a continuous operation handling level.

8. An electronic device for evaluating drug efficacy, comprising: processor; as well as A memory coupled to a processor, the memory having instructions stored therein, the instructions causing the electronic device to perform actions when executed by the processor, the actions including: Receive individual characteristics of one or more patient individuals and corresponding operational methods in a predetermined operation to be performed on the one or more patient individuals, wherein the individual characteristics include clinical characteristics of the one or more patient individuals, and the predetermined operation includes a drug intervention operation to be performed on the one or more patient individuals; Based on a prediction model and the individual characteristics of the one or more patient individuals, one or more subgroups to which the one or more patient individuals belong are determined. The prediction model has a hierarchical structure and includes gating nodes for grouping patient individuals and multiple different expert nodes for performing predictions based on different prediction methods. According to the corresponding operation method, expert nodes associated with the determined one or more subgroups in the prediction model are used to predict the corresponding operation results for the one or more individual patients, wherein the operation results include the treatment effect after performing the drug intervention operation for the one or more individual patients. The prediction model is predetermined through the following operations: An initial prediction model is established for a set of observation data, wherein the initial prediction model has a hierarchical structure and includes a gating node for grouping individual patients and multiple different expert nodes for performing predictions based on different prediction methods. The observation data includes multiple individual characteristics of multiple patient individuals, corresponding operations in the predetermined operations performed on the multiple patient individuals, and corresponding operation results. A target function for optimizing the initial prediction model is determined. Based on the observed data and the target function, the initial prediction model is optimized to determine the final prediction model. The determination of the optimization objective function for the initial prediction model includes: determining the optimization objective function based on the decomposition progressive Bayesian method, and introducing latent variables into the optimization objective function, wherein the latent variables indicate whether each observation in the observation data matches each expert node in the prediction model. The optimization of the initial prediction model based on the observed data and the optimization objective function includes: The latent variables are optimized using the observation data based on the optimization objective function to determine the optimization probability of matching the observation data with each expert node; Based on the determined optimization probability and the judgment threshold used to determine whether to remove expert nodes, the existing prediction model is reduced to obtain a simplified prediction model. Based on the aforementioned objective function, the gate nodes and expert nodes in the simplified prediction model are optimized to determine the optimized parameter values ​​for each gate node and expert node, thereby generating an optimized prediction model. The process of determining the optimized probability distribution, reducing the current prediction model, and optimizing the gate nodes and expert nodes is repeated until the optimization objective function converges, thereby determining the final prediction model that can be used for prediction. Whether each observation in the observation data matches each expert node in the prediction model is expressed by the following formula: In this formula, e indicates the index of the expert node, g indicates the index of the gate node, and ε e h indicates the indices of all gated nodes on the unique path from the root node to the e-th expert node. U (g, e) indicates whether a given observation is associated with the e-th expert node under the g-th gated node branch, where a value of 1 indicates that it is associated with the e-th expert node under the g-th gated node branch, and a value of 0 indicates that it is not associated with the e-th expert node. The initial prediction model mentioned therein is a hierarchical hybrid expert network model.

9. A computer program product tangibly stored on a computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to any one of claims 1 to 3.

10. A computer program product tangibly stored on a computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to claim 4.

Citation Information

Patent Citations

  • Feature selection method and equipment used for constructing HME (Hierarchical Mixtures of Expert) system

    CN106557451A

  • Auxiliary medication decision-making method and intelligent auxiliary medication system

    CN109859815A

  • Application method of hybrid expert system in lung adenocarcinoma classification

    CN110706804A