Model generation method, device and equipment, computer readable medium and program product
By detecting the value-transformed sample subset of sample distribution transformation and generating sample weights, model training and integration fusion, the training cost increase caused by sample distribution transformation is solved, and accurate value detection is achieved.
Patent Information
- Application Number
- CN202311529124.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art cannot effectively solve the problem when processing the transformation of sample distribution, resulting in increased training calculation costs and waste of resources.
By obtaining the value transformation sample set, detect the value transformation sample subset of the transformed sample distribution, determine the sample weight corresponding to each sample, perform model training, generate an ensemble model, and integrate it with the pre-generated value detection information generation model for model integration and integration.
Without wasting computing resources, an integrated model is generated to accurately predict samples whose sample distribution is transformed, which effectively solves the sample distribution transformation problem and improves the accuracy of value detection.
Smart Images

Figure CN120011799A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a model generation method, apparatus, device, computer-readable medium, and program product. Background Art
[0002] At present, with the vigorous development of big data and machine learning technology, more and more massive data and machine learning models are being applied by users in real environments. In the application of machine learning models, there is often a problem of sample distribution transformation. To deal with the problem of sample distribution transformation, the usual method is to re-acquire the training data set and retrain the relevant network model to solve the inaccurate network model output caused by the sample distribution transformation.
[0003] However, the inventors have found that when the above method is adopted, the following technical problems often occur:
[0004] It is not possible to adaptively solve the essence of the problem of sample distribution transformation, but only solve it by retraining the network model, which increases the training calculation cost and leads to a waste of resources.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0007] Some embodiments of the present disclosure propose a model generation method, apparatus, device, computer-readable medium, and program product to solve the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a model generation method, including: obtaining a value transformation sample set; detecting a value transformation sample subset having concept drift characteristics in the value transformation sample set; determining a sample weight corresponding to each value transformation sample in the value transformation sample subset for an initial integrated model to be added, and obtaining a sample weight set; performing model training on the initial integrated model to be added according to the value transformation sample subset and the sample weight set, and obtaining an integrated model; and integrating and fusing the integrated model with a pre-generated value detection information generation model to generate an integrated value detection information generation model.
[0009] Optionally, the above method also includes: obtaining model prediction information of the above integrated value detection information generation model within the future target time window; and generating model effect information corresponding to the above integrated value detection information generation model based on the above model prediction information.
[0010] Optionally, the value transformation sample set is a sample set stored in the form of data blocks; and the value transformation sample subset for detecting the existence of concept drift characteristics in the value transformation sample set includes: obtaining at least one historical area sample division result corresponding to at least one target data block, wherein the target data block is a storage data block corresponding to the historical value transformation sample set; determining the current area sample division result of the data block corresponding to the value transformation sample set; for each historical area sample division result in the at least one historical area sample division result, determining the area sample division difference information between the historical area sample division result and the current area sample division result; generating difference area information based on the at least one area sample division difference information obtained; determining at least one value transformation sample in the data block corresponding to the value transformation sample set, whose corresponding area information is the difference area information, as the value transformation sample subset.
[0011] Optionally, the above-mentioned determination of the current area sample division result of the data block corresponding to the above-mentioned value transformation sample set includes: determining the sample attribute set corresponding to the above-mentioned value transformation sample set; determining the first attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the above-mentioned value transformation sample set; screening out at least one sample attribute whose corresponding first attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set as at least one first sample attribute; screening out a first target sample attribute from the above-mentioned at least one first sample attribute; dividing the above-mentioned value transformation sample set into a sample set according to the above-mentioned first target sample attribute to obtain at least one first sample subset; for each first sample subset in the at least one first sample subset, performing the following determination steps: determining the second attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the first sample subset; screening out at least one sample attribute whose corresponding second attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set; This attribute is used as at least one second sample attribute; a second target sample attribute is screened out from the at least one second sample attribute; according to the second target sample attribute, the sample subset is divided into sample sets to obtain at least one second sample subset; for each second sample subset in the at least one second sample subset, third attribute division information corresponding to each sample attribute in the above sample attribute set and for the above second sample subset is determined; in response to determining that each third attribute division information in the at least one third attribute division information set obtained satisfies the corresponding attribute division condition, and the number of samples included in the second sample subset in the at least one second sample subset is less than a predetermined value, the sample division result corresponding to the at least one second sample subset is determined as the sample division result corresponding to the first sample subset; according to the at least one sample division result obtained, each area division result corresponding to each data area of the data block corresponding to the above value transformation sample set is determined as the current area sample division result.
[0012] Optionally, before determining the respective regional division results corresponding to the respective data regions for the data blocks corresponding to the above-mentioned value transformation sample set as the current regional sample division result based on at least one sample division result obtained, the method further includes: in response to determining that at least one third attribute division information set contains third attribute division information that does not satisfy the corresponding attribute division condition, and the number of samples included in the second sample subset of at least one second sample subset is greater than or equal to a predetermined value, determining a second sample subset group that does not satisfy the corresponding attribute division condition; determining the second sample subset group as at least one first sample subset, and continuing to execute the above-mentioned determination step.
[0013] Optionally, the value detection information generation model includes: a historical integration model sequence; and the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integration model to be added is determined to obtain a sample weight set, including: for the historical integration models in the historical integration model sequence, the following generation steps are performed: in response to determining that there is a model in the historical integration model sequence whose corresponding model position is located before the historical integration model, an adjacent historical integration model located before the historical integration model and adjacent to the historical integration model is determined; the sample weight corresponding to each value transformation sample in the value transformation sample subset for the adjacent historical integration model is determined as the adjacent sample weight to obtain an adjacent sample weight set; based on the adjacent sample weight set, a sample weight set for the historical integration model is generated; the sample weight set corresponding to the historical integration model at the target position in the historical integration model sequence is determined as the target sample weight set; based on the target sample weight set, a sample weight set for the initial integration model to be added is generated.
[0014] Optionally, the above-mentioned value detection information generation model includes: a historical integrated model sequence; and the above-mentioned integrated model is integrated and fused with the pre-generated value detection information generation model to generate an integrated value detection information generation model, including: determining the model diversity weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model diversity weight corresponding to the above-mentioned integrated model to obtain a model diversity weight set; determining the model discrimination weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model discrimination weight corresponding to the above-mentioned integrated model to obtain a model discrimination weight set; according to the above-mentioned model diversity weight set and the above-mentioned model discrimination weight set, the above-mentioned integrated model and the above-mentioned historical integrated model sequence are model integrated and fused to generate an integrated value detection information generation model.
[0015] In the second aspect, some embodiments of the present disclosure provide a model generation device, including: an acquisition unit, configured to acquire a value transformation sample set; a detection unit, configured to detect a value transformation sample subset having concept drift characteristics in the above value transformation sample set; a determination unit, configured to determine the sample weight corresponding to each value transformation sample in the above value transformation sample subset for the initial integrated model to be added, and obtain a sample weight set; a training unit, configured to perform model training on the above initial integrated model to be added according to the above value transformation sample subset and the above sample weight set, and obtain an integrated model; an integrated fusion unit, configured to perform model integration fusion on the above integrated model and a pre-generated value detection information generation model, so as to generate an integrated value detection information generation model.
[0016] Optionally, the device also includes: obtaining model prediction information of the above-mentioned integrated value detection information generation model within the future target time window; and generating model effect information corresponding to the above-mentioned integrated value detection information generation model based on the above-mentioned model prediction information.
[0017] Optionally, the value transformation sample set is a sample set stored in the form of a data block; and the detection unit can be configured to: obtain at least one historical area sample division result corresponding to at least one target data block, wherein the target data block is a storage data block corresponding to the historical value transformation sample set; determine the current area sample division result of the data block corresponding to the above value transformation sample set; for each historical area sample division result in the above at least one historical area sample division result, determine the area sample division difference information between the above historical area sample division result and the above current area sample division result; generate difference area information based on the at least one area sample division difference information obtained; determine at least one value transformation sample in the data block corresponding to the above value transformation sample set, whose corresponding area information is the difference area information, as a value transformation sample subset.
[0018] Optionally, the detection unit can be configured to: determine the sample attribute set corresponding to the above-mentioned value transformation sample set; determine the first attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the above-mentioned value transformation sample set; filter out at least one sample attribute whose corresponding first attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set as at least one first sample attribute; filter out the first target sample attribute from the above-mentioned at least one first sample attribute; divide the above-mentioned value transformation sample set into sample sets according to the above-mentioned first target sample attribute to obtain at least one first sample subset; for each first sample subset in at least one first sample subset, perform the following determination steps: determine the second attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the first sample subset; filter out at least one sample attribute whose corresponding second attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set as at least one first sample attribute The invention relates to a method for obtaining a sample set of at least one second sample attribute; filtering out a second target sample attribute from at least one second sample attribute; dividing the sample subset according to the second target sample attribute to obtain at least one second sample subset; for each second sample subset in at least one second sample subset, determining third attribute division information corresponding to each sample attribute in the above sample attribute set and for the above second sample subset; in response to determining that each third attribute division information in the at least one third attribute division information set obtained satisfies the corresponding attribute division condition, and the number of samples included in the second sample subset in at least one second sample subset is less than a predetermined value, determining the sample division result corresponding to the at least one second sample subset as the sample division result corresponding to the first sample subset; and determining, according to the at least one sample division result obtained, each area division result corresponding to each data area of the data block corresponding to the above value transformation sample set as the current area sample division result.
[0019] Optionally, the detection unit can be configured to: in response to determining that there is third attribute division information that does not satisfy the corresponding attribute division condition in at least one third attribute division information set, and the number of samples included in the second sample subset in at least one second sample subset is greater than or equal to a predetermined value, determine the second sample subset group that does not satisfy the corresponding attribute division condition; determine the second sample subset group as at least one first sample subset, and continue to perform the above determination step.
[0020] Optionally, the above-mentioned value detection information generation model includes: a historical integration model sequence; and the determination unit can be configured to: for the historical integration models in the above-mentioned historical integration model sequence, perform the following generation steps: in response to determining that there is a model in the above-mentioned historical integration model sequence whose corresponding model position is located before the above-mentioned historical integration model, determine an adjacent historical integration model that is located before the above-mentioned historical integration model and is adjacent to the above-mentioned historical integration model; determine the sample weight corresponding to each value transformation sample in the above-mentioned value transformation sample subset for the above-mentioned adjacent historical integration model as the adjacent sample weight, and obtain the adjacent sample weight set; based on the above-mentioned adjacent sample weight set, generate a sample weight set for the above-mentioned historical integration model; determine the sample weight set corresponding to the historical integration model located at the target position in the above-mentioned historical integration model sequence as the target sample weight set; based on the above-mentioned target sample weight set, generate a sample weight set for the above-mentioned initial integration model to be added.
[0021] Optionally, the above-mentioned value detection information generation model includes: a historical integrated model sequence; and the integration fusion unit can be configured to: determine the model diversity weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model diversity weight corresponding to the above-mentioned integrated model, and obtain a model diversity weight set; determine the model discrimination weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model discrimination weight corresponding to the above-mentioned integrated model, and obtain a model discrimination weight set; according to the above-mentioned model diversity weight set and the above-mentioned model discrimination weight set, perform model integration fusion on the above-mentioned integrated model and the above-mentioned historical integrated model sequence to generate an integrated value detection information generation model.
[0022] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0023] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0024] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner in the above-mentioned first aspect when executed by a processor.
[0025] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the model generation method of some embodiments of the present disclosure, without wasting computing resources, an integrated value detection information generation model can be obtained for accurately predicting samples with transformed sample distribution. Specifically, the reason for the inability to effectively solve the problem of sample distribution transformation is that it is not possible to adaptively solve the essence of the problem of sample distribution transformation, and only solve it by retraining the network model, resulting in an increase in training computing costs and a waste of resources. Based on this, the model generation method of some embodiments of the present disclosure first obtains a value transformation sample set for a sample set with a subsequent sample distribution transformation. Then, a value transformation sample subset with a transformed sample distribution in the above value transformation sample set is detected. Here, by screening out the value transformation sample subset with a transformed sample distribution, the sample distribution transformation problem can be effectively solved in the subsequent process, avoiding the situation where the value anomaly detection is inaccurate due to the sample distribution transformation problem. Next, the sample weight corresponding to each value transformation sample in the above value transformation sample subset for the initial integrated model to be added is determined to obtain a sample weight set. Here, by determining the sample weight corresponding to each value transformation sample, the initial integrated model to be added later is used for targeted model training, so that the initial integrated model can learn more characteristic information of the transformed sample distribution and generate more accurate value detection information. Furthermore, according to the above-mentioned value transformation sample subset and the above-mentioned sample weight set, the above-mentioned initial integrated model to be added is trained to obtain an integrated model that generates more accurate value detection information. Finally, the above-mentioned integrated model is integrated and fused with the pre-generated value detection information generation model to generate an integrated value detection information generation model. In summary, through the detection of the value transformation sample subset and the generation of the sample weight set, an integrated model that generates more accurate value detection information and can solve the sample distribution transformation problem in a targeted manner is trained. Therefore, through the integrated fusion of the integrated model and the value detection information generation model, not only can accurate value detection of the value transformation sample subset be guaranteed, but also accurate detection of the remaining value transformation sample sets can be guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0027] Figure 1 is a schematic diagram of an application scenario of the model generation method according to some embodiments of the present disclosure;
[0028] Figure 2is a flow chart of some embodiments of the model generation method according to the present disclosure;
[0029] Figure 3 are flow charts of other embodiments of the model generation method according to the present disclosure;
[0030] Figure 4 is a schematic diagram of the structure of some embodiments of the model generation device according to the present disclosure;
[0031] Figure 5 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0032] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0033] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0034] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0035] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0037] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0038] Figure 1 It is a schematic diagram of an application scenario of the model generation method according to some embodiments of the present disclosure.
[0039] exist Figure 1In the application scenario, first, the electronic device 101 can obtain the value transformation sample set 102. In this application scenario, the value transformation sample set 102 may include: a first value transformation sample 1021, a second value transformation sample 1022, a third value transformation sample 1023, and a fourth value transformation sample 1024. Then, the electronic device 101 can detect the value transformation sample subset 103 in which the sample distribution in the above value transformation sample set 102 has been transformed. In this application scenario, the sample distribution corresponding to the value transformation sample set 102 may be a normal distribution. The sample distribution corresponding to the value transformation sample subset 103 is a t distribution. The value transformation sample subset 103 may include: a first value transformation sample 1021, a second value transformation sample 1022, and a fourth value transformation sample 1024. Next, the electronic device 101 can determine the sample weight corresponding to each value transformation sample in the above value transformation sample subset 103 for the initial integrated model 105 to be added, and obtain the sample weight set 104. In this application scenario, the initial integrated model 105 to be added may be an initial XGBoost (eXtreme Gradient Boosting) model. The sample weight set 104 may include: a sample weight 1041 corresponding to the first value transformation sample, a sample weight 1042 corresponding to the second value transformation sample, and a sample weight 1043 corresponding to the third value transformation sample. Sample weight 1041 may be "0.3". Sample weight 1042 may be "0.4". Sample weight 1043 may be "0.5". Furthermore, the electronic device 101 may perform model training on the initial integrated model 105 to be added according to the above-mentioned value transformation sample subset 103 and the above-mentioned sample weight set 104 to obtain an integrated model 106. In this application scenario, the integrated model 106 may be an XGBoost model. Finally, the electronic device 101 may integrate and fuse the above-mentioned integrated model 106 with the pre-generated value detection information generation model 107 to generate an integrated value detection information generation model 108.
[0040] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.
[0041] It should be understood that Figure 1 The number of electronic devices in the embodiment is only for illustration. Any number of electronic devices may be provided according to implementation requirements.
[0042] Continue to refer Figure 2 , shows a process 200 of some embodiments of the model generation method according to the present disclosure. The model generation method comprises the following steps:
[0043] Step 201, obtaining a value transformation sample set.
[0044] In some embodiments, the execution entity of the above model generation method (for example Figure 1 The electronic device 101 shown) can obtain the value transformation sample set through a wired connection method or a wireless connection method. Among them, the value transformation samples in the value transformation sample set can be sample data that characterizes the value transformation. Specifically, the corresponding value sample set is different for different execution tasks. For example, for the risk control field, the execution task is a user credit information evaluation task, and the value transformation samples in the corresponding value transformation sample set may include user value transformation information (that is, user asset change information). For the risk control field, the execution task is a value lending credit evaluation task (that is, pre-loan credit evaluation), and the value transformation samples in the corresponding value transformation sample set may include user value lending information (lending information) and user value borrowing information (borrowing information).
[0045] Step 202, detecting a value transformation sample subset in which the sample distribution in the value transformation sample set has been transformed.
[0046] In some embodiments, the execution subject may detect a value transformation sample subset in which the sample distribution in the value transformation sample set is transformed. For example, the sample distribution is transformed from a normal distribution to an F (F-distribution) distribution. For another example, the sample distribution is transformed from a normal distribution to a t (t-distribution) distribution.
[0047] As an example, the execution entity may use the KS (Kolmogorov-Smirnov) test method to detect a value transformation sample subset in which the sample distribution in the value transformation sample set has been transformed.
[0048] In some optional implementations of some embodiments, the value transformation sample set is a sample set stored in the form of data blocks.
[0049] Optionally, the above-mentioned detecting the value transformation sample subset in which the sample distribution in the above-mentioned value transformation sample set has a transformation may include the following steps:
[0050] In the first step, the execution subject can obtain at least one historical area sample division result corresponding to at least one target data block. The target data block is a storage data block corresponding to the historical value transformation sample set. There is a one-to-one correspondence between the target data block in at least one target data block and the historical area sample division result in at least one historical area sample division result. The historical area sample division result can be the historical sample division result corresponding to each data area in the target data block. The historical sample division result can be the result of sample set division of the historical value transformation sample set for the corresponding data area. That is, the divided sample subset corresponding to the historical sample division result is stored in the corresponding data area. The historical value transformation sample set can be the value transformation sample set obtained historically.
[0051] In the second step, the execution subject can determine the current area sample division result of the data block corresponding to the value transformation sample set. Among them, the current area sample division result can characterize the sample storage situation of each data area corresponding to the data block corresponding to the value transformation sample set. For example, the value transformation sample set includes: the first value transformation sample, the second value transformation sample, the third value transformation sample, the fourth value transformation sample, and the fifth value transformation sample. The data areas corresponding to the data block corresponding to the value transformation sample set include: the first data area, the second data area, and the third data area. The current area sample division result can be: "First data area: the first value transformation sample, the second value transformation sample and the third value transformation sample, the second data area: the fourth value transformation sample, the third data area: the fifth value transformation sample".
[0052] The third step is to determine the regional sample division difference information between the historical regional sample division result and the current regional sample division result for each of the at least one historical regional sample division result. The regional sample division difference information may represent the sample difference information between the sample set corresponding to the historical regional sample division result and the sample set corresponding to the current regional sample division result. In practice, the sample difference information may be, but is not limited to, at least one of the following: sample quantity difference information, sample content difference information.
[0053] The fourth step is to generate difference area information based on the obtained difference information of at least one regional sample division. The difference area information may be data information of a data area where the sample set difference between at least one target data block and the corresponding data block of the value transformation sample set satisfies a corresponding difference condition. There is a positional correspondence between each data area corresponding to the target data block and each data area corresponding to the corresponding data block of the value transformation sample set. The difference condition may be a condition that the difference between the sample sets satisfies a corresponding difference threshold. For example, the difference between the sample sets may be the number of samples included in the sample sets. The difference threshold may be a preset sample difference number. The corresponding difference condition may be a condition that the difference between the sample sets is greater than the difference threshold.
[0054] As an example, first, the execution subject may select the regional sample division difference information whose corresponding difference information is greater than a preset threshold from the at least one regional sample division difference information as the target regional sample division difference information. Then, the regional information of the data region corresponding to the target regional sample division difference information is used as the difference regional information.
[0055] As another example, first, the execution subject may select a subset of regional sample division difference information whose difference size is within the first target number from at least one regional sample division difference information. Then, the data regional information subset corresponding to the regional sample division difference information subset is fused to generate fused regional information as the difference regional information.
[0056] In a fifth step, the execution entity may determine at least one value transformation sample in the data block corresponding to the value transformation sample set, whose corresponding region information is the difference region information, as a value transformation sample subset.
[0057] In some optional implementations of some embodiments, the above-mentioned determining the current area sample division result of the data block corresponding to the above-mentioned value transformation sample set includes:
[0058] The first step is to determine the sample attribute set corresponding to the above value transformation sample set.
[0059] The sample attributes in the sample data set may be attributes of the value transformation samples. For example, the value transformation samples are samples corresponding to the user credit information evaluation task. The sample attribute set may include but is not limited to at least one of the following: user value transformation time (user asset transformation time), user value transformation size (asset transformation size), user value transformation category (asset transformation category).
[0060] In the second step, based on the above value transformation sample set, the first attribute division information corresponding to each sample attribute in the above sample attribute set is determined. The first attribute division information may be attribute judgment basis information of the attribute for dividing the sample set into samples. The divisible attribute may be the sample attribute for dividing the sample set into samples. The attribute judgment basis information may be the judgment basis of the divisible attribute. For example, the attribute judgment basis information may be the difference between the maximum value and the minimum value of the sample attribute in the sample set.
[0061] Each sample attribute has corresponding judgment basis information. For example, the sample attribute set includes: a first sample attribute, a second sample attribute, and a third sample attribute. The first attribute division information corresponding to the first sample attribute can be: the difference between the maximum value and the minimum value of the first sample attribute in the value transformation sample set is 10. The first attribute division information corresponding to the second sample attribute can be: the difference between the maximum value and the minimum value of the second sample attribute in the value transformation sample set is 12. The first attribute division information corresponding to the third sample attribute can be: the attribute standard information corresponding to the third sample attribute can be: the difference between the maximum value and the minimum value of the third sample attribute in the value transformation sample set is greater than 15.
[0062] As an example, first, the execution subject may determine the attribute judgment basis corresponding to each sample attribute in the sample attribute set. For example, the attribute judgment basis may be the difference between the maximum value and the minimum value of the corresponding attribute. Then, based on the value transformation sample set, the attribute judgment basis information corresponding to the attribute judgment basis corresponding to each sample attribute is determined. That is, the attribute judgment basis information may be the difference between the maximum value and the minimum value of the attribute in the value transformation sample set.
[0063] In the third step, at least one sample attribute whose corresponding first attribute division information does not meet the attribute division condition is selected from the above sample attribute set as at least one first sample attribute. Each of the at least one first sample attribute may be a divisible attribute for the sample set. That is, the sample set may be divided into samples based on the divisible attributes. The attribute division condition may be that the value corresponding to the first attribute division information is less than the attribute threshold corresponding to the first attribute division attribute. The attribute threshold may be pre-set.
[0064] For example, the attribute standard information corresponding to the first sample attribute may be that the difference between the maximum value and the minimum value of the first sample attribute in the value transformation sample set is greater than 11. Then the first sample attribute is not a divisible attribute, and the value transformation sample set cannot be divided into samples based on the first sample attribute in the future. The attribute standard information corresponding to the second sample attribute may be that the difference between the maximum value and the minimum value of the second sample attribute in the value transformation sample set is greater than 10. Then the second sample attribute is a divisible attribute, and the value transformation sample set can be divided into samples based on the second sample attribute in the future. The difference between the maximum value and the minimum value of the third sample attribute in the value transformation sample set is 14. Then the third sample attribute is not a divisible attribute, and the value transformation sample set cannot be divided into samples based on the third sample attribute in the future.
[0065] The fourth step is to select a first target sample attribute from the at least one first sample attribute, wherein the first target sample attribute is an attribute to be divided for sample division of the sample set.
[0066] As an example, the execution subject may randomly select a first sample attribute from at least one first sample attribute as the first target sample attribute.
[0067] The fifth step is to divide the value transformation sample set into sample sets according to the first target sample attributes to obtain at least one first sample subset.
[0068] As an example, the execution subject may determine the median value of the first target sample attribute in the sample set, and then evenly divide the value transformation sample set according to the median value to obtain at least one first sample subset.
[0069] As another example, the execution subject may determine an average value of the first target sample attribute in the sample set, and then divide the value transformation sample set into uniform groups according to the average value to obtain at least one first sample subset.
[0070] Step 6: for each first sample subset in at least one first sample subset, perform the following determination steps:
[0071] Sub-step 1: Determine, based on the first sample subset, the second attribute division information corresponding to each sample attribute in the sample attribute set. The specific implementation will not be described in detail, see the generation of the first attribute division information.
[0072] Sub-step 2: Filter out at least one sample attribute whose corresponding second attribute division information does not meet the attribute division condition from the above sample attribute set as at least one second sample attribute. The specific implementation will not be described in detail, please refer to the generation of at least one first sample attribute.
[0073] Sub-step 3: Filter out a second target sample attribute from at least one second sample attribute. The specific implementation will not be described in detail, and please refer to the generation of the first target sample attribute.
[0074] Sub-step 4: divide the sample subsets according to the second target sample attribute to obtain at least one second sample subset. The specific implementation will not be repeated here, and please refer to the generation of at least one first sample subset.
[0075] Sub-step 5: for each second sample subset in at least one second sample subset, determine the third attribute division information corresponding to each sample attribute in the sample attribute set. The specific implementation will not be repeated here, see the generation of the first attribute division information.
[0076] Sub-step 6, in response to determining that each piece of third attribute division information in the obtained at least one third attribute division information set satisfies the corresponding attribute division condition, determining the sample division result corresponding to the at least one second sample subset as the sample division result corresponding to the first sample subset.
[0077] The seventh step is to determine, based on the at least one sample division result obtained, each area division result corresponding to each data area of the data block corresponding to the above-mentioned value transformation sample set as the current area sample division result.
[0078] As an example, the execution entity may determine, based on at least one sample division result and by using regional sample statistics, each regional division result corresponding to each data region of the data block corresponding to the value transformation sample set as the current regional sample division result.
[0079] Specifically, a KDQ-Tree (KD-Quad Tree, K-dimensional quadtree) algorithm may be used to detect a value transformation sample subset in which the sample distribution in the value transformation sample set has been transformed.
[0080] Optionally, after the "seventh step", the steps further include:
[0081] In the first step, in response to determining that there is third attribute division information that does not satisfy the corresponding attribute division condition in at least one third attribute division information set, and the number of samples included in the second sample subset in at least one second sample subset is greater than or equal to a predetermined value, a second sample subset group that does not satisfy the corresponding attribute division condition is determined. The predetermined value may be a pre-set value. For example, the predetermined value may be "10".
[0082] In the second step, the second sample subset group is determined as at least one first sample subset, and the above determination step is continued.
[0083] Optionally, after the "seventh step", the steps further include:
[0084] The first step is to determine a second sample subset group whose corresponding sample number is greater than or equal to a predetermined value in response to determining that at least one third attribute division information set contains third attribute division information that satisfies a corresponding attribute division condition, and at least one second sample subset contains a second sample subset whose corresponding sample number is greater than or equal to a predetermined value.
[0085] In the second step, each second sample subset in the second sample subset group is divided into samples to generate at least one third sample subset, thereby obtaining at least one third sample subset group, wherein the number of samples corresponding to each third sample subset is less than a predetermined value.
[0086] Optionally, after the "seventh step", the steps further include:
[0087] In response to determining that there is third attribute division information that does not satisfy the corresponding attribute division condition in at least one third attribute division information set, and there is a second sample subset in which the number of corresponding samples is less than a predetermined value in at least one second sample subset, the sample division result corresponding to the at least one second sample subset is determined as the sample division result corresponding to the first sample subset.
[0088] Step 203 , determining the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtaining a sample weight set.
[0089] In some embodiments, the above-mentioned execution entity may determine the sample weight corresponding to each value transformation sample in the above-mentioned value transformation sample subset for the initial integrated model to be added, and obtain a sample weight set. Among them, the sample weight may characterize the importance of the model learning the sample semantics. The above-mentioned sample weight may be a numerical value between 0 and 1. Specifically, the sample weight for the initial integrated model to be added may be a weight that characterizes the importance of the initial integrated model to be added learning the sample semantics. Among them, the initial integrated model to be added may be an initial integrated model to be subsequently integrated with the value detection information generation model for model integration. The initial integrated model may be an integrated model that has not yet been trained. In practice, the integrated model may be an XGBoost model.
[0090] In some optional implementations of some embodiments, the above-mentioned value detection information generation model includes: a historical integration model sequence. Among them, each historical integration model in the historical integration model sequence can be arranged according to the integration time. The integration time can be the model integration time. The historical integration model can be an integration model integrated and fused by history. For example, the historical integration model sequence may include: a first historical integration model, a second historical integration model, a third historical integration model and a fourth historical integration model. The model integration time corresponding to the first historical integration model is earlier than the model integration time corresponding to the second historical integration model. The model integration time corresponding to the second historical integration model is earlier than the model integration time corresponding to the third historical integration model. The model integration time corresponding to the third historical integration model is earlier than the model integration time corresponding to the fourth historical integration model. The historical integration models in the historical integration model sequence can be algorithm models based on integrated learning.
[0091] Optionally, the above-mentioned determination of the sample weight corresponding to each value transformation sample in the above-mentioned value transformation sample subset for the initial integrated model to be added to obtain the sample weight set may include the following steps:
[0092] In the first step, for the historical integration model in the above historical integration model sequence, the following generation steps are performed:
[0093] Sub-step 1, in response to determining that there is a model in the above-mentioned historical integrated model sequence whose corresponding model position is located before the above-mentioned historical integrated model, determine an adjacent historical integrated model that is located before the above-mentioned historical integrated model and adjacent to the above-mentioned historical integrated model.
[0094] For example, the historical integration model sequence may include: a first historical integration model, a second historical integration model, a third historical integration model and a fourth historical integration model. The above historical integration model is the second historical integration model. The adjacent historical integration model is the first historical integration model.
[0095] Sub-step 2: The execution entity may determine the sample weight corresponding to each value change sample in the value change sample subset for the adjacent historical integration model as the adjacent sample weight, and obtain an adjacent sample weight set.
[0096] Sub-step 3: The execution entity may generate a sample weight set for the historical integration model based on the adjacent sample weight set.
[0097] As an example, the execution subject may generate a sample weight set for the historical integration model based on a neighboring sample weight set and using the sample weight update formula of Adaboost.
[0098] As another example, the execution entity may generate a sample weight set for the historical integration model by using the following sample weight update formula:
[0099]
[0100] Among them, w m+1,i is the sample weight corresponding to the i-th sample in the m+1-th historical integrated model. m,i is the sample weight corresponding to the mth historical ensemble model and the ith sample. m is the model weight corresponding to the mth historical ensemble model. m It can be the normalization factor corresponding to the mth historical ensemble model. i It can be the i-th sample. G m (x i ) can be for x i The model output result of the mth historical ensemble model. i It can be the sample label corresponding to the i-th sample.
[0101] α m It is generated by the following formula:
[0102]
[0103] Among them, e m It can be the classification error rate corresponding to the mth historical integrated model. The classification error rate can be the weighted sum of the sample weights corresponding to the misclassified sample subset in the value transformation sample subset.
[0104] e m It is generated by the following formula:
[0105]
[0106] Here, N may be the number of samples included in the value transformation sample subset.
[0107] Z m It is generated by the following formula:
[0108]
[0109] Specifically, the derivation of the sample weight update formula is based on the sample label value of {-1,1}. For the business scenario of risk control pre-loan access, the label value is often {0,1}. Among them, "0" can represent a user with credit problems, and "1" can represent a user without credit problems. Therefore, the axis of symmetry is no longer "0", but 1 / 2. Therefore, the sample weight update formula needs to be adjusted. The adjusted formula is:
[0110]
[0111]
[0112] The specific parameter explanations are not repeated here.
[0113] The second step is to determine the sample weight set corresponding to the historical integrated model at the target position in the above historical integrated model sequence as the target sample weight set. Wherein, the historical integrated models in the historical integrated model sequence are arranged in order from early to late integration time, and the target position can be the first historical integrated model in the historical integrated model sequence. The historical integrated models in the historical integrated model sequence are arranged in order from late to early integration time, and the target position can be the last historical integrated model in the historical integrated model sequence.
[0114] The third step is to generate a sample weight set for the initial integrated model to be added based on the target sample weight set.
[0115] As an example, the execution subject may generate a sample weight set for the initial integrated model to be added based on the target sample weight set and using the sample weight update formula of Adaboost.
[0116] As another example, the execution entity may generate a sample weight set for the initial integrated model to be added based on the target sample weight set and using a sample weight update formula.
[0117] Step 204: Based on the value transformation sample subset and the sample weight set, the initial integrated model to be added is trained to obtain an integrated model.
[0118] In some embodiments, the execution subject may perform model training on the initial integrated model to be added according to the value transformation sample subset and the sample weight set to obtain an integrated model, wherein the integrated model may be a network model based on integrated learning after the model training is completed.
[0119] As an example, first, the execution entity may use the value transformation sample subset as a training data set and the sample weight set as a data weight set corresponding to the training data set to perform model training on the initial integration model to be added to obtain an integration model.
[0120] Step 205, integrating and fusing the above-mentioned integrated model with the pre-generated value detection information generation model to generate an integrated value detection information generation model.
[0121] In some embodiments, the above-mentioned execution entity may integrate the above-mentioned integration model with the pre-generated value detection information generation model to generate an integrated value detection information generation model. Among them, the value detection information generation model may be a model for generating value detection information. In practice, for the field of risk control, the execution task is a value lending credit assessment task, and the corresponding value detection information may be user credit value detection information. For the field of risk control, the execution task is a value lending credit assessment task, and the corresponding value detection information may be the user's value lending credit assessment information (loan credit assessment information).
[0122] In some optional implementations of some embodiments, after step 205, the steps further include:
[0123] The first step is to obtain the model prediction information of the above integrated value detection information generation model in the future target time window. In practice, the future target time window can be one month to six months after the time of the integrated value detection information generation model. The model prediction information can be the value detection information output by the above integrated value detection information generation model.
[0124] The second step is to generate model effect information corresponding to the integrated value detection information generation model based on the model prediction information. The model effect information can represent the model output effect (output accuracy) of the integrated value detection information generation model.
[0125] As an example, the above-mentioned execution entity can use the model prediction information as a data source according to the calculation logic of the model output effect, calculate the model output effect, and generate the model output effect.
[0126] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the model generation method of some embodiments of the present disclosure, without wasting computing resources, an integrated value detection information generation model can be obtained for accurately predicting samples with transformed sample distribution. Specifically, the reason for the inability to effectively solve the problem of sample distribution transformation is that it is not possible to adaptively solve the essence of the problem of sample distribution transformation, and only solve it by retraining the network model, resulting in an increase in training computing costs and a waste of resources. Based on this, the model generation method of some embodiments of the present disclosure first obtains a value transformation sample set for a sample set with a subsequent sample distribution transformation. Then, a value transformation sample subset with a transformed sample distribution in the above value transformation sample set is detected. Here, by screening out the value transformation sample subset with a transformed sample distribution, the sample distribution transformation problem can be effectively solved in the subsequent process, avoiding the situation where the value anomaly detection is inaccurate due to the sample distribution transformation problem. Next, the sample weight corresponding to each value transformation sample in the above value transformation sample subset for the initial integrated model to be added is determined to obtain a sample weight set. Here, by determining the sample weight corresponding to each value transformation sample, the initial integrated model to be added later is used for targeted model training, so that the initial integrated model can learn more characteristic information of the transformed sample distribution and generate more accurate value detection information. Furthermore, according to the above-mentioned value transformation sample subset and the above-mentioned sample weight set, the above-mentioned initial integrated model to be added is trained to obtain an integrated model that generates more accurate value detection information. Finally, the above-mentioned integrated model is integrated and fused with the pre-generated value detection information generation model to generate an integrated value detection information generation model. In summary, through the detection of the value transformation sample subset and the generation of the sample weight set, an integrated model that generates more accurate value detection information and can solve the sample distribution transformation problem in a targeted manner is trained. Therefore, through the integrated fusion of the integrated model and the value detection information generation model, not only can accurate value detection of the value transformation sample subset be guaranteed, but also accurate detection of the remaining value transformation sample sets can be guaranteed.
[0127] Further references Figure 3 , shows a process 300 of some other embodiments of the model generation method according to the present disclosure. The model generation method comprises the following steps:
[0128] Step 301, obtaining a value transformation sample set.
[0129] Step 302, detecting a value transformation sample subset in which the sample distribution in the value transformation sample set has been transformed.
[0130] Step 303 , determining the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtaining a sample weight set.
[0131] Step 304: Based on the value transformation sample subset and the sample weight set, the initial integrated model to be added is trained to obtain an integrated model.
[0132] In some embodiments, the specific implementation of steps 301-304 and the technical effects thereof can be referred to in Figure 2 The steps 201-204 in the corresponding embodiment are not described in detail here.
[0133] Step 305, determining the model diversity weight corresponding to each historical integrated model in the historical integrated model sequence and the model diversity weight corresponding to the integrated model, and obtaining a model diversity weight set.
[0134] In some embodiments, an execution entity (e.g. Figure 1 The electronic device 101 shown) can determine the model diversity weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model diversity weight corresponding to the above-mentioned integrated model to obtain a model diversity weight set. Among them, model diversity can characterize the generalization characteristics of the model. The better the model diversity, the stronger the generalization ability of the corresponding representation model. The model diversity weight can characterize the importance of the model in the subsequent integrated value detection information generation model. The model diversity weight can be a value between 0-1. The higher the model diversity weight, the more the corresponding model accounts for the output result in the subsequent integrated value detection information generation model.
[0135] As an example, the above execution entity can determine the model diversity weight by the following formula:
[0136]
[0137] Among them, div i It can represent the model diversity weight of the i-th model. E can be a fusion integration model set composed of the historical integration model sequence and the integration model. |E| can be the total number of models corresponding to the fusion integration model set. j It can be the jth model in the fusion ensemble model set. i It can be the i-th model in the fusion ensemble model set. ij It can be the diversity information between the j-th model and the i-th model.
[0138] div ij It can be generated by the following formula:
[0139]
[0140] Among them, ρ ij The model correlation between the j-th model and the i-th model can be characterized.
[0141] Step 306, determining the model discrimination weight corresponding to each historical integrated model in the historical integrated model sequence and the model discrimination weight corresponding to the integrated model, and obtaining a model discrimination weight set.
[0142] In some embodiments, the above-mentioned execution entity can determine the model discrimination weight corresponding to each historical integration model in the above-mentioned historical integration model sequence and the model discrimination weight corresponding to the above-mentioned integration model to obtain a model discrimination weight set. Among them, the model discrimination can characterize the discrimination of the model. The model discrimination weight can characterize the output proportion of the model in the value detection information generation model after subsequent integration. The higher the model discrimination weight, the higher the output proportion.
[0143] As an example, the above-mentioned execution entity can use the model discrimination measurement indicator-KS value to determine the model discrimination weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model discrimination weight corresponding to the above-mentioned integrated model to obtain a model discrimination weight set.
[0144] Step 307: Based on the model diversity weight set and the model discrimination weight set, the integrated model and the historical integrated model sequence are integrated and fused to generate an integrated value detection information generation model.
[0145] In some embodiments, the execution entity may perform model integration fusion on the integrated model and the historical integrated model sequence according to the model diversity weight set and the model discrimination weight set to generate an integrated value detection information generation model.
[0146] As an example, first, the execution entity may perform weighted summation of the model diversity weight set and the model discrimination weight set of the corresponding models to obtain a weighted sum weight set. Then, the weighted sum weight set is normalized to obtain a normalized value set. Finally, each normalized value in the normalized value set is used as the model weight of the corresponding integrated model to perform model integration fusion on the integrated model and the historical integrated model sequence to generate an integrated value detection information generation model.
[0147] As another example, the execution subject may then perform weighted summation of the weights of the corresponding models on the model diversity weight set and the model discrimination weight set to obtain a weighted summation weight set. Then, each weighted summation weight in the weighted summation weight set is used as the model weight of the corresponding integrated model to perform model integration fusion on the integrated model and the historical integrated model sequence to generate an integrated value detection information generation model.
[0148] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3 In the corresponding model generation method process 300 in some embodiments, by determining the model diversity weight and the model discrimination weight of the model, the model weight proportion of the model in all integrated models is considered from multiple angles, so that a more accurate integrated value detection information generation model can be output subsequently.
[0149] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a model generation device, and these device embodiments are Figure 2 Corresponding to the method embodiments shown, the model generation device can be specifically applied to various electronic devices.
[0150] like Figure 4 As shown, a model generation device 400 includes: an acquisition unit 401, a detection unit 402, a determination unit 403, a training unit 404 and an integration and fusion unit 405. The acquisition unit 401 is configured to acquire a value transformation sample set; the detection unit 402 is configured to detect a value transformation sample subset in which sample distribution in the value transformation sample set is transformed; the determination unit 403 is configured to determine the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtain a sample weight set; the training unit 404 is configured to perform model training on the initial integrated model to be added according to the value transformation sample subset and the sample weight set, and obtain an integrated model; the integration and fusion unit 405 is configured to integrate and fuse the integrated model with the pre-generated value detection information generation model to generate an integrated value detection information generation model.
[0151] In some optional implementations of some embodiments, the device 400 further includes: an information acquisition unit and a generation unit (not shown in the figure). The information acquisition unit may be configured to: acquire the model prediction information of the integrated value detection information generation model in the future target time window. The generation unit may be configured to: generate model effect information corresponding to the integrated value detection information generation model according to the model prediction information.
[0152] In some optional implementations of some embodiments, the above-mentioned value transformation sample set is stored in the form of data blocks; and the detection unit 402 can be further configured to: obtain at least one historical area sample division result corresponding to at least one target data block, wherein the target data block is a storage data block corresponding to the historical value transformation sample set; determine the current area sample division result of the data block corresponding to the above-mentioned value transformation sample set; for each historical area sample division result in the above-mentioned at least one historical area sample division result, determine the area sample division difference information between the above-mentioned historical area sample division result and the above-mentioned current area sample division result; generate difference area information based on the at least one area sample division difference information obtained; determine at least one value transformation sample in the data block corresponding to the above-mentioned value transformation sample set, whose corresponding area information is the difference area information, as a value transformation sample subset.
[0153] In some optional implementations of some embodiments, the detection unit 402 can be further configured to: determine the sample attribute set corresponding to the above-mentioned value transformation sample set; determine the first attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the above-mentioned value transformation sample set; filter out at least one sample attribute whose corresponding first attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set as at least one first sample attribute; filter out a first target sample attribute from the above-mentioned at least one first sample attribute; divide the above-mentioned value transformation sample set into a sample set according to the above-mentioned first target sample attribute to obtain at least one first sample subset; for each first sample subset in at least one first sample subset, perform the following determination steps: determine the second attribute division information corresponding to each sample attribute in the above-mentioned sample attribute set according to the first sample subset; filter out at least one sample attribute whose corresponding second attribute division information does not meet the attribute division condition from the above-mentioned sample attribute set attribute, as at least one second sample attribute; filter out a second target sample attribute from the at least one second sample attribute; divide the sample subset into sample sets according to the second target sample attribute to obtain at least one second sample subset; for each second sample subset in the at least one second sample subset, determine the third attribute division information corresponding to each sample attribute in the above sample attribute set for the above second sample subset; in response to determining that each third attribute division information in the at least one third attribute division information set obtained satisfies the corresponding attribute division condition, and the number of samples included in the second sample subset in the at least one second sample subset is less than a predetermined value, determine the sample division result corresponding to the at least one second sample subset as the sample division result corresponding to the first sample subset; and according to the at least one sample division result obtained, determine each area division result corresponding to each data area for the data block corresponding to the above value transformation sample set as the current area sample division result.
[0154] In some optional implementations of some embodiments, the detection unit 402 may be further configured to: in response to determining that there is third attribute division information that does not satisfy the corresponding attribute division condition in at least one third attribute division information set, and the number of samples included in the second sample subset in at least one second sample subset is greater than or equal to a predetermined value, determine a second sample subset group that does not satisfy the corresponding attribute division condition; determine the second sample subset group as at least one first sample subset, and continue to perform the above determination step.
[0155] In some optional implementations of some embodiments, the above-mentioned value detection information generation model includes: a historical integration model sequence; and the determination unit 403 can be further configured to: for the historical integration models in the above-mentioned historical integration model sequence, perform the following generation steps: in response to determining that there is a model in the above-mentioned historical integration model sequence whose corresponding model position is located before the above-mentioned historical integration model, determine the adjacent historical integration model that is located before the above-mentioned historical integration model and adjacent to the above-mentioned historical integration model; determine the sample weight corresponding to each value transformation sample in the above-mentioned value transformation sample subset for the above-mentioned adjacent historical integration model as the adjacent sample weight, and obtain the adjacent sample weight set; based on the above-mentioned adjacent sample weight set, generate the sample weight set for the above-mentioned historical integration model; determine the sample weight set corresponding to the historical integration model located at the target position in the above-mentioned historical integration model sequence as the target sample weight set; based on the above-mentioned target sample weight set, generate the sample weight set for the above-mentioned initial integration model to be added.
[0156] In some optional implementations of some embodiments, the above-mentioned value detection information generation model includes: a historical integrated model sequence; and the integration fusion unit 405 can be further configured to: determine the model diversity weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model diversity weight corresponding to the above-mentioned integrated model, and obtain a model diversity weight set; determine the model discrimination weight corresponding to each historical integrated model in the above-mentioned historical integrated model sequence and the model discrimination weight corresponding to the above-mentioned integrated model, and obtain a model discrimination weight set; according to the above-mentioned model diversity weight set and the above-mentioned model discrimination weight set, perform model integration fusion on the above-mentioned integrated model and the above-mentioned historical integrated model sequence to generate an integrated value detection information generation model.
[0157] It can be understood that the units recorded in the model generation device 400 are similar to those in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the model generation device 400 and the units included therein, and will not be described in detail here.
[0158] Reference below Figure 5 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of the electronic device 101)500. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0159] like Figure 5As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory 502 or a program loaded from a storage device 508 to a random access memory 503. Various programs and data required for the operation of the electronic device 500 are also stored in the random access memory 503. The processing device 501, the read-only memory 502, and the random access memory 503 are connected to each other via a bus 504. An input / output interface 505 is also connected to the bus 504.
[0160] Typically, the following devices may be connected to the input / output interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0161] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the read-only memory 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0162] It should be noted that the computer-readable medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0163] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0164] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a value transformation sample set; detects a value transformation sample subset in which the sample distribution in the value transformation sample set is transformed; determines the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtains a sample weight set; performs model training on the initial integrated model to be added according to the value transformation sample subset and the sample weight set, and obtains an integrated model; integrates the integrated model with the pre-generated value detection information generation model to generate an integrated value detection information generation model.
[0165] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0166] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0167] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The described units may also be set in a processor, for example, it may be described as: a processor includes an acquisition unit, a detection unit, a determination unit, and an integrated fusion unit. The names of these units do not constitute a limitation on the units themselves in some cases. For example, the acquisition unit may also be described as a "unit for acquiring a value transformation sample set."
[0168] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0169] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned model generation methods when executed by a processor.
[0170] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A model generation method, comprising: Get value transformation sample set; Detecting a value transformation sample subset in which sample distribution in the value transformation sample set has been transformed; Determine a sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtain a sample weight set; According to the value transformation sample subset and the sample weight set, model training is performed on the initial integrated model to be added to obtain an integrated model; The integrated model is integrated and fused with a pre-generated value detection information generation model to generate an integrated value detection information generation model.
2. The method according to claim 1, wherein: The method further comprises: Obtain model prediction information of the integrated value detection information generation model in a future target time window; Based on the model prediction information, model effect information corresponding to the integrated value detection information generation model is generated.
3. The method according to claim 1, wherein: The value transformation sample set is a sample set stored in the form of data blocks; as well as The detecting of the value transformation sample subset in which the sample distribution in the value transformation sample set is transformed comprises: Obtaining at least one historical area sample division result corresponding to at least one target data block, wherein the target data block is a storage data block corresponding to the historical value transformation sample set; Determine the current area sample division result of the data block corresponding to the value transformation sample set; For each historical area sample division result of the at least one historical area sample division result, determining area sample division difference information between the historical area sample division result and the current area sample division result; Divide the difference information according to the obtained at least one region sample to generate difference region information; At least one value transformation sample in the data block corresponding to the value transformation sample set, whose corresponding region information is the difference region information, is determined as a value transformation sample subset.
4. The method according to claim 3, wherein: The determining of the current area sample division result of the data block corresponding to the value transformation sample set includes: Determine a sample attribute set corresponding to the value transformation sample set; Determining first attribute division information corresponding to each sample attribute in the sample attribute set according to the value transformation sample set; At least one sample attribute whose corresponding first attribute division information does not satisfy the attribute division condition is selected from the sample attribute set as at least one first sample attribute; Filtering a first target sample attribute from the at least one first sample attribute; According to the first target sample attribute, dividing the value transformation sample set into sample sets to obtain at least one first sample subset; For each first sample subset in at least one first sample subset, the following determination steps are performed: Determine, according to the first sample subset, second attribute division information corresponding to each sample attribute in the sample attribute set; Filtering at least one sample attribute whose corresponding second attribute division information does not satisfy the attribute division condition from the sample attribute set as at least one second sample attribute; Filtering a second target sample attribute from at least one second sample attribute; Dividing the sample subsets according to the second target sample attribute to obtain at least one second sample subset; For each second sample subset in at least one second sample subset, determining third attribute division information corresponding to each sample attribute in the sample attribute set and for the second sample subset; In response to determining that each piece of third attribute division information in the obtained at least one third attribute division information set satisfies the corresponding attribute division condition, and the number of samples included in the second sample subset in the at least one second sample subset is less than a predetermined value, determining the sample division result corresponding to the at least one second sample subset as the sample division result corresponding to the first sample subset; According to the at least one sample division result obtained, each area division result corresponding to each data area of the data block corresponding to the value transformation sample set is determined as the current area sample division result.
5. The method according to claim 4, wherein: Before determining, based on the obtained at least one sample division result, each area division result corresponding to each data area for the data block corresponding to the value transformation sample set as the current area sample division result, the method further includes: In response to determining that there is third attribute division information that does not satisfy the corresponding attribute division condition in at least one third attribute division information set, and the number of samples included in the second sample subset in at least one second sample subset is greater than or equal to a predetermined value, determining a second sample subset group that does not satisfy the corresponding attribute division condition; The second sample subset group is determined as at least one first sample subset, and the determining step is continued.
6. The method according to claim 1, wherein: The value detection information generation model includes: a historical integration model sequence; and The determining of the sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added to obtain a sample weight set includes: For the historical integration model in the historical integration model sequence, the following generation steps are performed: In response to determining that there is a model in the historical integrated model sequence whose corresponding model position is located before the historical integrated model, determining an adjacent historical integrated model that is located before the historical integrated model and adjacent to the historical integrated model; Determine a sample weight corresponding to each value change sample in the value change sample subset for the adjacent historical integrated model as an adjacent sample weight, and obtain an adjacent sample weight set; Generating a sample weight set for the historical integration model according to the adjacent sample weight set; Determine a sample weight set corresponding to a historical integrated model located at a target position in the historical integrated model sequence as a target sample weight set; According to the target sample weight set, a sample weight set for the initial integrated model to be added is generated.
7. The method according to claim 1, wherein: The value detection information generation model includes: a historical integration model sequence; and The step of integrating the integrated model with a pre-generated value detection information generation model to generate an integrated value detection information generation model includes: Determine the model diversity weight corresponding to each historical integrated model in the historical integrated model sequence and the model diversity weight corresponding to the integrated model to obtain a model diversity weight set; Determine the model discrimination weight corresponding to each historical integrated model in the historical integrated model sequence and the model discrimination weight corresponding to the integrated model to obtain a model discrimination weight set; According to the model diversity weight set and the model discrimination weight set, the integrated model and the historical integrated model sequence are integrated and fused to generate an integrated value detection information generation model.
8. A model generation device, comprising: An acquisition unit, configured to acquire a value transformation sample set; A detection unit is configured to detect a value transformation sample subset in which a sample distribution in the value transformation sample set has been transformed; a determining unit configured to determine a sample weight corresponding to each value transformation sample in the value transformation sample subset for the initial integrated model to be added, and obtain a sample weight set; A training unit is configured to perform model training on the initial integrated model to be added according to the value transformation sample subset and the sample weight set to obtain an integrated model; The integration and fusion unit is configured to integrate and fuse the integration model with a pre-generated value detection information generation model to generate an integrated value detection information generation model.
9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.