Model generation method and device, electronic equipment and computer program product

By using a large language model and a phased training technique, the first training dataset is constructed and the target model is generated, which solves the problems of high model training complexity and high resource requirements, and achieves efficient and accurate prediction of multi-dimensional feature indicators.

CN120974191APending Publication Date: 2025-11-18INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511133380.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies involve high model training complexity, high computer performance and resource requirements, and the resulting models have low performance, making it difficult to meet business needs.

Method used

By responding to the training instructions of the large language model, the target sample data set and feature index set are extracted from the sample library to construct the first training dataset. The target model is then generated through phased training of the first and second preset models and knowledge distillation techniques.

Benefits of technology

This reduces the complexity of model training, decreases the requirements for computer performance and resources, improves the performance of the generated model, and ensures that the model can efficiently and accurately predict multi-dimensional feature indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974191A_ABST
    Figure CN120974191A_ABST
Patent Text Reader

Abstract

The invention discloses a model generation method and device, electronic equipment and a computer program product. Relates to the technical field of artificial intelligence, and the method comprises the steps: responding to model training instruction information, extracting a target sample data set from a sample library, and extracting a target feature index set from a feature index library, the model training instruction information being used for describing a model training target; for the target sample data set, determining a target feature value of each piece of target sample data under the target feature index set to obtain a target feature value subset, and determining each piece of target sample data and the corresponding target feature value subset as a piece of first training data to obtain a first training data set; and training a preset model according to the first training data set to obtain a target model. Through the method and the device, the problems of high complexity of model training, high requirements on computer performance and resources in the training process and low performance of the generated model in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a model generation method and device, electronic equipment and computer program product. BACKGROUND

[0002] Under the generation model, business personnel need to build a training data set and further train the model through the training data set. The complexity of model training is high, the training process has high requirements for computer performance and resources, and the performance of the generated model is not high, which is difficult to meet business needs.

[0003] For example, in the science and technology finance domain scenario, in the process of promoting science and technology finance work, it is necessary to form a precise evaluation of related institutions. However, in this scenario, the complexity of building a training data set and training a model is high, and due to factors such as industry differences, the model trained is difficult to accurately and intuitively evaluate related institutions.

[0004] In view of the high complexity of model training, the high requirements of the training process for computer performance and resources, and the low performance of the generated model in the related art, no effective solution has been proposed so far. SUMMARY

[0005] The main purpose of the present application is to provide a model generation method, device, electronic equipment and computer program product to solve the problem of high complexity of model training, high requirements of the training process for computer performance and resources, and low performance of the generated model in the related art.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a model generation method is provided. The method comprises: in response to model training instruction information, extracting a target sample data set from a sample library and extracting a target feature index set from a feature index library, wherein the model training instruction information is used to describe the model training target; for the target sample data set, determining the target feature value of each target sample data under the target feature index set to obtain a target feature value subset, and determining each target sample data and the corresponding target feature value subset as a first training data to obtain a first training data set; training a preset model according to the first training data set to obtain a target model.

[0007] Optionally, the preset model includes a first preset model and a second preset model, and training the preset model according to the first training data set to obtain the target model includes: training the first preset model according to the first training data set to obtain a first model; and determining the target model according to the training result of the first model and the second preset model.

[0008] Optionally, the first model includes M, and each first model is used to predict a characteristic index value of a target characteristic index in one dimension, and the target model is determined according to the training result of the first model and the second preset model includes: for the first training data set, a soft label of each target sample data in each first training data is obtained by the M first models respectively, to obtain a soft label set, and each target sample data and the corresponding soft label set are determined as a second training data, to obtain a second training data set, wherein a soft label contains a predicted characteristic value and a prediction probability of a target characteristic index by a first model; the second preset model is trained according to the second training data set, to obtain the target model.

[0009] Optionally, the first model includes M, and each first model is used to predict a characteristic index value of a target characteristic index in one dimension, and the second preset model is a fusion expression, and the target model is determined according to the training result of the first model and the second preset model includes: the output variable of each first model is determined as the training result of the first model; and the fusion expression of the output variables of the M first models is determined as the target model, wherein the fusion weight of each output variable is determined by the attribute of the target characteristic index in one dimension and the associated prediction probability.

[0010] Optionally, the first model is a time series prediction model, and is used to determine a predicted characteristic value and a prediction probability of a target characteristic index in multiple dimensions in a future time period, and the target model is determined according to the training result of the first model and the second preset model includes: for the first training data set, a soft label of each target sample data in each first training data is obtained by the first model, and each target sample data and the corresponding soft label are determined as a third training data, to obtain a third training data set, wherein a soft label contains a predicted characteristic value and a prediction probability of multiple target characteristic indexes in a future time period; the second preset model is trained according to the third training data set, to obtain the target model.

[0011] Optionally, the first model is a time series prediction model, and is used to determine a predicted characteristic value and a prediction probability of a target characteristic index in multiple dimensions in a future time period, and the second preset model is a fusion expression, and the target model is determined according to the training result of the first model and the second preset model includes: the output variable of the target characteristic index in multiple dimensions of the first model is determined as the training result of the first model; and the fusion expression of the output variables of the target characteristic index in multiple dimensions is determined as the target model, wherein the fusion weight of the output variable of each dimension is determined by the attribute of the target characteristic index in one dimension and the associated prediction probability.

[0012] Optionally, the preset model comprises a first preset model and a second preset model, the preset model is trained according to the first training data set, and obtaining the target model comprises: constructing a knowledge graph data according to the target sample data set, inputting the knowledge graph data into the first preset model, and obtaining an association relationship between each target sample data and the associated sample data; for the first training data set, determining a sample data sub-set according to a target sample data and the associated sample data in each first training data, determining a fusion feature value according to a target feature value under the sample data sub-set, determining each sample data sub-set and the corresponding fusion feature value as a fourth training data, and obtaining a fourth training data set; training the second preset model according to the fourth training data set to obtain the target model.

[0013] To achieve the above object, according to another aspect of the present application, a model generation device is provided. The device comprises: an extraction unit configured to extract a target sample data set from a sample library and a target feature index set from a feature index library in response to a model training instruction information, wherein the model training instruction information is used to describe a model training target; a determination unit configured to determine a target feature value of each target sample data in the target feature index set for the target sample data set, obtain a target feature value sub-set, and determine each target sample data and the corresponding target feature value sub-set as a first training data to obtain a first training data set; and a training unit configured to train a preset model according to the first training data set to obtain a target model.

[0014] To achieve the above object, according to another aspect of the present application, an electronic device is provided, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program performs any one of the model generation methods when running.

[0015] To achieve the above object, according to another aspect of the present application, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of any one of the model generation methods.

[0016] In the embodiment of the present application, the large language model response model training instruction information is adopted to extract the target sample data set from the sample library and extract the target feature index set from the feature index library, wherein the model training instruction information is used to describe the model training target; for the target sample data set, the target feature value of each target sample data under the target feature index set is determined to obtain a target feature value subset, each target sample data and the corresponding target feature value subset are determined as a first training data to obtain a first training data set; and the preset model is trained according to the first training data set to obtain a target model. The first training data set is constructed by the large language model response model training instruction information, and the preset model is obtained by training according to the first training data set, so as to reduce the model training complexity, thereby realizing the technical effect of reducing the requirement of the training process on the computer performance and resources and improving the performance of the generated model, and further solving the problems of high model training complexity, high requirement of the training process on the computer performance and resources, and low performance of the generated model. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings do not constitute an inappropriate limitation on the present application. In the drawings:

[0018] Figure 1 A hardware structure block diagram of a computer terminal for implementing the model generation method is shown;

[0019] Figure 2 A flowchart of the model generation method according to the embodiment of the present application is shown;

[0020] Figure 3 An optional model generation method according to the embodiment of the present application is shown;

[0021] Figure 4 A schematic diagram of a model generation device according to the embodiment of the present application is shown;

[0022] Figure 5 A structural block diagram of an electronic device according to the embodiment of the present application is shown. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present application are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal. For example, the system and the interface between the related users or institutions provide the user with a corresponding operation portal for the user to choose to agree or refuse the automatic decision result; if the user chooses to refuse, the expert decision process is entered. If the user chooses to agree, the user can view the data use purpose in real time through the authorization interface and has the right to withdraw authorization or delete data at any time. After withdrawing authorization, the system will terminate the related data processing within 24 hours.

[0026] Embodiment 1

[0027] According to the embodiments of the present application, a model generation method embodiment is also provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.

[0028] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the model generation method is shown. As shown in Figure 1 The computer terminal 10 (or mobile device) can include one or more (as shown in the figure) Figure 1(Illustrated using 102a, 102b, ..., 102n) Processor 102 (processor 102 may include, but is not limited to, a microprocessor (MCU) or a field-programmable gate array (FPGA) processing device), memory 104 for storing data, and transmission device 106 for communication functions. In addition, it may include: a display, input / output interface (I / O interface), Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), network interface, power supply, and / or camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0029] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned model generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0031] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal 10. In an example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0032] The display can be a touch screen liquid crystal display (LCD) that enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0033] In the above operating environment, the present application provides a model generation method as shown in Figure 2 Figure 2 is a flowchart of the model generation method according to an embodiment of the present application.

[0034] In step S202, a target sample data set is extracted from a sample library and a target feature index set is extracted from a feature index library in response to a model training instruction information, wherein the model training instruction information is used to describe a model training target.

[0035] It should be noted that the execution subject of the present embodiment can be a large language model or an intelligent agent service based on the large language model. The large language model determines a specific model training target by analyzing the model training instruction information, and then extracts corresponding samples and feature indexes from a pre-constructed database, thereby ensuring that the model training can focus on a specific technical field and improving the pertinence and efficiency of subsequent model training.

[0036] The model training instruction information can be specific training target description content input by a user. The training instruction information can describe what kind of model to train, for example, it can be "please construct an industry analysis model for 8 target technology industries".

[0037] The target sample data set refers to a group of samples selected according to the model training instruction information, which should be representative and diverse. For example, if the model training target is to train an industry analysis model, the target sample data set is a group of enterprise samples that can cover typical features and various possible situations of the target industry. If the model training target is to train a semiconductor industry analysis model, the target sample data set contains data of various enterprises in the semiconductor industry to fully reflect the characteristics of the industry.

[0038] ​The target feature index set is a multi-dimensional feature index selected according to the training target and the sample type in the target sample data set, and is used to describe and quantify the key attributes of the target sample. For example, if the model training target is to train a semiconductor industry analysis model, the target sample data can include "circuit design diagram submission quantity in the past 2 years", and the target feature index can include innovation capability, which can quantify the technical innovation capability of the semiconductor industry based on the target sample data.

[0039] In step S204, for the target sample data set, the target feature value of each target sample data under the target feature index set is determined to obtain a target feature value sub-set, each target sample data and the corresponding target feature value sub-set are determined as a first training data to obtain a first training data set.

[0040] It should be noted that the large language model performs feature engineering on the target sample, converts the target sample data set into specific numerical features, i.e., target feature values, which provides a data basis for constructing the first training data set.

[0041] In this process, a target sample data is mapped to a set of target feature values (i.e., a target feature value sub-set), and a target feature value sub-set is calculated based on the target feature index set for a target sample data. The target feature value sub-set constitutes an important part of the first training data, and the first training data set is a data set composed of all target sample data and their corresponding target feature value sub-sets, which ensures that the model can learn the relationship between sample data and feature values from the first training data set, thereby improving the efficiency and effectiveness of model training.

[0042] For example, a target sample data is the establishment time, number of employees, R&D investment proportion, number of patents, etc. of a technology enterprise, and the feature index can include enterprise basic situation, scientific and technological competitiveness, industrial development, ecological support, operating capacity, financial risk, etc. For example, for an enterprise that has been established for 10 years, if the total number of employees is more than 200 and the R&D investment proportion is 10%, the feature value of the enterprise basic situation feature index can be 4 points.

[0043] In step S206, the preset model is trained according to the first training data set to obtain a target model.

[0044] The preset model can be one or more models to be trained, and the target model is a trained model that can predict the index value of the to-be-tested sample data under the multi-dimensional feature index. For example, in the case of an industry analysis model as the target model, the index value of the target enterprise under the feature index of innovation capability, growth potential, financial health degree, etc. can be predicted.

[0045] Exemplarily, the preset model can be a teaching model, including a teacher model and a student model. Through a large number of training iterations, the internal parameters of the teacher model are gradually optimized and the prediction error is minimized, and the teacher model can generate a soft label after training, that is, a label based on the prediction probability of the teacher model. The student model is trained based on the soft label, and the trained student model inherits the prediction knowledge of the teacher model and can be independently applied to an actual scene. The target model refers to the student model after knowledge distillation training.

[0046] The model generation method provided by the embodiments of the present application extracts a target sample data set from a sample library and a target feature index set from a feature index library by responding to model training instruction information through a large language model, wherein the model training instruction information is used to describe the model training target; for the target sample data set, the target feature value of each target sample data under the target feature index set is determined to obtain a target feature value subset, each target sample data and the corresponding target feature value subset are determined as a first training data to obtain a first training data set; and a preset model is trained according to the first training data set to obtain a target model. By responding to the model training instruction information through the large language model to construct the first training data set and training the preset model according to the first training data set, the purpose of reducing the model training complexity is achieved, thereby realizing the technical effect of reducing the requirement of the training process on the computer performance and resources and improving the performance of the generated model, and further solving the problems of high model training complexity, high requirement of the training process on the computer performance and resources, and low performance of the generated model.

[0047] To reduce the requirement of the training process on the computer performance and resources and improve the performance of the target model, optionally, in the model generation method provided by the embodiments of the present application, the preset model includes a first preset model and a second preset model, and the target model is obtained by training the preset model according to the first training data set, including: training the first preset model according to the first training data set to obtain a first model; and determining the target model according to the training result of the first model and the second preset model.

[0048] The first preset model can be a first preset model, the first preset model can use a complex neural network or other deep learning architecture with strong learning ability, and can accept multi-modal information input. Its task is to deeply mine and learn the complex patterns and relationships implied in the first training data set after receiving the first training data set, so as to accurately predict the index values of the sample data under the multi-dimensional feature indexes as much as possible. For example, the sample is patent data, R&D investment, financial statements and other information from multiple sources of technology-based enterprises, which can predict the innovation ability, growth and financial stability of technology-based enterprises under the feature indexes.

[0049] In the training process, the first preset model aims to minimize the prediction error and generate a series of soft labels, i.e., an output containing a prediction probability distribution for each sample. The soft labels can be adjusted using temperature scaling techniques to ensure that they not only reflect the feature values of the sample under the feature indicators and eliminate the noise information covered by the original labels, but also contain subtle differences between similar samples and their degree of association with the feature indicators, which is conducive to the accurate understanding and learning of the student model.

[0050] The second preset model can be a student model, which can use a model with a simple structure, intuitive and easy to interpret, such as linear regression, logistic regression, score card, etc. Its training process learns and inherits the soft knowledge of the first preset model through knowledge distillation, with the help of its own structure and the soft label output of the first preset model. Specifically, the soft labels generated by the first preset model can be combined with the hard labels in the first training data set to generate a more detailed training data set. The second preset model is trained through this training data set, so that it not only focuses on accurate prediction, but also pays attention to the interpretability and efficiency of the model.

[0051] The target model refers to the second preset model after knowledge distillation training. Through knowledge distillation, the target model can maintain high efficiency reasoning while inheriting the prediction accuracy of the first preset model. In online service, the target model can quickly give analysis results with low computational resource consumption, while ensuring that the prediction quality does not decrease significantly.

[0052] The present embodiment realizes the effective transfer from the deep prediction ability of the first preset model to the high transparency and low operation cost of the second preset model through phased training and knowledge distillation technology. The final target model not only can accurately predict the index values of the sample data under the multi-dimensional feature indicators, but also outputs intuitive and clear results, which are easy for business personnel to understand and use.

[0053] To accurately predict the index values of the sample under multiple dimensions of feature indicators, optionally, in the model generation method provided in the present application, the first model includes M, and each first model is used to predict the feature indicator value of a target feature indicator of one dimension. According to the training result of the first model and the second preset model, the target model is determined as follows: for the first training data set, the soft labels of M first models for each target sample data in each first training data are obtained respectively to obtain a soft label set, and each target sample data and the corresponding soft label set are determined as a second training data to obtain a second training data set, wherein a soft label contains a predicted feature value and a prediction probability of a first model for a target feature indicator; the second preset model is trained according to the second training data set to obtain the target model.

[0054] In this embodiment, M first models are established, which can adopt advanced deep learning architectures (such as neural networks or ensemble learning models), and each first model is trained for a specific target feature indicator. The M first models can deeply mine and learn complex representations and patterns closely related to their respective target feature indicators from the first training data set. For example: innovation, growth potential, financial stability, market competitiveness, etc. For example, one first model can focus on evaluating the innovation of an enterprise, which is used to deeply analyze patent data, R&D investment, and R&D team size, etc. to predict the innovation ability, and another first model can focus on the growth potential of the enterprise, which is used to predict the future development prospects of the enterprise by analyzing the income growth rate, market expansion speed, etc.

[0055] After the training of the M first models is completed, for each target sample data in each first training data in the first training data set, the M first models respectively predict it to generate a soft label set, which contains the prediction results and their prediction probabilities of the feature values under the target feature indicators of each first model. For example, the first model includes a plurality of first models. For a technology-based enterprise, one first model can predict its innovation score to be 85 points with a prediction probability of 0.9, and another first model can predict the growth score to be 70 points with a prediction probability of 0.85, forming a soft label set containing two soft labels.

[0056] The second training data set is composed of each target sample data and its M soft label sets, which not only includes the original hard label information, but more importantly, incorporates the multi-dimensional knowledge and uncertainty measurement of the first model prediction, providing multi-level and multi-angle information for the training of the second model. Next, according to the constructed second training data set, the second preset model is trained. During the training process, the second model not only needs to learn to predict the hard label, but also needs to absorb and integrate the soft label set of the M first models through distillation learning to inherit the soft knowledge of the second model. Through learning and understanding the prediction knowledge of multiple dimensions, the second model can reflect the predicted feature values of the sample in multiple dimensions of the feature indicators (for example, the evaluation results of the enterprise in multiple aspects) in the final output.

[0057] In this embodiment, M first models are established, which can adopt advanced deep learning architectures (such as neural networks or ensemble learning models), and each first model is trained for a specific target feature indicator. The M first models can deeply mine and learn complex representations and patterns closely related to their respective target feature indicators from the first training data set. For example: innovation, growth potential, financial stability, market competitiveness, etc. For example, one first model can focus on evaluating the innovation of an enterprise, which is used to deeply analyze patent data, R&D investment, and R&D team size, etc. to predict the innovation ability, and another first model can focus on the growth potential of the enterprise, which is used to predict the future development prospects of the enterprise by analyzing the income growth rate, market expansion speed, etc.

[0058] In the case that the first model includes M, there are multiple ways to determine the target model. In the model generation method provided in the embodiments of the present application, the first model includes M, each first model is used to predict the characteristic index value of a dimension of target characteristic index, and the second preset model is a fusion expression. The target model is determined according to the training result of the first model and the second preset model, including: determining the output variable of each first model as the training result of the first model; and determining the fusion expression of the output variables of the M first models as the target model, wherein the fusion weight of each output variable is determined by the attribute of the target characteristic index of a dimension and the associated prediction probability.

[0059] In the embodiments, M first models are established. The M first models can deeply mine and learn the complex representation and pattern closely related to the respective target characteristic index from the first training data set, so as to predict the characteristic value of a specific dimension of target characteristic index in the sample. After the M first models are trained, the output of each model becomes the predicted characteristic value of a dimension and the prediction probability thereof. The output variables not only contain the score of a specific dimension of the sample, but also carry the confidence information of the model on the score, that is, the prediction probability.

[0060] The second preset model is a fusion expression. In order to generate a comprehensive target model, the fusion expression needs to assign weights to each output variable, combine the M output variables in a suitable manner, and these weights reflect the attribute of the target characteristic index and the credibility of the model prediction. For example, the first first model is used to predict the innovation ability of an enterprise. If the innovation ability is considered to be a very key dimension in the analysis of a technology-based enterprise, the output variable of the first first model will be assigned a higher weight. At the same time, the prediction probability of the model will also play a certain adjusting role in the weight determination, ensuring that the dimensions with more accurate predictions occupy a larger proportion in the final output result. The weight allocation method based on the prediction probability enables the target model to automatically identify and pay attention to the dimensions with more reliable information and more accurate predictions, thereby enhancing the accuracy and effectiveness of the output result of the target model.

[0061] The target model is a fusion expression that integrates the output variables of the M first models in the form of weighted fusion to form a unified scoring system. The specific fusion method can include weighted summation, weighted average, joint learning or other algorithms to ensure that the prediction results of each dimension are reasonably included in the output results of the target model according to their importance and prediction probability. For example, the target model is a scorecard with four dimensions of scores corresponding to "innovation", "growth potential", "financial condition" and "market competitiveness", and each dimension score and its prediction probability constitutes an output variable. In determining the target model, the following fusion expression can be used: [Comprehensive score = w1 x Innovation score + w2 x Growth potential score + w3 x Financial condition score + w4 x Market competitiveness score] where (w1) to (w4) are the fusion weights of each dimension, which are determined by considering the intrinsic properties of the target feature indicators and the prediction probabilities of the first models, ensuring that the scorecard can comprehensively and accurately reflect the performance of technology-based enterprises in multiple key areas, while also retaining the transparency and interpretability of each dimension score.

[0062] The target model is a fusion expression that integrates the output variables of the M first models in the form of weighted fusion to form a unified scoring system. The specific fusion method can include weighted summation, weighted average, joint learning or other algorithms to ensure that the prediction results of each dimension are reasonably included in the output results of the target model according to their importance and prediction probability. For example, the target model is a scorecard with four dimensions of scores corresponding to "innovation", "growth potential", "financial condition" and "market competitiveness", and each dimension score and its prediction probability constitutes an output variable. In determining the target model, the following fusion expression can be used: [Comprehensive score = w1 x Innovation score + w2 x Growth potential score + w3 x Financial condition score + w4 x Market competitiveness score] where (w1) to (w4) are the fusion weights of each dimension, which are determined by considering the intrinsic properties of the target feature indicators and the prediction probabilities of the first models, ensuring that the scorecard can comprehensively and accurately reflect the performance of technology-based enterprises in multiple key areas, while also retaining the transparency and interpretability of each dimension score.

[0063] To improve the timeliness of the target model analysis, optionally, in the model generation method provided in the present application, the first model is a time series prediction model for determining the predicted feature values and prediction probabilities of the target feature indicators in multiple dimensions in the future time period, and the target model is determined according to the training results of the first model and the second preset model. For the first training data set, the soft label of the first model for each target sample data in each first training data is obtained, and each target sample data and the corresponding soft label are determined as a third training data to obtain a third training data set, wherein a soft label contains predicted feature values and prediction probabilities of target feature indicators in a future time period; the second preset model is trained according to the third training data set to obtain the target model.

[0064] The first model in this embodiment can be a time series prediction model. By analyzing historical data, the time series prediction model can understand potential trends and cycles, thereby predicting the predicted feature values and their prediction probabilities of the target feature indicators of the sample in multiple dimensions in the future time period, providing forward-looking information for subsequent second model training. For example, the sample is a technology-based enterprise, and the predicted feature values of the target feature indicators in multiple dimensions in the future can be predicted, including but not limited to the innovation ability, growth, financial health, and market influence of the enterprise, etc.

[0065] After obtaining the prediction results of the target feature indicators of the first model in multiple dimensions in the future time period (i.e., soft labels), each target sample data is combined with its corresponding soft label to form a new third training data set, which not only retains the information of the original historical data, but also introduces the consideration of future prediction, providing a forward-looking learning data source for the second model. For example, for a technology-based enterprise, its soft label can include a predicted innovation ability of 5 points in the next year with a prediction probability of 0.75, and a predicted growth of 6 points with a prediction probability of 0.8. These soft labels together with the sample data of the enterprise constitute the third training data set.

[0066] Based on the constructed third training data set, the second preset model is continuously trained to obtain a target model. In this process, the second preset model learns how to interpret and use the soft labels generated by the first model, i.e., the predicted feature values and prediction probability information. The second preset model can be a simple and intuitive model that inherits the time series prediction ability of the first model through knowledge distillation technology, while maintaining its high interpretability and low computational cost. For example, the second preset model is an enterprise scorecard. In training, the predicted feature values in each soft label are fitted, and the uncertainty of the prediction probability is considered to optimize its prediction ability and adjust the feature weights. The final scorecard can independently predict the feature indicator values of the technology-based enterprise in multiple dimensions in the future and make a comprehensive score based on it.

[0067] This embodiment improves the prediction timeliness and accuracy of the target model by combining the forward-looking prediction ability of the time series prediction model and the efficient interpretability of the second preset model.

[0068] In the case that the first model is a time series prediction model, there are multiple ways to determine the target model. Optionally, in the model generation method provided in the embodiments of the present application, the first model is a time series prediction model, which is used to determine the predicted feature values and prediction probabilities of the target feature indicators in multiple dimensions in the future time period. The second preset model is a fusion expression. The target model is determined according to the training result of the first model and the second preset model. The output variables of the target feature indicators in multiple dimensions of the first model are determined as the training result of the first model. The fusion expression of the output variables of the target feature indicators in multiple dimensions is determined as the target model. The fusion weight of the output variable of each dimension is determined by the attribute of the target feature indicator of one dimension and the associated prediction probability.

[0069] The first model in the embodiments can be a time series prediction model. By analyzing historical data, the time series prediction model can understand potential trends and cycles, thereby predicting the predicted feature values and prediction probabilities of the target feature indicators in multiple dimensions of the sample in the future time period, to provide forward-looking information for subsequent second model training. The training result of the model is not only a simple prediction value, but also includes a prediction probability, reflecting the confidence of the model in the prediction accuracy.

[0070] After obtaining the prediction results of the target feature indicators in multiple dimensions by the first model, the training phase of the second preset model is entered. In order to generate a comprehensive target model, the fusion expression needs to assign weights to each predicted variable, combine multiple predicted variables according to certain rules, and form the target model. The fusion weight reflects the inherent attribute of the target feature indicator and the prediction probability associated therewith. The attribute of the target feature indicator in one dimension can include the historical volatility of the indicator and the correlation with enterprise performance. The prediction probability represents the confidence of the model in the prediction result. A higher prediction probability means that the model has stronger confidence in the prediction result. Through the fusion expression, the attention degree of the model to different dimensions can be adjusted according to the characteristics of each dimension and the confidence of the prediction, so as to obtain a more accurate and comprehensive comprehensive analysis result.

[0071] For example, the following fusion expression might be used as the target model: [\text{Comprehensive Score}=w_{\text{Finance}}\times(\text{Financial Forecast Value}\times p_{\text{Finance}})+w_{\text{Market}}\times(\text{Market Forecast Value}\times p_{\text{Market}})+w_{\text{Innovation}}\times(\text{Innovation Forecast Value}\timesp_{\text{Innovation}})] where (w_{\text{Finance}}), (w_{\text{Market}}), and (w_{\text{Innovation}}) are the fusion weights of the three dimensions of finance, market, and innovation, respectively, while (p_{\text{Finance}}), (p_{\text{Market}}), and (p_{\text{Innovation}}) are the predicted probabilities of the prediction results under these three dimensions.

[0072] This embodiment combines a time series prediction model with a fusion expression, enabling the target model to predict future trends based on historical data. Furthermore, by adjusting the weights of the prediction results for each dimension, it can accurately reflect the true state and expected development of the sample in different aspects, avoiding the bias or information gaps that may result from single-dimensional prediction and improving the performance of the target model.

[0073] To avoid the information limitations that may arise from relying solely on data from a single sample, optionally, in the model generation method provided in this application embodiment, the preset model includes a first preset model and a second preset model. Training the preset model based on the first training dataset to obtain the target model includes: constructing knowledge graph data based on the target sample data set, and inputting the knowledge graph data into the first preset model to obtain the association relationship between each target sample data and associated sample data; for the first training dataset, determining a sample data subset based on one target sample data and associated sample data in each first training dataset, determining a fusion feature value based on the target feature value under the sample data subset, and determining each sample data subset and the corresponding fusion feature value as a fourth training data to obtain a fourth training dataset; training the second preset model based on the fourth training dataset to obtain the target model.

[0074] First, based on the target sample data set, a knowledge graph data is constructed to describe the relationship between each sample data and other sample data, each sample data is an entity, and the relationship between sample data is an edge. Illustratively, the target sample data set is a set of enterprise sample data, and the constructed knowledge graph data depicts the association relationship between each target technology-based enterprise and other enterprises, including but not limited to supply chain relationship, competitor relationship, cooperation partner relationship, etc. The relationship between these entities is represented by a graph structure, each node represents an enterprise, and the edge represents the association between enterprises. The purpose of constructing the knowledge graph is to extract and understand the position and influence of technology-based enterprises in their industry ecosystem, providing a more comprehensive and detailed data perspective for the model.

[0075] Subsequently, the constructed knowledge graph data is input into the first preset model, which is a complex large-scale machine learning model such as a graph neural network, capable of processing structured graph data and unstructured enterprise information. The task of the first preset model is to analyze the association relationship between the nodes in the knowledge graph in depth to determine the complex relationship between each target sample data and the associated sample data, and how these relationships affect the feature values under the feature indicators of the sample. For example, for target technology enterprise A, the knowledge graph includes its direct relationship with supplier B, customer C, competitor D, etc. Through graph analysis, the first preset model can reveal that the stability of supplier B's supply chain directly affects the production efficiency and cost control of enterprise A, while the market share fluctuation of customer C is an important reference for the market sales forecast of enterprise A.

[0076] Based on the association relationship between the target sample data and the associated sample data, for each first training data in the first training data set, a sub-set containing the target sample data and its directly associated sample data is determined. Subsequently, according to the feature indicators of the target sample data and the associated sample data under the sub-set, a fused feature value is calculated, which integrates the multi-dimensional information of a target sample itself and its surrounding ecosystem, and can more comprehensively reflect the condition of the target sample.

[0077] For example, if the feature indicators of enterprise A include the number of patents, the proportion of R&D investment, etc., and the feature indicators of associated samples B, C, and D include their financial health status, market performance, etc. By analyzing the feature values of enterprise A and these associated enterprises under the feature indicators, a fused feature value (e.g., in the form of weighted sum) can be calculated, which not only reflects the independent performance of enterprise A, but also integrates the influence of its supply chain stability, market competitiveness, etc.

[0078] The fourth training data set is constructed based on the association relationship and the fusion feature value output by the first preset model, and provides a detailed data source containing the target sample and the ecosystem information for the second model training.

[0079] Next, the fourth training data set is used to train the second preset model, which can be a simple and easy-to-interpret model such as linear regression or score card. Through distillation learning, the core features of the target sample data are learned, and the complex interaction relationship between the target sample data and the associated sample data is also learned. For example, when analyzing a target enterprise, the average risk level of the industry group in which the enterprise is located or the performance of associated companies is obtained, and then the systemic risk and innovation impact of the enterprise are more accurately evaluated, ensuring that the target model can comprehensively consider the influence of the enterprise ecosystem when predicting the status of a technology-based enterprise.

[0080] The present embodiment, through the construction and analysis of the knowledge graph, enables the target model to comprehensively consider the status and mutual influence of the target sample and its associated samples in the knowledge graph, providing a more comprehensive and three-dimensional analysis perspective and improving the performance of the target model.

[0081] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0082] Embodiment 2

[0083] The present application also provides a model generation method, Figure 3 is a schematic diagram of an optional model generation method according to an embodiment of the present application, as Figure 3 shown, the method comprises:

[0084] Model input stage: description of the model training scene, such as: "please construct an industry analysis model for 8 hard technology industries". Sample library, enterprise sample library for subsequent model training, in the intelligent evaluation scene of technology-based enterprises, the sample library can store data of different types of enterprises. Feature index library, which can include indexes of technology-based enterprise scientific and technological competitiveness, industrial development, ecological support, operating capacity, financial risk, public opinion information, etc.

[0085] The training preparation stage: according to the description of the model training scene, the large model sequentially performs modeling scene understanding, modeling sample selection, and input modeling feature index selection. For example, when training a scientific and technological enterprise evaluation model suitable for the semiconductor industry, the large model will automatically select enterprises in the semiconductor industry and select more relevant feature indexes in the semiconductor industry for subsequent model training construction.

[0086] The model training stage: according to the modeling samples and input modeling indexes selected by the large model in the previous step, training is performed based on the teacher-student model training framework, and the final obtained simple model is converted into an expert scoring card. Among them, the teacher model adopts a complex large architecture, aims to maximize the prediction performance, can learn rich feature representations, and can be trained in multiple tasks to simultaneously output scores and probabilities under multiple indexes; the student model is designed as a small, lightweight and interpretable model, which can inherit the teacher's soft knowledge while maintaining interpretability and efficient reasoning. For example, the teacher model can be trained as a multi-layer neural network or a gradient boosting tree ensemble model to capture complex patterns; the student model adopts a linear or segmented rule form to output the final score.

[0087] The model deployment and adjustment stage: after the teacher model training is completed, the teacher model no longer participates in online services, and the student model independently runs, so that only low computing resources are required during reasoning. After the model training is completed, the large model verifies the index weight, and adjusts the weight of the student model according to the actual business demand to meet the actual situation, and adjusts the weight of the student model when the mathematical statistics result and the business experience do not match the index weight.

[0088] The embodiment takes the large model as the technical support and takes the teacher-student model as the core algorithm, considers factors such as industry differences, and can deploy a scientific and technological enterprise analysis model in one sentence, reduces the construction complexity of the analysis model, and improves the performance of the analysis model.

[0089] Embodiment 3

[0090] The embodiment of the application also provides a model generation device. It should be noted that the model generation device of the embodiment of the application can be used to execute the model generation method provided by the embodiment of the application. The model generation device provided by the embodiment of the application is introduced as follows.

[0091] According to the embodiment of the application, a device for implementing the above-mentioned model generation method is also provided, Figure 4 is a schematic diagram of the model generation device provided by the embodiment of the application, as Figure 4 shown, the device comprises:

[0092] The extraction unit 402 is configured to extract, in response to the model training instruction information, a target sample data set from a sample library and a target feature index set from a feature index library, where the model training instruction information is used to describe a model training target.

[0093] The determination unit 404 is configured to determine, for the target sample data set, a target feature value of each piece of target sample data under the target feature index set, obtain a target feature value sub-set, determine each piece of target sample data and the corresponding target feature value sub-set as a piece of first training data, and obtain a first training data set.

[0094] The training unit 406 is configured to train a preset model according to the first training data set, and obtain a target model.

[0095] The model generation apparatus provided by the embodiment of the present application comprises the extraction unit 402, the determination unit 404, and the training unit 406. The extraction unit 402 is configured to extract, in response to the model training instruction information, a target sample data set from a sample library and a target feature index set from a feature index library, where the model training instruction information is used to describe a model training target. The determination unit 404 is configured to determine, for the target sample data set, a target feature value of each piece of target sample data under the target feature index set, obtain a target feature value sub-set, determine each piece of target sample data and the corresponding target feature value sub-set as a piece of first training data, and obtain a first training data set. The training unit 406 is configured to train a preset model according to the first training data set, and obtain a target model. The model generation apparatus provided by the embodiment of the present application can solve the problems of high complexity of model training, high requirement of the training process on computer performance and resources, and low performance of the generated model in the related art. The first training data set is constructed by the large language model in response to the model training instruction information, and the preset model is trained according to the first training data set, so as to achieve the purpose of reducing the complexity of model training, thereby realizing the technical effects of reducing the requirement of the training process on computer performance and resources and improving the performance of the generated model, and further solving the problems of high complexity of model training, high requirement of the training process on computer performance and resources, and low performance of the generated model.

[0096] It should be noted that the above units and modules correspond to the steps in Embodiment 1, and the units and modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (for example, the memory 104) and processed by one or more processors (for example, the processors 102a, 102b, …, 102n), and the above modules can also be run in the computer terminal 10 provided in Embodiment 1 as part of the apparatus.

[0097] Embodiment 4

[0098] The embodiment of the present application can provide an electronic device,Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present application. As shown in the figure, the electronic device can include one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, wherein the peripheral interface is connected with a radio frequency module, an audio module, and a display. Figure 5 Figure 5 The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, i.e., implements the above-mentioned methods. The memory can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0099] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: in response to model training instruction information, extracting a target sample data set from a sample library and a target feature index set from a feature index library, wherein the model training instruction information is used to describe a model training target; for the target sample data set, determining a target feature value of each target sample data under the target feature index set to obtain a target feature value subset, determining each target sample data and the corresponding target feature value subset as a first training data to obtain a first training data set; training a preset model according to the first training data set to obtain a target model.

[0100] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: the preset model includes a first preset model and a second preset model, training the preset model according to the first training data set to obtain the target model includes: training the first preset model according to the first training data set to obtain a first model; determining the target model according to the training result of the first model and the second preset model.

[0101] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: the preset model includes a first preset model and a second preset model, training the preset model according to the first training data set to obtain the target model includes: training the first preset model according to the first training data set to obtain a first model; determining the target model according to the training result of the first model and the second preset model.

[0102] ​The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the first model includes M, and each first model is used to predict a characteristic index value of a target characteristic index in one dimension; and the target model is determined according to a training result of the first model and a second preset model, including: for a first training data set, soft labels of one target sample data in each first training data are obtained respectively by M first models to obtain a soft label set, and each target sample data and the corresponding soft label set are determined as a second training data to obtain a second training data set, wherein one soft label contains a predicted characteristic value and a prediction probability of one target characteristic index by one first model; and the second preset model is trained according to the second training data set to obtain the target model.

[0103] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the first model includes M, and each first model is used to predict a characteristic index value of a target characteristic index in one dimension; and the target model is determined according to a training result of the first model and a second preset model, including: the output variable of each first model is determined as the training result of the first model; and the fusion expression of the output variables of the M first models is determined as the target model, wherein the fusion weight of each output variable is determined by the attribute of the target characteristic index in one dimension and the associated prediction probability.

[0104] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the first model is a time series prediction model, and is used to determine predicted characteristic values and prediction probabilities of target characteristic indexes in multiple dimensions in a future time period; and the target model is determined according to a training result of the first model and a second preset model, including: for a first training data set, soft labels of one target sample data in each first training data are obtained by the first model, and each target sample data and the corresponding soft label are determined as a third training data to obtain a third training data set, wherein one soft label contains predicted characteristic values and prediction probabilities of target characteristic indexes in a future time period; and the second preset model is trained according to the third training data set to obtain the target model.

[0105] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the first model is a time series prediction model, used to determine predicted characteristic values and prediction probabilities of the plurality of target characteristic indicators in a future time period; and the second preset model is a fusion expression, and the target model is determined according to the training result of the first model and the second preset model, including: determining the output variables of the plurality of dimensions of the target characteristic indicators in the first model as the training result of the first model; and determining the fusion expression of the output variables of the plurality of dimensions of the target characteristic indicators as the target model, wherein the fusion weight of each dimension of the output variables is determined by the attribute of the target characteristic indicator of one dimension and the associated prediction probability.

[0106] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: the preset model includes a first preset model and a second preset model, the preset model is trained according to the first training data set to obtain the target model, including: constructing a knowledge graph data according to the target sample data set, and inputting the knowledge graph data into the first preset model to obtain the association relationship between each target sample data and the associated sample data; for the first training data set, determining a sample data sub-set according to a target sample data and associated sample data in each first training data, determining a fusion characteristic value according to the target characteristic value under the sample data sub-set, and determining each sample data sub-set and the corresponding fusion characteristic value as a fourth training data to obtain a fourth training data set; and training the second preset model according to the fourth training data set to obtain the target model.

[0107] Those skilled in the art can understand that, Figure 5 The structure shown is only schematic, and the electronic device can also be a smart phone, a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD (tablet computer), and the like. Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0108] Those skilled in the art can understand that all or part of the steps in the various methods of the above-mentioned embodiments can be instructed by a program to complete the related hardware of the terminal device, and the program can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.

[0109] Embodiment 5

[0110] The embodiment of the present application further provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the model generation method provided in the first embodiment.

[0111] Optionally, in the embodiment, the storage medium can be located in any one of computer terminals in a computer terminal group in a computer network, or in any one of mobile terminals in a mobile terminal group.

[0112] The present application further provides a computer program product adapted to execute the steps of the model generation method when executed on a data processing device.

[0113] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0114] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0115] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the embodiment described above is only a schematic, for example, the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.

[0116] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.

[0117] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0118] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk and various media that can store program codes.

[0119] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for generating a model, characterized in that, include: In response to model training instruction information, a target sample data set is extracted from the sample library, and a target feature index set is extracted from the feature index library, wherein the model training instruction information is used to describe the model training objective; For the target sample data set, determine the target feature value of each target sample data under the target feature index set to obtain a target feature value subset. Determine each target sample data and the corresponding target feature value subset as a first training data to obtain the first training dataset. The target model is obtained by training a preset model based on the first training dataset.

2. The method according to claim 1, characterized in that, The preset model includes a first preset model and a second preset model. The target model is obtained by training the preset model based on the first training dataset. The first preset model is trained based on the first training dataset to obtain the first model; The target model is determined based on the training results of the first model and the second preset model.

3. The method according to claim 2, characterized in that, The first model comprises M models, each used to predict the feature index value of a target feature index in one dimension. The target model is determined based on the training results of the first model and the second preset model, including: For the first training dataset, obtain the soft labels of one target sample data in each first training dataset for each of the M first models, and obtain a soft label set. Then, determine each target sample data and the corresponding soft label set as a second training dataset to obtain the second training dataset. Here, a soft label contains the predicted feature value and predicted probability of a target feature index by a first model. The second preset model is trained based on the second training dataset to obtain the target model.

4. The method according to claim 2, characterized in that, The first model comprises M models, each used to predict the feature value of a target feature indicator in one dimension. The second preset model is a fusion expression. The target model is determined based on the training results of the first models and the second preset model, including: The output variable of each first model is determined as the training result of the first model; The target model is determined by the fusion expression of the output variables of the M first models, wherein the fusion weight of each output variable is determined by the attribute of a target feature index in one dimension and the associated prediction probability.

5. The method according to claim 2, characterized in that, The first model is a time series prediction model used to determine the predicted feature values ​​and predicted probabilities of target feature indicators across multiple dimensions for future time periods. The target model is determined based on the training results of the first model and the second preset model, including: For the first training dataset, the soft label of one target sample data in each first training data is obtained by the first model, and each target sample data and its corresponding soft label are determined as a third training data to obtain the third training dataset. Here, a soft label contains the predicted feature value and predicted probability under multiple target feature indicators in the future time period. The second preset model is trained based on the third training dataset to obtain the target model.

6. The method according to claim 2, characterized in that, The first model is a time series prediction model used to determine the predicted feature values ​​and prediction probabilities for multiple target feature indicators in the future time period. The second preset model is a fusion expression, and the target model is determined based on the training results of the first model and the second preset model, including: The output variables of the first model under the target feature indicators of multiple dimensions are determined as the training results of the first model; The fusion expression of the output variables under the target feature indicators of the multiple dimensions is determined as the target model, wherein the fusion weight of the output variable of each dimension is determined by the attribute of the target feature indicator of one dimension and the associated prediction probability.

7. The method according to claim 1, characterized in that, The preset model includes a first preset model and a second preset model. The target model is obtained by training the preset model based on the first training dataset. A knowledge graph is constructed based on the target sample data set, and the knowledge graph is input into the first preset model to obtain the association relationship between each target sample data and related sample data. For the first training dataset, a subset of sample data is determined based on a target sample data and associated sample data in each first training dataset. A fusion feature value is determined based on the target feature value under the subset of sample data. Each subset of sample data and the corresponding fusion feature value are determined as a fourth training dataset to obtain the fourth training dataset. The second preset model is trained based on the fourth training dataset to obtain the target model.

8. A model generation apparatus, characterized in that, include: An extraction unit is used to respond to model training instruction information, extract a target sample data set from a sample library, and extract a target feature index set from a feature index library, wherein the model training instruction information is used to describe the model training objective; The determining unit is used to determine the target feature value of each target sample data under the target feature index set for the target sample data set, obtain a subset of target feature values, and determine each target sample data and the corresponding subset of target feature values ​​as a first training data to obtain a first training dataset; The training unit is used to train a preset model based on the first training dataset to obtain the target model.

9. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 7.