Business evaluation model iterative optimization method, business evaluation method, and pipeline system

By continuously collecting and optimizing enterprise characteristic data, and combining iterative optimization of enterprise evaluation models with domain expert evaluation rules, the problems of model accuracy and consistency on big data platforms have been solved, and efficient application of enterprise evaluation models has been achieved.

CN115759810BActive Publication Date: 2026-08-04BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2022-10-24
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing enterprise assessment models suffer from accuracy and consistency issues on big data platforms and are difficult to optimize through automation, resulting in lagging training data and data drift, which fails to guarantee the accuracy and effectiveness of enterprise assessments.

Method used

By continuously collecting enterprise characteristic data, processing and optimizing it, and storing it in the target database, a dataset is formed for model training and validation. The model is then evaluated and deployed in conjunction with evaluation rules set by domain experts, and feedback is used to correct the data to achieve iterative optimization of the model.

Benefits of technology

It has enabled automated iterative optimization of enterprise evaluation models, improved the accuracy and efficiency of model application, reduced the difficulty of involving domain experts, and enhanced the accuracy and effectiveness of enterprise evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115759810B_ABST
    Figure CN115759810B_ABST
Patent Text Reader

Abstract

The application provides an enterprise evaluation model iterative optimization method, an enterprise evaluation method and a pipeline system. The optimization method comprises the following steps: in the current iterative optimization process of the enterprise evaluation model, the data set and the target version of the enterprise evaluation model are evaluated according to the evaluation rule data set by the domain expert in advance, and the enterprise state evaluation result is output to enable the domain expert to judge whether the corresponding enterprise feature data is corrected according to the enterprise state evaluation result; the enterprise feature data corrected by the domain expert and the enterprise feature data corresponding to the correct enterprise state evaluation result are both stored in the optimized data set to be used for the iterative optimization process of the enterprise evaluation model of the next data version. The application can effectively improve the application accuracy and optimization efficiency of the enterprise evaluation model, and can reduce the difficulty of the domain expert to participate, thereby improving the accuracy and effectiveness of the enterprise state evaluation by using the enterprise evaluation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to iterative optimization methods for enterprise evaluation models, enterprise evaluation methods, and pipeline systems. Background Technology

[0002] Enterprise valuation is crucial in business decision-making and in determining enterprise value for third parties. In real-world economic situations, enterprises are often transferred or merged as a whole, such as mergers, acquisitions, sales, restructuring, joint ventures, equity operations, and guarantees. All of these involve assessing the overall value of the enterprise. In such cases, it's necessary to evaluate the enterprise's development stage, risk profile, and growth potential to determine the price for the joint venture or resale. Machine learning techniques can be used to enhance the automation of enterprise valuation. However, when using enterprise valuation models integrated into big data platforms, the accuracy and consistency often fall short of laboratory results. This is because big data platforms are characterized by large data volumes, rapid data generation and processing, and real-time requirements. This can lead to data lag during the integration of the enterprise valuation model with the platform, resulting in a model that performs well on the original dataset at the time of delivery but performs poorly on newly collected enterprise-specific data.

[0003] Currently, solutions for enterprise evaluation using enterprise assessment models primarily address accuracy and consistency issues by increasing model size. However, these solutions require significant time and resources from experienced machine learning professionals. Furthermore, even with manually adjusted models, data lag and drift issues persist as new data accumulates over time. Moreover, with the increasing volume and variety of enterprise feature data, the initially trained enterprise assessment model becomes incapable of utilizing novel data never observed before model generation and collection.

[0004] In other words, the existing enterprise evaluation process lacks an intelligent optimization mechanism for the enterprise evaluation model, resulting in low efficiency in model optimization, lagging training data, and difficulty in involving domain experts in the optimization process. Therefore, the accuracy and effectiveness of enterprise evaluation cannot be guaranteed. Summary of the Invention

[0005] In view of this, embodiments of this application provide an iterative optimization method for enterprise evaluation models, an enterprise evaluation method, and a pipeline system to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of this application provides an iterative optimization method for a business evaluation model, comprising:

[0007] Continuously collect raw enterprise characteristic data, and after determining that the iterative optimization process for the enterprise evaluation model has started, process and optimize the raw enterprise characteristic data to obtain corresponding enterprise characteristic data, and store the enterprise characteristic data in the target database;

[0008] Enterprise feature data is extracted from the target database to form the dataset used for training and validating the enterprise evaluation model. Based on this dataset, the previous data version of the enterprise evaluation model is trained and optimized to obtain the target version of the enterprise evaluation model after training.

[0009] The dataset and the target version of the enterprise evaluation model are evaluated according to the evaluation rules data set in advance by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, the target version of the enterprise evaluation model is deployed online, and the enterprise feature data in the target database is input into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the corresponding enterprise status evaluation result.

[0010] The system stores and outputs the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results. If the system receives corrected enterprise feature data from the domain expert or a message indicating that the enterprise status assessment results are correct, both the corrected enterprise feature data from the domain expert and the enterprise feature data corresponding to the correct enterprise status assessment results are stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version.

[0011] In some embodiments of this application, the continuous collection of raw enterprise characteristic data, and after determining that the iterative optimization process for the enterprise evaluation model has started, processing and optimizing the raw enterprise characteristic data to obtain corresponding enterprise characteristic data, and storing the enterprise characteristic data in the target database, includes:

[0012] Original enterprise feature data is collected based on a preset periodic update strategy and / or dynamic update strategy, and the original enterprise feature data is preprocessed and stored in the corresponding Kafka data topic unit.

[0013] After determining that the iterative optimization process for the enterprise evaluation model has been initiated, the original enterprise characteristic data is obtained from the Kafka data topic unit and processed to obtain the corresponding enterprise characteristic data. The enterprise characteristic data is then stored in the target database, and the data changes in the target database are monitored. The target database includes a Hive database.

[0014] In some embodiments of this application, the step of evaluating the dataset and the trained enterprise evaluation model according to evaluation rules pre-set by domain experts, and if the dataset and the trained enterprise evaluation model pass the evaluation, then deploying the enterprise evaluation model online, and inputting the enterprise feature data from the target database into the online deployed enterprise evaluation model so that the enterprise evaluation model outputs the corresponding enterprise status evaluation result, includes:

[0015] The dataset is evaluated according to the evaluation rules set in advance by domain experts, and the datasets that pass the evaluation are saved to the MySQL database; and the performance of the trained enterprise evaluation model is evaluated, and the trained enterprise evaluation model is validated according to the evaluation rules set in advance by domain experts and the validation set, and the parameter information and performance information corresponding to the enterprise evaluation model that passes the performance evaluation and the validation are saved to the MySQL database.

[0016] The enterprise assessment model is deployed on a server, and enterprise feature data is extracted from the Hive database. Each enterprise feature data is then input into the enterprise assessment model deployed on the server, so that the enterprise assessment model outputs enterprise status assessment results corresponding to each enterprise feature data. Each enterprise status assessment result is stored in the Hive database, and the correspondence between the enterprise feature data and the enterprise status assessment results is stored in a preset model calculation result index.

[0017] In some embodiments of this application, the enterprise evaluation model includes: Gradient Boosting Decision Tree (GBDT) model;

[0018] The enterprise assessment model is used to output corresponding enterprise status assessment results based on the input enterprise characteristic data. The enterprise status assessment results include: enterprise development stage prediction results, enterprise growth assessment results, or enterprise risk assessment results.

[0019] The predicted enterprise development stages include: seed stage, startup stage, growth stage, expansion stage, maturity stage, or decline stage.

[0020] Another aspect of this application provides a business evaluation methodology, including:

[0021] Receive enterprise evaluation model parameters and store the target version of the enterprise evaluation model corresponding to the enterprise evaluation model parameters locally, wherein the target version of the enterprise evaluation model is generated in advance based on the enterprise evaluation model iterative optimization method.

[0022] Obtain enterprise characteristic data of the target company;

[0023] The enterprise characteristic data is input into the target version of the enterprise assessment model so that the target version of the enterprise assessment model outputs the enterprise status assessment result of the target enterprise, wherein the enterprise status assessment result includes: enterprise development stage prediction result, enterprise growth assessment result, or enterprise risk assessment result.

[0024] Another aspect of this application provides an iterative optimization pipeline system for enterprise evaluation models, comprising:

[0025] The data acquisition and processing module is used to continuously collect raw enterprise characteristic data, and after determining that the iterative optimization process for the enterprise evaluation model has started, it processes and optimizes the raw enterprise characteristic data to obtain the corresponding enterprise characteristic data, and stores the enterprise characteristic data in the target database.

[0026] The model training and optimization module is used to extract enterprise feature data from the target database to form the dataset used for training and validating the enterprise evaluation model, and to train and optimize the previous data version of the enterprise evaluation model based on the dataset to obtain the target version of the enterprise evaluation model after training.

[0027] The model evaluation and deployment module is used to evaluate the dataset and the target version of the enterprise evaluation model according to the evaluation rules data set in advance by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, the target version of the enterprise evaluation model is deployed online, and the enterprise feature data in the target database is input into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the corresponding enterprise status evaluation result.

[0028] The model result feedback module is used to store and output the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results. If the domain expert corrects the enterprise feature data or a message indicating that the enterprise status assessment result is correct is received, the domain expert corrects the enterprise feature data and the enterprise feature data corresponding to the correct enterprise status assessment result are both stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version.

[0029] In some embodiments of this application, the data acquisition and processing module includes:

[0030] The data acquisition module is used to collect raw enterprise feature data based on a preset periodic update strategy and / or dynamic update strategy, preprocess the raw enterprise feature data, and store the preprocessed raw enterprise feature data in the corresponding Kafka data topic unit.

[0031] The data processing module is used to obtain raw enterprise feature data from the Kafka data topic unit and process the data after determining that the iterative optimization process for the enterprise evaluation model has started, to obtain the corresponding enterprise feature data, and to store the enterprise feature data in the target database, and to monitor the data changes in the target database, wherein the target database includes: a Hive database.

[0032] In some embodiments of this application, the model evaluation and deployment module includes:

[0033] The model evaluation module is used to evaluate the dataset according to the evaluation rules data set in advance by domain experts, and save the dataset that passes the evaluation to the MySQL database; and to evaluate the performance of the trained enterprise evaluation model, and to verify the trained enterprise evaluation model according to the evaluation rules data set in advance by domain experts and the validation set, and to save the parameter information and performance information of the enterprise evaluation model that passes the performance evaluation and the validation to the MySQL database.

[0034] The model deployment and calculation module is used to deploy the enterprise evaluation model to the server, extract enterprise feature data from the Hive database, input each enterprise feature data into the enterprise evaluation model deployed on the server, so that the enterprise evaluation model outputs the enterprise status evaluation results corresponding to each enterprise feature data, stores each enterprise status evaluation result in the Hive database, and stores the correspondence between the enterprise feature data and the enterprise status evaluation results in a preset model calculation result index.

[0035] Another aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the iterative optimization method for the enterprise evaluation model, or to implement the enterprise evaluation method.

[0036] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the iterative optimization method for the enterprise evaluation model, or implements the enterprise evaluation method.

[0037] The enterprise evaluation model iterative optimization method provided in this application continuously collects original enterprise feature data. After determining that the iterative optimization process for the current enterprise evaluation model has started, the original enterprise feature data is processed and optimized to obtain corresponding enterprise feature data, which is then stored in a target database. Enterprise feature data is extracted from the target database to form a dataset used for training and validating the enterprise evaluation model. Based on this dataset, the previous data version of the enterprise evaluation model is trained and optimized to obtain a trained target version of the enterprise evaluation model. The dataset and the target version of the enterprise evaluation model are evaluated according to evaluation rules pre-set by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, the target version of the enterprise evaluation model is deployed online, and the enterprise feature data from the target database is input into the target database. A version of the enterprise assessment model is provided to output corresponding enterprise status assessment results. The model stores and outputs the correspondence between the enterprise feature data and the enterprise status assessment results, allowing domain experts to determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results. If corrected enterprise feature data or a message indicating that the enterprise status assessment result is correct is received from a domain expert, both the corrected enterprise feature data and the enterprise feature data corresponding to the correct enterprise status assessment result are stored in the optimized dataset for iterative optimization of the next data version of the enterprise assessment model. This enables automated iterative optimization of the enterprise assessment model, effectively improving the application accuracy and optimization efficiency of the enterprise assessment model, reducing the difficulty for domain experts to participate, and thus improving the accuracy and effectiveness of applying the enterprise assessment model for enterprise status assessment.

[0038] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0039] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0040] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings:

[0041] Figure 1 This is a schematic diagram of the overall process of the iterative optimization method for the enterprise evaluation model in one embodiment of this application.

[0042] Figure 2 This is a schematic diagram illustrating a specific process of the iterative optimization method for the enterprise evaluation model in one embodiment of this application.

[0043] Figure 3 This is a schematic diagram of the structure of the enterprise evaluation model iterative optimization pipeline system in another embodiment of this application.

[0044] Figure 4 This is a schematic diagram of a specific structure of an iterative optimization pipeline system for enterprise evaluation model in another embodiment of this application.

[0045] Figure 5 This is a flowchart illustrating the enterprise evaluation method in another embodiment of this application.

[0046] Figure 6 This is a schematic diagram illustrating an example of the data iteration process provided in the application examples of this application.

[0047] Figure 7 This is a schematic diagram illustrating an example of a model iteration pipeline provided in the application examples of this application.

[0048] Figure 8 This is a schematic diagram illustrating an example of the pipeline architecture provided in the application examples of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0050] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the scheme according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0051] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0052] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0053] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0054] To address the issues of low efficiency in enterprise evaluation model construction and optimization, outdated training data, and difficulty in involving domain experts in the optimization process, thus compromising the accuracy and effectiveness of enterprise evaluation, this application provides an iterative optimization method for enterprise evaluation models that can be executed by an iterative optimization pipeline system. (See [link to relevant documentation]). Figure 1 The iterative optimization method for the enterprise evaluation model specifically includes the following:

[0055] Step 100: Continuously collect raw enterprise feature data, and after determining that the iterative optimization process for the enterprise evaluation model has started, process and optimize the raw enterprise feature data to obtain the corresponding enterprise feature data, and store the enterprise feature data in the target database.

[0056] In one or more embodiments of this application, continuous data collection refers to periodic or real-time data collection, specifically determined based on the data type or actual application requirements. The domain expert refers to a professional with experience in enterprise evaluation.

[0057] The original enterprise feature data refers to the enterprise feature data obtained after collection and preprocessing. The enterprise feature data can be determined according to the purpose of the enterprise evaluation model. For example, if the enterprise evaluation model is a machine learning model used to predict the enterprise development stage, then the enterprise feature data will use feature data related to the prediction of the enterprise development stage, such as: enterprise registered capital, number of patents, number of social security employees, number of court hearings, and duration of existence.

[0058] In one or more embodiments of this application, the iterative optimization process for the enterprise evaluation model refers to the iterative optimization process for the enterprise evaluation model that is started on a timed basis or according to a trigger condition. The specific implementation method for determining the start of the iterative optimization process for the enterprise evaluation model can be: receiving a message to start the iterative optimization process for the enterprise evaluation model, obtaining the enterprise evaluation model of the previous data version obtained in the previous iterative optimization process for the enterprise evaluation model according to the message, and adding 1 to the previous round number value N as the round number value N+1 of the iterative optimization process for the enterprise evaluation model.

[0059] Step 200: Extract enterprise feature data from the target database to form the dataset used for training and validating the enterprise evaluation model. Based on this dataset, train and optimize the enterprise evaluation model of the previous data version to obtain the target version of the enterprise evaluation model after training.

[0060] In step 200, the dataset required for model training can be extracted from the Hive database using the Spark computing engine. Then, the optimal hyperparameters of the model on this dataset are found using the GridSearchCV grid search algorithm and cross-validation, and the positive and negative samples are adjusted to finally complete the construction of the enterprise evaluation model.

[0061] In one or more embodiments of this application, the target version of the enterprise evaluation model refers to the enterprise evaluation model trained in the iterative optimization process of the enterprise evaluation model. If the identifier of the previous version of the enterprise evaluation model is denoted as M, then the identifier of the target version of the enterprise evaluation model can be denoted as M+1.

[0062] The enterprise evaluation model refers to a machine learning model used to predict or evaluate the state of an enterprise, such as the Gradient Boosting Decision Tree (GBDT) model.

[0063] Step 300: Evaluate the dataset and the target version of the enterprise evaluation model according to the evaluation rules data set in advance by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, deploy the target version of the enterprise evaluation model online and input the enterprise feature data in the target database into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the corresponding enterprise status evaluation result.

[0064] In step 300, the evaluation rule data pre-set by the domain expert can refer to: the evaluation rule data pre-received by the enterprise evaluation model iterative optimization pipeline system from the client device held by the domain expert, which includes at least evaluation data for the dataset and for the enterprise evaluation model.

[0065] It is understood that the example of deploying the target version of the enterprise assessment model online refers to sending the model parameters of the target version of the enterprise assessment model to a server (e.g., a server specifically used to demonstrate calculation results), and conducting enterprise assessment on that server using the current target version of the enterprise assessment model.

[0066] Step 400: Store and output the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results; if the domain expert corrects the enterprise feature data or a message is received to indicate that the enterprise status assessment results are correct, then both the domain expert corrected enterprise feature data and the enterprise feature data corresponding to the correct enterprise status assessment results are stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version.

[0067] In one or more embodiments of this application, outputting the enterprise status assessment results may refer to directly sending each corresponding enterprise status assessment result and each enterprise characteristic data to a client device held by a domain expert, so that the domain expert can view these enterprise status assessment results on their client device and determine whether to correct the corresponding enterprise characteristic data.

[0068] In addition, outputting the enterprise status assessment results can also refer to directly sending each corresponding enterprise status assessment result and each enterprise characteristic data to a preset display screen for display, so that domain experts can view these enterprise status assessment results through the display screen and determine whether to correct the corresponding enterprise characteristic data.

[0069] In step 400, if the domain expert corrected enterprise characteristic data or a message used to feedback that the enterprise status assessment result is correct is received, specifically, if the enterprise assessment model iterative optimization pipeline system receives the domain expert corrected enterprise characteristic data or a message used to feedback that the enterprise status assessment result is correct sent by the client device.

[0070] For example, the results of Elasticsearch model calculations can be indexed to provide applications with a display of model calculation results. Domain experts can monitor the model calculation results through the front-end application to confirm whether the calculation results are correct and whether they need to be corrected. Validated or corrected data will be exported to form a new training dataset and supplemented in the next data iteration.

[0071] As can be seen from the above description, the enterprise evaluation model iterative optimization method provided in this application embodiment can realize automated iterative optimization of the enterprise evaluation model, effectively improve the application accuracy and optimization efficiency of the enterprise evaluation model, and reduce the difficulty of domain experts to participate, thereby improving the accuracy and effectiveness of applying the enterprise evaluation model to evaluate the enterprise status.

[0072] To further improve the reliability and effectiveness of iterative optimization of enterprise evaluation models, an iterative optimization method for enterprise evaluation models is provided in this application embodiment. (See also...) Figure 2 Step 100 in the iterative optimization method of the enterprise evaluation model specifically includes the following:

[0073] Step 110: Collect raw enterprise feature data based on the preset periodic update strategy and / or dynamic update strategy, perform data preprocessing on the raw enterprise feature data, and store the preprocessed raw enterprise feature data in the corresponding Kafka data topic unit.

[0074] Specifically, third-party interface information can be obtained from data in the MySQL database, and data can be acquired and stored in the corresponding data topic in Kafka using a Python data acquisition program.

[0075] In particular, a regular update strategy is adopted for key interfaces, while a dynamic update strategy is adopted for non-key interfaces. At the same time, the data is cleaned, transformed, classified and compared in a unified manner during collection to ensure that the data can be collected correctly and completely.

[0076] Step 120: After determining that the iterative optimization process for the enterprise evaluation model has started, the original enterprise feature data is obtained from the Kafka data topic unit and processed to obtain the corresponding enterprise feature data. The enterprise feature data is then stored in the target database, and the data changes in the target database are monitored. The target database includes a Hive database.

[0077] Specifically, the Kafka consuming unit is responsible for subscribing to the corresponding Kafka topic and locating the consumption position. It then retrieves data from the relevant Kafka topic, processes the data, writes it to a newly created file, and finally uses the Hive load command to update the data in the file to Hive, updating the offset and crawler record table. The Hive change monitoring unit can perform daily monitoring tasks, querying newly added data in some Hive tables for the day, retrieving related tables to populate the required information, sorting by dt to get the latest results, and finally importing the filtered new data into the tables.

[0078] To further improve the reliability and effectiveness of evaluating the dataset and the trained enterprise evaluation model, an iterative optimization method for the enterprise evaluation model is provided in this application embodiment, see [link to relevant documentation]. Figure 2 Step 300 in the iterative optimization method of the enterprise evaluation model specifically includes the following:

[0079] Step 310: Evaluate the dataset according to the evaluation rules set in advance by domain experts, and save the dataset that passes the evaluation to the MySQL database; and evaluate the performance of the trained enterprise evaluation model, and validate the trained enterprise evaluation model according to the evaluation rules set in advance by domain experts and the validation set, and save the parameter information and performance information corresponding to the enterprise evaluation model that passes the performance evaluation and validation to the MySQL database.

[0080] Specifically, the training dataset can be evaluated based on domain experts' or domain research reports, and the dataset files used for each training session can be organized and saved, with the dataset information stored in a MySQL database. The accuracy, precision, recall, and F1 score of the trained model can also be calculated using the Sklearn framework, and validated using data validation sets provided by domain experts. Finally, the model information and model performance information are saved in a MySQL database.

[0081] Step 320: Deploy the enterprise assessment model to the server, extract enterprise feature data from the Hive database, input each enterprise feature data into the enterprise assessment model deployed on the server, so that the enterprise assessment model outputs enterprise status assessment results corresponding to each enterprise feature data, store each enterprise status assessment result in the Hive database, and store the correspondence between the enterprise feature data and the enterprise status assessment results in a preset model calculation result index.

[0082] Specifically, trained and performance-compliant models can be deployed to a server. The Spark computing engine and shell scripts can be used to extract model computation data from the Hive database and perform computations. After the computation is completed, the computation results are saved to the Hive database, and each model input data and computation result is saved to the Elasticsearch model computation result index.

[0083] To further improve the applicability of the iterative optimization method for enterprise evaluation models, in an embodiment of this application, the enterprise evaluation model includes: Gradient Boosting Decision Tree (GBDT) model.

[0084] The enterprise assessment model is used to output corresponding enterprise status assessment results based on the input enterprise characteristic data. The enterprise status assessment results include: enterprise development stage prediction results, enterprise growth assessment results, or enterprise risk assessment results. Correspondingly, the enterprise assessment model may include: enterprise development stage prediction model, enterprise growth assessment model, and enterprise risk assessment model.

[0085] The predicted enterprise development stages include: seed stage, startup stage, growth stage, expansion stage, maturity stage, or decline stage.

[0086] Specifically, to ensure the accuracy and consistency of using the Gradient Boosting Decision Tree (GBDT) model to assess enterprise growth and risk, it is necessary to iteratively optimize the decision tree model. However, since the data changes continuously over time, this process is inefficient and unreliable, which in turn affects the accuracy and timeliness of enterprise growth and risk assessment.

[0087] For example, the enterprise development stage prediction model mainly uses the gradient boosting decision tree model. When evaluating the enterprise development stage, the characteristics of enterprises at different development stages differ significantly, often by several orders of magnitude. Therefore, the feature binning method is chosen to discretize the enterprise features, and the missing feature values ​​of unlisted enterprises are treated as a separate class to improve the stability and generalization ability of the model. Cross entropy is used as the model's loss function, as shown in the following formula:

[0088]

[0089] Where M is the number of categories, y ic Let c be the sign function; it takes the value 1 if the true class of sample i is equal to c, and 0 otherwise. ic This is the predicted probability that observed sample i belongs to category c. Here, the category refers to the stage of the enterprise (seed stage, startup stage, growth stage, expansion stage, maturity stage, decline stage).

[0090] From a software perspective, this application also provides a pipeline system for executing all or part of the iterative optimization pipeline for the enterprise evaluation model, see [link to relevant documentation]. Figure 3 The enterprise evaluation model iterative optimization pipeline system specifically includes the following:

[0091] The data acquisition and processing module 10 is used to continuously collect original enterprise characteristic data, and after determining that the iterative optimization process for the enterprise evaluation model has started, it processes and optimizes the original enterprise characteristic data to obtain the corresponding enterprise characteristic data, and stores the enterprise characteristic data in the target database.

[0092] The model training and optimization module 20 is used to extract enterprise feature data from the target database to form a dataset for training and validating the enterprise evaluation model, and to train and optimize the previous data version of the enterprise evaluation model based on the dataset to obtain the trained target version of the enterprise evaluation model.

[0093] The model evaluation and deployment module 30 is used to evaluate the dataset and the target version of the enterprise evaluation model according to the evaluation rules data set in advance by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, the target version of the enterprise evaluation model is deployed online, and the enterprise feature data in the target database is input into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the corresponding enterprise status evaluation result.

[0094] The model result feedback module 40 is used to store and output the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results; if the domain expert corrects the enterprise feature data or a message indicating that the enterprise status assessment result is correct is received, the domain expert corrects the enterprise feature data and the enterprise feature data corresponding to the correct enterprise status assessment result are both stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version.

[0095] The embodiment of the enterprise evaluation model iterative optimization pipeline system provided in this application can be used to execute the processing flow of the embodiment of the enterprise evaluation model iterative optimization method in the above embodiment. Its function will not be repeated here, but can be referred to the detailed description of the above embodiment of the enterprise evaluation model iterative optimization method.

[0096] The iterative optimization portion of the enterprise evaluation model pipeline system can be executed on a server. Alternatively, in another practical application scenario, all operations can be completed on the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed on the client device, the client device may further include a processor for the specific processing of the enterprise evaluation model iterative optimization.

[0097] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0098] The server and the client device can communicate using any suitable network protocol, including those not yet developed as of the date of this application. Such network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Furthermore, such network protocols may also include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer Protocol) protocols used on top of the aforementioned protocols.

[0099] As can be seen from the above description, the enterprise evaluation model iterative optimization pipeline system provided in this application embodiment can realize automated iterative optimization of the enterprise evaluation model, effectively improve the application accuracy and optimization efficiency of the enterprise evaluation model, and reduce the difficulty of domain experts to participate, thereby improving the accuracy and effectiveness of applying the enterprise evaluation model to evaluate the enterprise status.

[0100] To further improve the reliability and effectiveness of iterative optimization of enterprise evaluation models, this application provides an embodiment of an iterative optimization pipeline system for enterprise evaluation models, see [link to relevant documentation]. Figure 4 The data acquisition and processing module 10 in the enterprise evaluation model iterative optimization pipeline system specifically includes the following components:

[0101] The data acquisition module 11 is used to collect raw enterprise feature data based on a preset periodic update strategy and / or dynamic update strategy, preprocess the raw enterprise feature data, and store the preprocessed raw enterprise feature data in the corresponding Kafka data topic unit.

[0102] The data processing module 12 is used to obtain the original enterprise feature data from the Kafka data topic unit and process the data after determining that the iterative optimization process for the enterprise evaluation model has started, to obtain the corresponding enterprise feature data, and to store the enterprise feature data in the target database, and to monitor the data changes in the target database, wherein the target database includes: a hive database.

[0103] To further improve the reliability and effectiveness of evaluating the dataset and the trained enterprise evaluation model, an iterative optimization pipeline system for the enterprise evaluation model is provided in this application embodiment. (See also...) Figure 4 The model evaluation and deployment module 30 in the enterprise evaluation model iterative optimization pipeline system specifically includes the following components:

[0104] The model evaluation module 31 is used to evaluate the dataset according to the evaluation rules data set in advance by domain experts, and save the dataset that passes the evaluation to the MySQL database; and to evaluate the performance of the trained enterprise evaluation model, and to verify the trained enterprise evaluation model according to the evaluation rules data set in advance by domain experts and the validation set, and to save the parameter information and performance information corresponding to the enterprise evaluation model that passes the performance evaluation and the validation to the MySQL database.

[0105] The model deployment and calculation module 32 is used to deploy the enterprise evaluation model to the server, extract enterprise feature data from the Hive database, input each enterprise feature data into the enterprise evaluation model deployed on the server, so that the enterprise evaluation model outputs the enterprise status evaluation results corresponding to each enterprise feature data, stores each enterprise status evaluation result in the Hive database, and stores the correspondence between the enterprise feature data and the enterprise status evaluation results in a preset model calculation result index.

[0106] To address the issues of low efficiency in enterprise evaluation model construction and optimization, outdated training data, and difficulty in involving domain experts in the optimization process, thus compromising the accuracy and effectiveness of enterprise evaluation, this application also provides an enterprise evaluation method, see [link to relevant documentation]. Figure 5 The enterprise evaluation method specifically includes the following:

[0107] Step 500: Receive the enterprise evaluation model parameters and store the target version of the enterprise evaluation model corresponding to the enterprise evaluation model parameters locally.

[0108] In step 500, the target version of the enterprise evaluation model is generated in advance based on the enterprise evaluation model iterative optimization method provided in the foregoing embodiments;

[0109] Step 600: Obtain the enterprise characteristic data of the target enterprise.

[0110] Step 700: Input the enterprise characteristic data into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the enterprise status evaluation result of the target enterprise, wherein the enterprise status evaluation result includes: enterprise development stage prediction result, enterprise growth assessment result, or enterprise risk assessment result.

[0111] To further illustrate this solution, this application also provides a specific application example of an iterative optimization method for enterprise evaluation models, and also relates to a general model iterative optimization method. Specifically, machine learning is the process by which computers learn new experiences and knowledge by studying the inherent regularities in data, thereby improving the computer's intelligence and enabling it to make decisions like humans. With the development of big data, on the one hand, data features are becoming increasingly comprehensive, and on the other hand, the data volume is large enough to guarantee the implementation of machine learning models. Therefore, many big data platforms hope to use machine learning models in their applications.

[0112] Currently, machine learning models integrated into big data platforms often exhibit lower accuracy or consistency compared to laboratory results. This is because big data platforms are characterized by large data volumes, rapid data generation and processing, and real-time requirements. This can lead to data lag during the integration of machine learning models with the platform, resulting in models performing well on the original datasets delivered by developers but performing poorly on newly collected data. Current methods to address accuracy or consistency issues in machine learning models primarily involve increasing model size. For example, in decision trees and random forests, low-accuracy classifiers are integrated and weighted to achieve more accurate results; or better model hyperparameters and architectures are sought. For instance, after selecting a model, hyperparameters can be fine-tuned using grid search or random search, and in neural networks, the architecture can be adjusted for different datasets. However, all these solutions require significant time and resources from experienced machine learning professionals, and even after adjusting the model, data lag and data drift issues persist as new data arrives over time within the big data streaming model. Meanwhile, as business grows and the collected information becomes more complete, the original model cannot utilize novel data that was never observed before the model was generated and collected, and this data can have a profound impact on model performance. Therefore, Danilo Sato et al. proposed Continuous Delivery for Machine Learning (CD4ML), which aims to ensure the accuracy and stability of machine learning models by continuously iterating and optimizing them.

[0113] (I) General Model Iterative Optimization Method

[0114] This application example aims to design a data-driven iterative optimization method for machine learning models. By discovering, collecting, transforming, and understanding model data, the dataset is iterated. At the same time, expert feedback on the model results is introduced to form new labeled data for human-computer collaboration. The method also realizes an automated pipeline for model training optimization and deployment, solving the problems of slow model building, data lag, and difficulty for domain experts to participate in verification.

[0115] The purpose of this application example is to provide a data-driven iterative optimization method for machine learning model datasets and a machine learning model iterative optimization pipeline based on this method. The method continuously collects model data through a data acquisition module, processes the collected raw data to form a model dataset through a data processing module, evaluates the model dataset by introducing domain expert opinions in the model evaluation module, and finally completes model construction using a model training and optimization module. The entire process is automated into a pipeline, realizing the sustainable delivery of machine learning models and solving the problems of slow model construction, data lag, and difficulty in involving domain experts in verification in existing models.

[0116] See Figure 6 Data-driven machine learning model dataset iterative optimization methods mainly include the design of data iteration space and data iteration cycle:

[0117] The data iteration space treats each row of the training dataset as an instance, each column as a feature, and the final learning target as a label. All data iteration spaces can be enumerated by adding, modifying, and deleting instances, features, and labels from the dataset.

[0118] Optionally, the data iteration cycle can be iterated in the time dimension at the granularity of hours, days, months, quarters or years.

[0119] Optionally, the data iteration cycle can be iterated at the granularity of the percentage of data change and the amount of data fed back by the model results, in terms of the degree of change.

[0120] See Figure 7 The machine learning model iterative optimization pipeline based on the above method includes a data acquisition module, a data processing module, a model training and optimization module, a model evaluation module, a model deployment and computation module, and a model result feedback module.

[0121] The data acquisition module obtains third-party interface information based on data in the MySQL database, acquires data through a Python data acquisition program, and stores it in the corresponding data topic in Kafka.

[0122] Specifically, the data acquisition module adopts a periodic update strategy for key interfaces and a dynamic update strategy for non-key interfaces. At the same time, the data is cleaned, transformed, classified and compared uniformly during the acquisition process to ensure that the data can be collected correctly and completely.

[0123] The data processing module includes a Kafka consumption unit and a Hive change monitoring unit:

[0124] The Kafka consuming unit is responsible for subscribing to the corresponding Kafka topic and locating the consumption position. Then, it retrieves data from the corresponding Kafka topic, processes the data, writes it to a newly created file, and finally uses the Hive load command to update the data in the file to Hive, updating the offset and the crawler record table.

[0125] The monitoring unit for Hive changes executes a monitoring task on a daily basis, queries some Hive tables for newly added data for the day, retrieves related tables to populate the required information, sorts the data by dt to get the latest results, and finally imports the filtered new data into the tables.

[0126] The model training and optimization module extracts the dataset required for model training from the Hive database through the Spark computing engine. Then, it uses the GridSearchCV grid search algorithm and cross-validation to find the optimal hyperparameters of the model on this dataset, adjusts the positive and negative samples, and finally completes the model construction.

[0127] The model evaluation module includes a data evaluation unit and a model performance evaluation unit:

[0128] The data evaluation unit evaluates the training dataset based on domain experts or domain research reports, organizes and saves the dataset files used for each training session, and saves the dataset information to a MySQL database.

[0129] The model performance evaluation unit calculates the accuracy, precision, recall, and F1 score of the trained model using the Sklearn framework, and validates the model using data validation sets provided by domain experts. Finally, the model information and model performance information are saved to a MySQL database.

[0130] The model deployment and computation module deploys the trained and performance-compliant model to the server. It uses the Spark computing engine and shell scripts to extract model computation data from the Hive database and perform computations. After the computation is completed, the computation results are saved to the Hive database, and each model input data and computation result is saved to the Elasticsearch model computation result index.

[0131] The model result feedback module provides the application with a display of model calculation results based on the Elasticsearch model calculation result index. Domain experts monitor the model calculation results through the front-end application to confirm whether the calculation results are correct and whether they need to be corrected. Validated or corrected data will be exported to form a new training dataset and supplemented in the next data iteration.

[0132] Specifically, the execution process of the model iteration pipeline is as follows:

[0133] S1: The data acquisition module performs third-party data acquisition.

[0134] S2: The data processing module retrieves data from Kafka and monitors changes in Hive data to provide data for model calculation and training.

[0135] S3: The model training and optimization module searches for the best model hyperparameters based on the new data and trains the model.

[0136] S4: The model evaluation module records the model training results and model validation results.

[0137] S5: The model deployment and computation module provides model computation services.

[0138] S6: The model result feedback module generates new training data based on expert feedback on the model calculation results and provides it to the model training and optimization module.

[0139] This paper presents a data-driven iterative optimization method for machine learning models and a pipeline for iterative optimization based on this method. By introducing the concepts of data iteration space and data iteration cycle, the paper clarifies the data-driven dataset iterative optimization method and manages model versions based on data versions. This enables lifecycle management and continuous iterative optimization of machine learning models, providing continuous model computation and iterative optimization services for big data platforms. It also supports domain experts in validating and providing feedback on model results, allowing domain experts without machine learning experience to participate in the development and training of machine learning models, and providing new solutions for model optimization.

[0140] Additionally, see Figure 8 This application also provides examples of pipeline architecture, which mainly consists of 5 layers:

[0141] The first layer of data sources includes data from third-party interfaces, enterprise basic data, and industry data.

[0142] The second data storage layer stores the data from the data source and the results of data processing and calculation.

[0143] The third layer, data processing and computation: This layer utilizes tools such as Python, Spark, and Hive SQL to consume Kafka data and monitor HIV. e It automates the functions of modification, model data processing, model data calculation, and log file collection, while providing model training and calculation data for upper-layer services.

[0144] The fourth layer, model training and computation layer, builds a framework based on existing models for model training and also supports computation of deployed models.

[0145] The fifth layer, the data query and display layer, displays the model's iterative optimization status and calculation results to end users.

[0146] (II) Iterative Optimization Methods for Enterprise Evaluation Models

[0147] To ensure the accuracy and consistency of using Gradient Boosting Decision Tree (GBDT) to assess a company's growth and risk, iterative optimization of the decision model is necessary. However, since the data changes over time, this process is inefficient and unreliable, which in turn affects the accuracy and timeliness of the assessment of a company's growth and risk.

[0148] This application refers to the above-mentioned model iterative optimization method, uses data acquisition tools to prepare data, introduces experts to verify the data, and takes the enterprise development stage prediction model as a specific implementation example.

[0149] The enterprise development stage prediction model primarily uses a gradient boosting decision tree model. When evaluating enterprise development stages, the characteristics of enterprises at different stages differ significantly, often by several orders of magnitude. Therefore, a feature binning method is chosen to discretize enterprise features, and missing feature values ​​for unlisted enterprises are treated as a separate category to improve the model's stability and generalization ability. Cross-entropy is used as the model's loss function, as shown in the formula below:

[0150]

[0151] Where M is the number of categories, y ic Let c be the sign function; it takes the value 1 if the true class of sample i is equal to c, and 0 otherwise. ic This is the predicted probability that observed sample i belongs to category c. Here, the category refers to the stage of the enterprise (seed stage, startup stage, growth stage, expansion stage, maturity stage, decline stage).

[0152] The prediction model for enterprise development stages utilizes the aforementioned machine learning model to iteratively optimize the pipeline. It employs 15 features, including enterprise registered capital, number of patents, number of employees covered by social security, number of court hearings, and duration of existence, along with 6 labels: seed stage, startup stage, growth stage, maturity stage, expansion stage, and decline stage, to construct a total of 56,393 instances.

[0153] The GridSearch method is mainly used to search for hyperparameters such as learning rate, number of iterations, maximum depth, number of trees (n_estimators), and number of leaves on the trees (num_leaves) to complete model training.

[0154] Since the above features change over time, 6,000 complete and representative enterprises were used as the validation set for model validation. At the same time, based on the above 56,393 instances, the above machine learning model dataset iterative optimization method was applied. Multiple models were trained with the same model construction framework but different model parameters, and the same validation set was used for validation.

[0155] This application example introduces a data-driven iterative optimization method for machine learning models, and introduces a human-machine collaboration approach to enable non-machine learning experts to participate in the iterative optimization process. We also automated this process in our embodiments and deployed a machine learning model iterative optimization pipeline based on this method. In actual production, we found that the data-driven approach can effectively improve model performance over time and continuously provide us with better models.

[0156] It is foreseeable that with the development of big data and the continuous increase in data volume, the data features that machine learning models can utilize will become richer. Therefore, iteratively optimizing machine learning models through data-driven methods will provide a new approach to model optimization.

[0157] This application also provides an electronic device (i.e., an electronic device), such as a central server. This electronic device may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the enterprise evaluation model iterative optimization method and / or enterprise evaluation method mentioned in the above embodiments. The processor and memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and memory via wired or wireless means. The electronic device can receive real-time motion data from sensors in the wireless multimedia sensor network and receive raw video sequences from the video acquisition device.

[0158] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0159] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the enterprise evaluation model iterative optimization method and / or enterprise evaluation method in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the enterprise evaluation model iterative optimization method and / or enterprise evaluation method in the above method embodiments.

[0160] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0161] The one or more modules are stored in the memory, and when executed by the processor, they execute the enterprise evaluation model iterative optimization method and / or the enterprise evaluation method in the embodiment.

[0162] In some embodiments of this application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected via a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0163] As one implementation method, the functions of the receiver and transmitter in this application can be implemented by transceiver circuits or dedicated transceiver chips, and the processor can be implemented by dedicated processing chips, processing circuits or general-purpose chips.

[0164] As another implementation approach, the server provided in this application embodiment can be implemented using a general-purpose computer. That is, the program code implementing the processor, receiver, and transmitter functions is stored in memory, and the general-purpose processor implements the processor, receiver, and transmitter functions by executing the code in memory.

[0165] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned iterative optimization method for enterprise evaluation models and / or enterprise evaluation methods. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0166] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0167] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0168] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0169] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An iterative optimization method for an enterprise evaluation model, characterized in that, include: Raw enterprise feature data is collected based on a preset periodic update strategy and / or dynamic update strategy. The raw enterprise feature data is preprocessed and stored in the corresponding Kafka data topic unit. Specifically, third-party interface information is obtained from data in the MySQL database, and data is acquired and stored in the corresponding Kafka data topic through a Python data acquisition program. After the iterative optimization process for the enterprise evaluation model is initiated, raw enterprise characteristic data is obtained from the Kafka data topic unit and processed to obtain corresponding enterprise characteristic data. This enterprise characteristic data is then stored in the target database, and changes in the target database are monitored. The target database includes a Hive database. The Kafka consuming unit is responsible for subscribing to the corresponding Kafka topic and locating the consumption position. Then, it retrieves data from the corresponding Kafka topic, processes the data, and writes it to a newly created file. Finally, the Hive database's load command is used to update the data in the file to the Hive database, updating the offset and crawler record table. The Hive change monitoring unit executes a monitoring task daily, queries the Hive database for newly added data that day, retrieves related tables to populate the required information, sorts the results by dt to get the latest results, and finally imports the filtered new data into the table. Enterprise feature data is extracted from the target database to form the dataset used for training and validating the enterprise evaluation model. Based on this dataset, the previous data version of the enterprise evaluation model is trained and optimized to obtain the target version of the enterprise evaluation model after training. The dataset is evaluated according to the evaluation rules set in advance by domain experts, and the datasets that pass the evaluation are saved to the MySQL database; and the performance of the trained enterprise evaluation model is evaluated, and the trained enterprise evaluation model is validated according to the evaluation rules set in advance by domain experts and the validation set, and the parameter information and performance information corresponding to the enterprise evaluation model that passes the performance evaluation and the validation are saved to the MySQL database. The enterprise evaluation model is deployed to a server, and enterprise feature data is extracted from the Hive database. Each enterprise feature data is then input into the enterprise evaluation model deployed to the server, so that the model outputs enterprise status evaluation results corresponding to each enterprise feature data. These enterprise status evaluation results are stored in the Hive database, and the correspondence between the enterprise feature data and the enterprise status evaluation results is stored in a preset model calculation result index. Specifically, the trained and performance-compliant model is deployed to the server, and the Spark computing engine and shell scripts are used to extract model calculation data from the Hive database and perform calculations. After calculation, the results are saved to the Hive database, and each model input data and calculation result are saved to the Elasticsearch model calculation result index. The system stores and outputs the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results. If the system receives the corrected enterprise feature data from the domain expert or a message indicating that the enterprise status assessment results are correct, then the corrected enterprise feature data from the domain expert and the enterprise feature data corresponding to the correct enterprise status assessment results are both stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version. The enterprise evaluation model includes: Gradient Boosting Decision Tree Model (GBDT); The feature binning method is chosen to discretize the enterprise feature data, and the missing feature values ​​of unlisted enterprises are treated as a separate class to improve the gradient and enhance the stability and generalization ability of the decision tree model GBDT. Cross-entropy is used as the loss function L for the Gradient Boosting Decision Tree (GBDT) model, as shown in the following formula: where M is the number of classes, is the indicator function that takes 1 if the sample has a true class equal to and 0 otherwise; is the predicted probability that the sample belongs to class .

2. The iterative optimization method for the enterprise evaluation model according to claim 1, characterized in that, The enterprise assessment model is used to output corresponding enterprise status assessment results based on the input enterprise characteristic data. The enterprise status assessment results include: enterprise development stage prediction results, enterprise growth assessment results, or enterprise risk assessment results. The predicted enterprise development stages include: seed stage, startup stage, growth stage, expansion stage, maturity stage, or decline stage.

3. A method of assessing the state of an enterprise, characterized by, include: Receive enterprise evaluation model parameters and store the target version of the enterprise evaluation model corresponding to the enterprise evaluation model parameters locally, wherein the target version of the enterprise evaluation model is generated in advance based on the enterprise evaluation model iterative optimization method described in claim 1 or 2; Obtain enterprise characteristic data of the target company; The enterprise characteristic data is input into the target version of the enterprise assessment model so that the target version of the enterprise assessment model outputs the enterprise status assessment result of the target enterprise, wherein the enterprise status assessment result includes: enterprise development stage prediction result, enterprise growth assessment result, or enterprise risk assessment result.

4. An enterprise evaluation model iterative optimization pipeline system, comprising: The system is used to perform the iterative optimization method for the enterprise evaluation model as described in claim 1 or 2; the system includes: The data acquisition and processing module is used to continuously collect raw enterprise characteristic data, and after determining that the iterative optimization process for the enterprise evaluation model has started, it processes and optimizes the raw enterprise characteristic data to obtain the corresponding enterprise characteristic data, and stores the enterprise characteristic data in the target database. The model training and optimization module is used to extract enterprise feature data from the target database to form the dataset used for training and validating the enterprise evaluation model, and to train and optimize the previous data version of the enterprise evaluation model based on the dataset to obtain the target version of the enterprise evaluation model after training. The model evaluation and deployment module is used to evaluate the dataset and the target version of the enterprise evaluation model according to the evaluation rules data set in advance by domain experts. If the dataset and the target version of the enterprise evaluation model pass the evaluation, the target version of the enterprise evaluation model is deployed online, and the enterprise feature data in the target database is input into the target version of the enterprise evaluation model so that the target version of the enterprise evaluation model outputs the corresponding enterprise status evaluation result. The model result feedback module is used to store and output the correspondence between the enterprise feature data and the enterprise status assessment results, so that domain experts can determine whether to correct the corresponding enterprise feature data based on the enterprise status assessment results; if the corrected enterprise feature data or a message indicating that the enterprise status assessment result is correct is received from the domain expert, the corrected enterprise feature data and the enterprise feature data corresponding to the correct enterprise status assessment result are both stored in the optimized dataset for iterative optimization of the enterprise assessment model for the next data version; The data acquisition and processing module includes: The data acquisition module is used to collect raw enterprise feature data based on a preset periodic update strategy and / or dynamic update strategy, preprocess the raw enterprise feature data, and store the preprocessed raw enterprise feature data in the corresponding Kafka data topic unit. The data processing module is used to obtain raw enterprise feature data from the Kafka data topic unit and process the data after determining that the iterative optimization process for the enterprise evaluation model has started, to obtain the corresponding enterprise feature data, and to store the enterprise feature data in the target database and monitor the data changes in the target database. The target database includes a Hive database. The overall module for model evaluation and deployment includes: The model evaluation module is used to evaluate the dataset according to the evaluation rules data set in advance by domain experts, and save the dataset that passes the evaluation to the MySQL database; and to evaluate the performance of the trained enterprise evaluation model, and to verify the trained enterprise evaluation model according to the evaluation rules data set in advance by domain experts and the validation set, and to save the parameter information and performance information of the enterprise evaluation model that passes the performance evaluation and the validation to the MySQL database. The model deployment and calculation module is used to deploy the enterprise evaluation model to the server, extract enterprise feature data from the Hive database, input each enterprise feature data into the enterprise evaluation model deployed on the server, so that the enterprise evaluation model outputs the enterprise status evaluation results corresponding to each enterprise feature data, stores each enterprise status evaluation result in the Hive database, and stores the correspondence between the enterprise feature data and the enterprise status evaluation results in a preset model calculation result index. The enterprise evaluation model includes: Gradient Boosting Decision Tree Model (GBDT); The feature binning method is chosen to discretize the enterprise feature data, and the missing feature values ​​of unlisted enterprises are treated as a separate class to improve the gradient and enhance the stability and generalization ability of the decision tree model GBDT. Cross-entropy is used as the loss function L for the Gradient Boosting Decision Tree (GBDT) model, as shown in the following formula: where M is the number of classes, is the indicator function that takes 1 if the sample has a true class equal to and 0 otherwise; is the predicted probability that the sample belongs to class .

5. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the iterative optimization method for the enterprise evaluation model as described in claim 1 or 2, or the enterprise status evaluation method as described in claim 3.

6. A computer-readable storage medium having stored thereon a computer program, characterized in that, When executed by a processor, the computer program implements the iterative optimization method for the enterprise evaluation model as described in claim 1 or 2, or the enterprise status evaluation method as described in claim 3.