A machine learning-based enterprise operation health degree evaluation method

By time-labeling and sample verification of the enterprise health assessment model, the problem of insufficient verification of similar historical samples in existing models is solved, which improves the accuracy of identifying enterprise operating status and risk identification ability, and enhances the stability of assessment conclusions.

CN122492016APending Publication Date: 2026-07-31NANJING JIYANYUN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING JIYANYUN INTELLIGENT TECH CO LTD
Filing Date
2026-05-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing business health assessment models lack a further verification mechanism for similar historical samples, resulting in insufficient accuracy and explanatory power in identifying borderline enterprises, enterprises with abnormal fluctuations, and enterprises with potential risks.

Method used

By acquiring multi-cycle operating data and subsequent operating status information of historical enterprise samples, time-annotated data is generated to determine the lag relationship between the original health status label and the risk status, generating status stratification information. Based on this, an enterprise operating health assessment model is trained, supporting and counterexample samples are retrieved, the initial health assessment information is corrected, and a more comprehensive sample validation set is formed.

Benefits of technology

It enhances the ability to identify complex business conditions and risks, strengthens the stability and credibility of business health assessment conclusions, and provides more comprehensive historical sample data to verify the initial assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492016A_ABST
    Figure CN122492016A_ABST
Patent Text Reader

Abstract

This invention discloses a machine learning-based method for assessing the operational health of enterprises, belonging to the field of data assessment technology. The method includes: acquiring multi-period operational data and corresponding subsequent operational status information of historical enterprise samples; time-annotating the multi-period operational data and subsequent operational status information to determine the lag relationship between the original health label and the risk status; using the lag relationship of the risk status as the basis for sample stratification to generate status stratification information for each historical enterprise sample; and calibrating the original health labels with delayed risk labels to form a historical calibration training sample set; determining the risk of misjudgment in the initial health assessment information based on the sample validation set, correcting the initial health assessment information, and obtaining the operational health assessment conclusion of the enterprise to be assessed. This invention characterizes similar business models and status differentiation features through the sample validation set, identifies potential risk boundaries, realizes the verification of initial assessment results, and improves the ability to identify complex operational statuses and risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data evaluation technology, and in particular to a method for evaluating the health of business operations based on machine learning. Background Technology

[0002] With the development of enterprise operation data collection capabilities and machine learning modeling technology, enterprise operation health assessment is gradually shifting from judging by a single financial indicator to multi-cycle, multi-dimensional feature analysis. By performing feature processing on data such as revenue changes, cash flow status, performance capability, cost structure, and credit performance, the assessment model can identify non-linear correlations in the enterprise's operating status, providing quantitative basis for risk identification, credit evaluation, and business decision-making.

[0003] In existing technologies, the assessment of business health usually focuses on the model output results themselves, lacking a further verification mechanism for similar historical samples. For companies with similar business characteristics but different subsequent health statuses, the model may easily ignore the state differentiation boundary, making the initial assessment conclusions lack sufficient sample reference support, which in turn affects the accuracy and interpretability of identifying boundary-type companies, companies with abnormal fluctuations, and companies with potential risks. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a machine learning-based method for assessing business health to address the problem of insufficient identification of state differentiation boundaries in business health assessment, which leads to a high risk of misjudgment.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a machine learning-based method for assessing the health of business operations, which includes: acquiring multi-period operating data and corresponding subsequent operating status information of historical enterprise samples; time-annotating the multi-period operating data and subsequent operating status information; and determining the lag relationship between the original health label and the risk status. Using the lag relationship of risk status as the basis for sample stratification, status stratification information of each historical enterprise sample is generated, and the original health status label is calibrated with delayed risk label to form a set of historical calibration training samples. The enterprise business health assessment model is trained based on a historical calibration training sample set. The current business data of the enterprise to be assessed is characterized into current business characteristics and then input into the enterprise business health assessment model to obtain initial health assessment information. Based on the current operating characteristics and initial health assessment information, support samples that are similar to the current operating characteristics and have the same health status, as well as counterexample samples that are similar to the current operating characteristics but deviate from the health status, are retrieved from the historical calibration training sample set to form a sample verification set; Based on the sample validation set, the risk of misjudgment of the initial health assessment information is determined, and the initial health assessment information is corrected to obtain the conclusion of the business health assessment of the enterprise to be assessed.

[0007] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the step of time-annotating multi-period operating data and subsequent operating status information refers to obtaining multi-period operating data and corresponding subsequent operating status information of historical enterprise samples, and time-annotating the multi-period operating data and subsequent operating status information according to the chronological relationship between the multi-period operating data and the subsequent operating status information to obtain time-annotated information.

[0008] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, wherein determining the lag relationship between the original health label and the risk status includes: Based on multi-period operating data of historical enterprise samples, the corresponding operating status is determined and initial health status is marked to obtain the original health status label; Based on time-labeled information, the original health status label is sequentially associated with subsequent business status information. The lag of risk status in subsequent business status information relative to the original health status label is identified and associated with the original health status label to obtain the risk status lag relationship.

[0009] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the step of determining the lag relationship between the original health label and the risk status includes: Based on multi-period operating data of historical enterprise samples, the corresponding operating status is determined and initial health status is marked to obtain the original health status label; Based on time-labeled information, the original health status label is sequentially associated with subsequent business status information. The lag of risk status in subsequent business status information relative to the original health status label is identified and associated with the original health status label to obtain the risk status lag relationship.

[0010] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the specific steps for forming the historical calibration training sample set are as follows: Based on the status stratification information, we identify historical enterprise samples that are in the risk delay layer and extract the corresponding original health status labels, subsequent operating status information and risk status lag relationship. Based on the lag relationship of risk status, determine the lag of risk status in subsequent operational status information relative to the original health status label, and calibrate the original health status label corresponding to the historical enterprise sample in the risk lag layer to the delayed risk label according to the lag. Based on the state stratification information, historical enterprise samples in the health and stability layer and the risk manifestation layer are determined, and the original health labels of historical enterprise samples in the health and stability layer and the risk manifestation layer are retained. By associating multi-period operating data, time-stamped information, status stratification information, delay risk labels, and retained original health labels, a set of historical calibration training samples containing training labels is formed.

[0011] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the step of training the enterprise health assessment model based on a historical calibration training sample set includes: The multi-cycle operating data in the historical calibration training sample set is characterized to obtain historical operating features, which are then associated with the delayed risk label and the retained original health label to form a health training sample. The enterprise operational health assessment model is trained based on the health training samples to obtain the health training output for the health training samples. Based on state hierarchical information, the deviation of the health training output in the health stability layer, risk manifestation layer and risk delay layer is verified in a hierarchical manner to determine the evaluation bias of the enterprise business health assessment model. Based on the assessment bias, identify historical enterprise samples with a tendency to misjudge from the historical calibration training sample set, and use these historical enterprise samples with a tendency to misjudge to continue training the enterprise business health assessment model, thus obtaining the trained enterprise business health assessment model.

[0012] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the method includes: acquiring the current operating data of the enterprise to be assessed, performing feature processing on the current operating data, and obtaining the current operating features; The current operating characteristics are input into the trained enterprise operating health assessment model, which identifies the health status and risk status of the current operating characteristics, obtains health status judgment information and risk status judgment information, and forms the initial health assessment information of the enterprise to be assessed.

[0013] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the specific steps for forming the sample verification set are as follows: Align the current business characteristics with the historical business characteristics in the historical calibration training sample set by feature dimensions to obtain business characteristic comparison information. Based on the business characteristic comparison information, retrieve historical enterprise samples that are similar to the current business characteristics in the historical calibration training sample set to obtain similar historical enterprise samples. The training labels of similar historical enterprise samples are compared with the initial health assessment information to identify supporting samples with similar current operating characteristics and consistent health status. By comparing the training labels of similar historical enterprise samples with the initial health assessment information, negative examples with similar current operating characteristics but deviating health status are identified. The supporting samples and their corresponding state levels are used as the consistency verification content, while the negative samples, their corresponding state levels, and the risk state lag relationship of the negative samples are used as the deviation verification content, thus forming a sample verification set.

[0014] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the step of determining the risk of misjudgment of initial health assessment information based on a sample validation set includes: Based on the consistency and deviation verification content in the sample verification set, the degree of support and deviation for the initial health assessment information are determined respectively. Based on the state level corresponding to the counterexample samples, determine the distribution of counterexample samples in the risk delay layer, and combine the risk state lag relationship of the counterexample samples to identify the omission of delayed risks in the initial health assessment information. By correlating the degree of support, the degree of deviation, and the omission of delay risk, the risk of misjudgment in the initial health assessment information is obtained.

[0015] As a preferred embodiment of the machine learning-based enterprise health assessment method of the present invention, the step of obtaining the enterprise health assessment conclusion refers to determining, based on the risk of misjudgment in the initial health assessment information, the health status judgments that may have been missed due to delays in the initial health assessment information. The initial health assessment information is corrected according to the type of misjudgment to obtain the corrected health assessment information, which is then used as the conclusion of the business health assessment of the enterprise to be assessed.

[0016] The beneficial effects of this invention are as follows: Based on the current operating characteristics and initial health assessment information retrieval support samples and counterexample samples, a sample verification set with the ability to verify consistent status and identify deviation status can be formed, providing more sufficient historical sample basis for the enterprise to be assessed. The sample verification set describes similar operating models and status differentiation characteristics, marks potential risk boundaries, realizes the verification of the initial assessment results, improves the ability to identify complex operating status and risk identification, and enhances the stability and credibility of the operating health assessment conclusion. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a machine learning-based method for assessing the health of a company's operations.

[0019] Figure 2 This is a flowchart of the initial assessment process for companies to be evaluated.

[0020] Figure 3 A flowchart for constructing a similar sample validation set.

[0021] Figure 4 A flowchart for correcting health status misjudgments. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a machine learning-based method for assessing the health of business operations, including the following steps: S1: Obtain multi-period operating data and corresponding subsequent operating status information of historical enterprise samples, time-annotate the multi-period operating data and subsequent operating status information, and determine the lag relationship between the original health label and the risk status.

[0026] S1.1: Obtain multi-period operating data and corresponding subsequent operating status information of historical enterprise samples. Based on the chronological relationship between the multi-period operating data and the subsequent operating status information, time-label the multi-period operating data and the subsequent operating status information to obtain time-labeled information.

[0027] Using the enterprise's unified social credit code as the unique identifier, we obtain basic enterprise information, multi-period operating data, and subsequent operating status information of historical enterprise samples.

[0028] Basic enterprise information includes the enterprise's unified social credit code and industry category. Multi-period operating data includes operating revenue, operating costs, net profit, net cash flow from operating activities, total assets, total liabilities, accounts receivable balance, inventory balance, tax declaration status, business registration status, administrative penalty records, and records of abnormal operations. Subsequent operating status information includes normal existence, abnormal operations, serious violations of laws and regulations, being subject to enforcement, being subject to enforcement for dishonesty, business restrictions due to administrative penalties, abnormal taxation, bankruptcy reorganization, bankruptcy liquidation, revocation, and deregistration. The above information is obtained through legally authorized channels or legally publicly available channels; when authorization is required for data, it is used only after obtaining authorization from the relevant parties; all information is used for legitimate purposes with the user's consent.

[0029] For multi-period operating data, the enterprise's unified social credit code, industry category, operating cycle start date, operating cycle end date, and data formation date are entered. For subsequent operating status information, the enterprise's unified social credit code, subsequent operating status type, subsequent operating status occurrence date, and subsequent operating status source are entered. Among these, the data formation date for financial statement data is the end date of the operating cycle to which the financial statements belong; the data formation date for tax declaration status is the end date of the tax declaration period; and the data formation date for business registration status, administrative penalty records, and abnormal operation records is the formation date stated in the corresponding public record.

[0030] The multi-period operating data and subsequent operating status information are cleaned and deduplicated separately: multi-period operating data with empty enterprise unified social credit code, industry category, operating cycle end date, or data formation date are removed; subsequent operating status information with empty enterprise unified social credit code, subsequent operating status type, or subsequent operating status occurrence date is also removed; for multiple multi-period operating data with the same enterprise unified social credit code, the same operating cycle, and the same data item, the record with the latest data formation date and the highest source priority is retained; for multiple subsequent operating status information with the same enterprise unified social credit code, the same subsequent operating status type, and the same subsequent operating status occurrence date, the record with the highest source priority is retained.

[0031] Based on the enterprise's unified social credit code, multi-period operating data is matched with subsequent operating status information; the end date of the operating cycle is taken as the cycle time point, and the date of occurrence of the subsequent operating status is taken as the status time point; when the date of occurrence of the subsequent operating status is later than the end date of the operating cycle, the number of natural days between the two is determined as the time interval, and a time-marked record is generated that includes the enterprise's unified social credit code, industry category, start date of the operating cycle, end date of the operating cycle, data formation date, type of subsequent operating status, date of occurrence of subsequent operating status, source of subsequent operating status, and time interval, and the time-marked information is obtained by summarizing.

[0032] S1.2: Determine the corresponding operating status based on the multi-period operating data of historical enterprise samples, and perform initial health labeling to obtain the original health label.

[0033] Based on time-marked information, the system extracts multi-cycle operating data of the same historical enterprise sample within the same operating cycle according to the enterprise's unified social credit code, industry category, start date of operating cycle, and end date of operating cycle, and determines the operating status based on the multi-cycle operating data.

[0034] Operating status is determined based on revenue size, profitability, cash flow, debt repayment, accounts receivable collection, inventory holding, and external operating constraints.

[0035] Specifically, the revenue scale status is determined based on operating revenue, the profitability status is determined based on net profit, the cash flow status is determined based on net cash flow from operating activities, the debt repayment status is determined based on the ratio of total liabilities to total assets, the accounts receivable collection status is determined based on the ratio of accounts receivable balance to operating revenue, the inventory status is determined based on the ratio of inventory balance to operating costs, and the external operating constraint status is determined based on tax declaration status, business registration status, administrative penalty records, and abnormal operation records.

[0036] When total assets are zero, the risk level of debt repayment status is determined by whether total liabilities are zero; when operating revenue is zero, the risk level of accounts receivable collection status is determined by whether accounts receivable balance is zero; when operating costs are zero, the risk level of inventory occupation status is determined by whether inventory balance is zero. A zero balance is recorded as the minimum risk level, and a non-zero balance is recorded as the maximum risk level.

[0037] Based on industry category and operating cycle cutoff date, historical enterprise samples were divided into peer-to-peer (P2P) sample sets. Within these P2P sample sets, revenue size, profitability, and cash flow status were ranked from best to worst performance. Debt repayment status, accounts receivable collection status, and inventory status were ranked from lowest to highest risk level. The ranking results were then divided into top, middle, and bottom segments based on sample size. Indicators ranked in the top segment were labeled as good, those in the middle segment as normal, and those in the bottom segment as weak. When enterprises with the same ranking were at the segment boundaries, they were further divided according to the segment that best reflected the risk.

[0038] Assess the status of external operational constraints. A "good" status is defined as follows: normal tax filing status, normal business registration status, no administrative penalty records, and no abnormal business operation records. A "normal" status is defined as follows: normal status is defined as follows: abnormal tax filing status, abnormal business registration status, incomplete administrative penalty records, or unremoved abnormal business operation records.

[0039] The number of indicators categorized as "Good," "Normal," and "Weak" for the seven business status indicators is statistically analyzed. The indicator with the highest number of indicators is determined as the business status. If the number of indicators is the same, the business status is determined in the order of "Weak," "Normal," and "Good." When the business status is "Good," "Normal," or "Weak," the original health label is marked as "Good," "Normal," or "Weak," respectively.

[0040] The original health label is obtained by combining the enterprise's unified social credit code, industry category, start date of the operating cycle, end date of the operating cycle, operating status, and original health label. All original health label records are then aggregated to obtain the original health label.

[0041] S1.3: Based on time-labeled information, the original health status label is sequentially associated with subsequent business status information to identify the lag of risk status in subsequent business status information relative to the original health status label, and then associated with the original health status label to obtain the risk status lag relationship.

[0042] The original health status records are matched with the time-marked records based on the enterprise's unified social credit code, the start date of the operating cycle, and the end date of the operating cycle. Risk status is then filtered from subsequent operating status types. Risk status includes abnormal operation, serious violation of laws and regulations, being subject to enforcement, being subject to enforcement for dishonesty, administrative penalties leading to operational restrictions, abnormal taxation, bankruptcy reorganization, bankruptcy liquidation, revocation, and deregistration; and normal existence without action risk status.

[0043] When the date of the subsequent operating status is later than the end date of the operating cycle, and the type of the subsequent operating status is a risk status, the risk status will be determined as the risk status that occurred after the original health status label.

[0044] When the same historical enterprise sample corresponds to multiple risk states within the same operating cycle, the risk state with the earliest occurrence time is selected as the earliest associated risk state according to the order of the occurrence dates of subsequent operating states; according to the degree of impact of the risk state on the enterprise's ability to continue operating, credit status and operating qualifications, the risk state with the highest risk level is determined from multiple risk states as the highest risk associated state; when multiple risk states have the same risk level, the risk state with the earliest occurrence time is selected as the highest risk associated state.

[0045] The initial risk delay is determined based on the time interval corresponding to the earliest associated risk status. The number of calendar days between the start and end dates of the operating cycle is defined as the calendar day length of the operating cycle. When the time interval corresponding to the earliest associated risk status is less than or equal to the calendar day length of the operating cycle, the number of operating cycles in which the initial risk is delayed is recorded as one operating cycle. When the time interval corresponding to the earliest associated risk status is greater than the length of the natural days of the operating cycle, the number of operating cycles that the first risk lags behind the operating cycle is determined according to the number of operating cycles that the time interval spans. The remaining natural days that are less than one complete operating cycle are recorded as one operating cycle.

[0046] Based on the date of occurrence of the risk status corresponding to the highest risk associated status, determine the lag of the highest risk status relative to the original health status label, and record the type of the highest risk status, the date of occurrence of the highest risk status, the time interval between the highest risk and the number of operating cycles that the highest risk lags behind.

[0047] The risk state that occurs earliest is selected as the earliest associated risk state, and the risk state with the highest risk level is determined as the highest risk associated state.

[0048] The risk status lag relationship is recorded by combining the enterprise's unified social credit code, industry category, start date of operating cycle, end date of operating cycle, original health status label, earliest associated risk status type, earliest risk status occurrence date, first risk time interval, number of operating cycles after the first risk, highest risk associated status type, highest risk status occurrence date, highest risk time interval, and number of operating cycles after the highest risk.

[0049] S2: Using the lag relationship of risk status as the basis for sample stratification, generate status stratification information for each historical enterprise sample, and calibrate the original health status label with delayed risk label to form a set of historical calibration training samples.

[0050] S2.1: Classify the historical enterprise samples according to the lag relationship of risk status to obtain the historical enterprise sample groups corresponding to different lag situations.

[0051] Using the enterprise's unified social credit code, the start date of the operating cycle, and the end date of the operating cycle as sample identification fields, the risk status lag relationship is matched with historical enterprise samples.

[0052] When a historical enterprise sample is matched with a risk status lag relationship, it is classified according to the matched risk status type, the date of occurrence of the risk status, the time interval, and the number of lag operating cycles, resulting in a group of historical enterprise samples with risk status lag; when a historical enterprise sample is not matched with a risk status lag relationship, it is classified into the risk-free status group.

[0053] By summarizing the historical enterprise sample groups with delayed risk status and the risk-free status group, we can obtain the historical enterprise sample groups corresponding to different delay situations.

[0054] S2.2: Combining the original health status labels and subsequent operational status information in the historical enterprise sample grouping, the historical enterprise samples are divided into the corresponding status levels of the stable health layer, the risk manifestation layer, and the risk delay layer.

[0055] Based on the sample identification fields, historical enterprise samples are grouped, and their original health labels and subsequent operating status information are associated to obtain stratified judgment data for each historical enterprise sample within the corresponding operating cycle.

[0056] When the initial health status label is "good" or "normal," and no risk status is found in subsequent operational status information, the corresponding historical enterprise sample is classified into the "healthy and stable" layer. The "healthy and stable" layer indicates that the historical enterprise sample has not shown a weak operational status during the current operating cycle, and no risk status has appeared in subsequent operational status information.

[0057] When the original health status label is weak, the corresponding historical enterprise samples will be classified into the risk manifestation layer.

[0058] The risk visibility layer is used to indicate that the historical enterprise sample has shown a weak operating state in the current operating cycle, and the division of the risk visibility layer does not require the existence of a risk state in subsequent operating status information as a necessary condition.

[0059] When the original health status label is good or normal, but the subsequent business status information shows a risk status, and the corresponding historical enterprise sample belongs to the historical enterprise sample group with a delayed risk status, the corresponding historical enterprise sample will be classified into the risk delay layer.

[0060] The same historical enterprise sample is divided into different status levels for different operating cycles, and the status levels corresponding to different operating cycles are not interchangeable.

[0061] S2.3: Associate the state hierarchy with the corresponding historical enterprise samples to generate state hierarchy information for each historical enterprise sample.

[0062] The unified social credit code, industry category, start date of the operating cycle, end date of the operating cycle, original health label, subsequent operating status type, risk status type, date of occurrence of risk status, number of delayed operating cycles, and status level of enterprises are associated to form a status stratification record. All status stratification records are summarized to generate status stratification information for each historical enterprise sample.

[0063] S2.4: Based on the status stratification information, identify historical enterprise samples in the risk delay layer, and extract the corresponding original health status labels, subsequent operating status information, and risk status lag relationship.

[0064] Based on the state hierarchy information, the state level corresponding to each historical enterprise sample is read, and the historical enterprise samples with the state level of risk delay layer are identified as historical enterprise samples to be calibrated.

[0065] Based on the sample identification fields, the historical enterprise samples to be calibrated are matched with the original health status label, subsequent operating status information, and risk status lag relationship, respectively, and the original health status label, subsequent operating status type, risk status type, risk status occurrence date, time interval, and number of lag operating cycles corresponding to the historical enterprise samples to be calibrated are extracted.

[0066] S2.5: Based on the risk status lag relationship, determine the lag of the risk status in subsequent operating status information relative to the original health status label, and calibrate the original health status label corresponding to the historical enterprise sample in the risk lag layer to the delayed risk label according to the lag.

[0067] For historical enterprise samples in the risk delay layer, when the original health label is good or normal, and there is a risk status type and a number of delayed operating cycles in the risk status lag relationship, the original health label of the historical enterprise sample is calibrated to a delayed risk label.

[0068] The delayed risk label is used to indicate that a historical enterprise sample was not marked as weak by the original health label during the current operating cycle, but a delayed risk status appeared in subsequent operating status information.

[0069] The risk status type, the date of occurrence of the risk status, the time interval, and the number of delayed operating cycles are used as the calibration basis for the delayed risk label, and are associated with the delayed risk label to generate a delayed risk label record.

[0070] S2.6: Based on the state stratification information, determine the historical enterprise samples in the health and stability layer and the risk manifestation layer, and retain the original health status labels of the historical enterprise samples in the health and stability layer and the risk manifestation layer.

[0071] Based on the state stratification information, the state level corresponding to each historical enterprise sample is read, and the historical enterprise samples with the state level of healthy and stable layer or risk manifestation layer are identified as historical enterprise samples with retained labels.

[0072] For historical enterprise samples in the healthy and stable layer, retain their original health status labels of good or normal; for historical enterprise samples in the risk manifestation layer, retain their original health status labels of weak.

[0073] The original health labels are retained as calibrated health labels for the corresponding historical enterprise samples.

[0074] S2.7: Associate multi-period operating data, time-labeled information, status stratification information, delay risk labels, and retained original health labels to form a set of historical calibration training samples containing training labels.

[0075] Using the enterprise's unified social credit code, industry category, start date of the operating cycle, and end date of the operating cycle as associated fields, multi-cycle operating data, time-marked information, and status hierarchical information are linked to obtain the basic records of historical enterprise samples.

[0076] When the status level corresponding to the basic record of the historical enterprise sample is the risk delay layer, the delayed risk label is used as the training label for the corresponding historical enterprise sample, and the risk status type, the date of occurrence of the risk status, the time interval, and the number of delayed operating cycles are used as the calibration basis for the training label.

[0077] When the status level corresponding to the basic record of the historical enterprise sample is the healthy and stable layer or the risk manifestation layer, the original health label is retained as the training label of the corresponding historical enterprise sample.

[0078] The unified social credit code of the enterprise, industry category, start date of the operating cycle, end date of the operating cycle, multi-cycle operating data, time labeling information, status stratification information, training label and calibration basis of the training label are used to form historical calibration training sample records, which are then summarized to obtain a historical calibration training sample set containing training labels.

[0079] S3: Train the enterprise business health assessment model based on the historical calibration training sample set, and input the current business data of the enterprise to be assessed into the current business characteristics into the enterprise business health assessment model to obtain the initial health assessment information.

[0080] S3.1: Perform feature processing on the multi-cycle operating data in the historical calibration training sample set to obtain historical operating features, and associate them with the delayed risk label and the retained original health label to form a health training sample.

[0081] Based on the historical calibration training sample set, the system retrieves multi-cycle operating data, status stratification information, training labels, and calibration basis for each historical enterprise sample, categorized by enterprise unified social credit code, industry category, start date of operating cycle, and end date of operating cycle. Training labels include delayed risk labels and retained original health labels. The calibration basis for training labels includes risk status type, date of risk status occurrence, time interval, and number of delayed operating cycles.

[0082] Based on industry category and operating cycle cutoff date, historical enterprise samples are divided into peer-to-peer (PTP) sample sets, and multi-cycle operating data are then subjected to feature processing within these PTP sample sets. Operating revenue, operating costs, net profit, net cash flow from operating activities, total assets, total liabilities, accounts receivable balance, and inventory balance are all normalized using percentile rank. Within a sample set of peers and during the same period, the same operating data item is sorted from smallest to largest. When identical values ​​exist, the average rank of the corresponding rankings is taken as the final ranking. The quantile characteristic value is expressed as follows: ; In the formula, For the first quantile characteristic values ​​of a sample of historical enterprises under the corresponding operating data item. For the first The ranking of historical enterprise samples within a set of industry peers and period-specific samples is obtained by sorting the corresponding operating data items from smallest to largest. This represents the total number of samples in the same industry and period sample set. This refers to the historical enterprise sample number in the sample set of peers with the same business cycle.

[0083] When multiple historical enterprise samples have the same value under the same operating data item, the ranking order is... Take the average rank of the rankings corresponding to the same value; combine the quantile feature values ​​corresponding to each operating data item in a fixed field order to obtain the operating numerical features.

[0084] Tax declaration status, business registration status, administrative penalty records, and abnormal business operation records are classified and coded. Tax declaration status is divided into normal declaration status and abnormal declaration status; business registration status is divided into normal existence status and abnormal existence status; administrative penalty records are divided into no administrative penalty records, completed administrative penalty records, and incomplete administrative penalty records; and abnormal business operation records are divided into no abnormal business operation records, removed abnormal business operation records, and not removed abnormal business operation records. Each status is represented by a fixed coding position. When a status matches the corresponding coding position, the corresponding coding position is recorded as one, and the remaining coding positions are recorded as zero. The coding results corresponding to each business status item are combined in a fixed field order to obtain the business status characteristics.

[0085] The system performs adjacent operating cycle variation processing on operating revenue, net profit, and net cash flow from operating activities. It reads the operating revenue, net profit, and net cash flow from operating activities for the current operating cycle and the adjacent previous operating cycle from the same historical enterprise sample. It calculates the increase / decrease difference of the current operating cycle data relative to the adjacent previous operating cycle data. Within the same industry and same cycle sample set, it performs percentile rank normalization on the increase / decrease difference to obtain the corresponding change quantile feature value. When data from the adjacent previous operating cycle is missing, the corresponding increase / decrease difference is recorded as zero, and a previous cycle missing identifier is generated. When data from the adjacent previous operating cycle is not missing, the previous cycle missing identifier is recorded as zero. The change quantile feature values ​​corresponding to each cycle variation item and the previous cycle missing identifier are combined in a fixed field order to obtain the cycle variation characteristics.

[0086] Historical operating characteristics are obtained by concatenating the operating numerical characteristics, operating status characteristics, and periodic change characteristics in a fixed field order. The fixed field order is as follows: operating numerical characteristics, operating status characteristics, and periodic change characteristics.

[0087] The training weights of the corresponding historical enterprise samples are determined based on the calibration criteria of the training labels. Among them, the training labels are historical enterprise samples with delayed risk labels, and the training weights of the samples are determined based on the risk status type, time interval, and number of delayed operating cycles. The higher the degree of risk status, the shorter the time interval, or the fewer the number of delayed operating cycles, the higher the training weight of the sample. The training labels are historical enterprise samples with retained original health labels, and the basic sample training weights are used.

[0088] By associating historical operating characteristics, status stratification information, training labels, and sample training weights, a health training sample is formed.

[0089] S3.2: Based on the health training samples, perform machine learning training on the enterprise operation health assessment model to obtain the health training output for the health training samples.

[0090] The model structure of the enterprise business health assessment model is determined, and the enterprise business health assessment model adopts a multi-layer feedforward neural network structure.

[0091] The hierarchical connections of the enterprise operational health assessment model are established. The input layer is fully connected to the first hidden layer, the first hidden layer is fully connected to the second hidden layer, and the second hidden layer is fully connected to the output layer. The input layer receives historical operational features arranged in a fixed field order. The first hidden layer extracts primary operational characteristics from these historical features. The second hidden layer performs a combined mapping of the primary operational characteristics. The output layer maps the secondary operational characteristics to the classification probabilities corresponding to each calibrated health label.

[0092] The number of input nodes in the input layer is consistent with the number of fields in the historical operating characteristics, and the order of the input nodes is consistent with the order of the fixed fields; the number of output nodes in the output layer is consistent with the number of categories of the calibrated health label, and each output node corresponds to one type of calibrated health label.

[0093] The number of input nodes, the number of calibrated health label categories, and the number of health training samples are used as the basis for determining the number of nodes, constructing candidate combinations for the number of nodes in the first and second hidden layers. The enterprise operational health assessment model is trained using different candidate combinations, and the classification loss of the validation samples for each candidate combination is calculated. The candidate combination with the lowest classification loss of the validation samples is determined as the number of nodes in the first and second hidden layers.

[0094] When multiple candidate combinations have the same classification loss for the validation samples, select the candidate combination with the smaller total number of hidden layer nodes to reduce the risk of overfitting.

[0095] The health score training samples are divided into training samples and validation samples based on the enterprise's unified social credit code. Samples corresponding to the same enterprise's unified social credit code are only assigned to either the training samples or the validation samples to avoid samples from the same enterprise participating in both training and validation simultaneously.

[0096] Historical operating characteristics from the training samples are input into the enterprise operating health assessment model. The first hidden layer performs a weighted summation of the historical operating characteristics and adds the bias parameters of the first hidden layer to obtain the first weighted result. The first hidden layer performs modified linear unit activation processing on the first weighted result, setting values ​​less than zero to zero and keeping values ​​greater than or equal to zero unchanged, to obtain the first operating characterization result.

[0097] The second hidden layer performs a weighted summation on the first business representation result and adds the bias parameters of the second hidden layer to obtain a second weighted result. The second hidden layer performs a modified linear unit activation process on the second weighted result, setting values ​​less than zero to zero and keeping values ​​greater than or equal to zero unchanged, thus obtaining the second business representation result.

[0098] The output layer performs a weighted summation of the second operational representation results and superimposes the bias parameters of the output layer to obtain the output results corresponding to each calibrated health label. The output layer performs classification normalization processing on each output result, converting each output result into a corresponding classification probability, and determines the calibrated health label with the highest classification probability as the health training output.

[0099] The calibrated health labels are converted into corresponding supervised targets. Cross-entropy loss is used to calculate the classification loss between the health training output and the supervised targets. The classification probability of the enterprise operational health assessment model for the actual calibrated health label output is read, and the negative logarithm of the classification probability is calculated. This negative logarithmic loss is then weighted according to the training weights corresponding to the health training samples, and finally averaged over the training samples to obtain the classification loss. If no training weights are set for the health training samples, the corresponding training weights are recorded as the base weights.

[0100] Based on classification loss, the weight parameters and bias parameters of the enterprise business health assessment model are updated in reverse. The weight parameters include the weight parameters from the input layer to the first hidden layer, the weight parameters from the first hidden layer to the second hidden layer, and the weight parameters from the second hidden layer to the output layer; the bias parameters include the bias parameters of the first hidden layer, the bias parameters of the second hidden layer, and the bias parameters of the output layer.

[0101] Repeatedly execute the process of inputting historical business features, generating health training output, calculating classification loss, and updating model parameters, and record the classification loss of the validation samples and the health training output of the validation samples after each training round.

[0102] The classification loss of the validation samples in the current training epoch is compared with the recorded lowest classification loss of the validation samples. If the classification loss of the validation samples in the current training epoch is lower than the recorded lowest classification loss of the validation samples, the lowest classification loss of the validation samples is updated, and the model parameter update for the next training epoch is performed.

[0103] For example, if the classification loss of the validation samples in three consecutive training rounds is not lower than the lowest recorded classification loss of the validation samples, and the health training output corresponding to the same validation sample in three consecutive training rounds remains consistent, the enterprise business health assessment model is determined to have reached convergence, training is stopped, and the preliminary trained enterprise business health assessment model is obtained; otherwise, the model parameter update continues.

[0104] The historical operating characteristics from the health training samples are input into the pre-trained enterprise operating health assessment model to obtain the health training output for the health training samples.

[0105] S3.3: Based on state hierarchical information, perform hierarchical verification of the deviation of the health training output in the health stability layer, risk manifestation layer and risk delay layer to determine the evaluation bias of the enterprise operation health assessment model.

[0106] Based on the health training output, state hierarchical information, and training labels, the data is associated with the enterprise's unified social credit code, the start date of the operating cycle, and the end date of the operating cycle to obtain the state level, training label, and health training output corresponding to each health training sample.

[0107] Based on the state level, the health training samples are assigned to the health stability layer verification set, the risk manifestation layer verification set, and the risk delay layer verification set, respectively. In each verification set, the health training output is compared with the corresponding training label. If the two are consistent, it is recorded as verification consistency; if they are inconsistent, it is recorded as verification deviation.

[0108] For health training samples that exhibit verification deviations, the risk level is determined in the order of good, normal, weak, and delayed risk labels. When the risk level corresponding to the health training output is lower than the risk level corresponding to the training label, it is recorded as a risk underestimation deviation. When the risk level corresponding to the health training output is higher than the risk level corresponding to the training label, it is recorded as a risk overestimation deviation.

[0109] The verification deviation results of each state level are summarized to obtain the deviation sample, deviation state level, health training output, training label and deviation direction, and the summarized results are determined as the evaluation deviation of the enterprise business health assessment model.

[0110] S3.4: Based on the assessment bias, identify historical enterprise samples with a tendency to misjudge from the historical calibration training sample set, and use these historical enterprise samples with a tendency to misjudge to continue training the enterprise business health assessment model, thus obtaining the trained enterprise business health assessment model.

[0111] Based on the deviation samples, deviation status levels, health training outputs, training labels, and deviation directions in the evaluation bias, historical enterprise samples with a tendency to misjudge are identified.

[0112] Historical enterprise samples with deviations in the direction of risk underestimation are identified as historical enterprise samples with a tendency to misjudge; historical enterprise samples with verification deviations in the risk delay layer are identified as historical enterprise samples with a tendency to misjudge; when the same historical enterprise sample appears repeatedly, only one historical enterprise sample record with a tendency to misjudge is retained.

[0113] Based on the enterprise's unified social credit code, the start date of the operating cycle, and the end date of the operating cycle, historical operating characteristics, status stratification information, and training labels corresponding to historical enterprise samples with a tendency to misjudge are extracted from the historical calibration training sample set to form training samples with a tendency to misjudge.

[0114] Based on the training samples with a tendency to misjudge, the initially trained enterprise health assessment model is further trained. The further training uses the same classification loss, optimization method, and stopping condition as the previous model training, without adjusting the training weights by repeatedly retaining samples. After the further training is completed, the trained enterprise health assessment model is obtained.

[0115] S3.5: Obtain the current operating data of the enterprise to be evaluated, perform feature processing on the current operating data, and obtain the current operating characteristics.

[0116] Obtain the Unified Social Credit Code, industry category, current operating cycle, and current operating data of the company to be evaluated. Current operating data includes operating revenue, operating costs, net profit, net cash flow from operating activities, total assets, total liabilities, accounts receivable balance, inventory balance, tax declaration status, business registration status, administrative penalty records, and records of abnormal operations.

[0117] Based on the industry category and current operating cycle of the company to be evaluated, historical company samples with the same industry category and operating cycle cutoff date corresponding to the current operating cycle are selected from the historical calibration training sample set to form the peer sample set corresponding to the company to be evaluated; if there is no corresponding operating cycle cutoff date, the peer sample set with the operating cycle cutoff date earlier than the current operating cycle and the closest time is selected as the peer sample set corresponding to the company to be evaluated.

[0118] The current operating data is processed using the same feature processing method as the historical operating features to obtain the current operating numerical features, current operating status features, and current periodic change features. The current operating numerical features, current operating status features, and current periodic change features are then concatenated in a fixed field order to obtain the current operating features.

[0119] S3.6: Input the current operating characteristics into the trained enterprise operating health assessment model, identify the health status and risk status of the current operating characteristics, obtain health status judgment information and risk status judgment information, and form the initial health assessment information of the enterprise to be assessed.

[0120] The current business characteristics are input into the trained enterprise business health assessment model, and processed layer by layer in the order of connection between the input layer, the first hidden layer, the second hidden layer and the output layer to obtain the classification output result corresponding to the current business characteristics.

[0121] The health status assessment information is determined based on the classification output results: when the classification output results correspond to good, normal, or weak, the health status assessment information is recorded as good, normal, or weak, respectively; when the classification output results correspond to the delay risk label, the health status assessment information is recorded as having delay risk.

[0122] The risk status judgment information is determined based on the classification output results: when the classification output results correspond to good or normal, the risk status judgment information is recorded as no risk status is identified; when the classification output results correspond to weak, the risk status judgment information is recorded as a manifest risk status is identified; when the classification output results correspond to the delayed risk label, the risk status judgment information is recorded as a delayed risk status is identified.

[0123] The initial health assessment information of the enterprise to be assessed is formed by associating its unified social credit code, industry category, current operating cycle, current operating characteristics, health status assessment information, and risk status assessment information.

[0124] It should be noted that the health status assessment information includes good, normal, weak, and risk of delay. After correction by the sample validation set, the health status assessment information may also include concern about the risk of delay.

[0125] S4: Based on the current operating characteristics and initial health assessment information, retrieve support samples that are similar to the current operating characteristics and have the same health status from the historical calibration training sample set, as well as counterexample samples that are similar to the current operating characteristics but have a different health status, to form a sample verification set.

[0126] S4.1: Align the current business characteristics with the historical business characteristics in the historical calibration training sample set by feature dimension to obtain business characteristic comparison information. Based on the business characteristic comparison information, retrieve historical enterprise samples that are similar to the current business characteristics in the historical calibration training sample set to obtain similar historical enterprise samples.

[0127] Based on the current operating characteristics and the historical operating characteristics, state stratification information, training labels, subsequent operating status information and risk status lag relationship in the historical calibration training sample set, the operating characteristics are compared.

[0128] Align the current business feature with each historical business feature in a fixed field order; retain the same fields when they exist; and fill in zero values ​​in the business feature that does not contain a field if any business feature has a field that another business feature does not contain.

[0129] Calculate the feature matching degree for the current operating feature and each historical operating feature after feature dimension alignment. The cosine matching degree is used for determination, expressed as: ; In the formula, The feature matching degree between current operating characteristics and historical operating characteristics. The first of the current operating characteristics The field values ​​corresponding to each field The first in historical management characteristics The field values ​​corresponding to each field; To complete the alignment of the operational feature field sequence number, The total number of operational feature fields after feature dimension alignment.

[0130] When all field values ​​in the current or historical operating characteristics are zero, the corresponding historical enterprise sample will not be identified as a similar historical enterprise sample.

[0131] The business feature comparison information is composed of the enterprise's unified social credit code, industry category, start date of business cycle, end date of business cycle, current business features after feature dimension alignment, historical business features after feature dimension alignment, and feature matching degree.

[0132] The feature matching condition is that the historical enterprise samples are sorted from high to low according to the feature matching degree and are located in the first quarter position range; when the number of historical enterprise samples cannot be divided into four equal parts, the first quarter position range is used as the standard; when there are the same feature matching degree at the dividing position, the historical enterprise samples corresponding to the same feature matching degree all meet the feature matching condition.

[0133] Historical enterprise samples that meet the feature matching criteria are identified as similar historical enterprise samples.

[0134] S4.2: Compare the training labels of similar historical enterprise samples with the initial health assessment information to identify supporting samples with similar current operating characteristics and consistent health status.

[0135] Consistency comparison is performed based on the training labels, state hierarchical information, subsequent operating status information, risk status lag relationship and feature matching degree corresponding to similar historical enterprise samples.

[0136] Establish a correspondence between training labels and health status judgment information. When the training label is good, normal, or weak, it corresponds to good, normal, or weak in the health status judgment information, respectively. When the training label is a delay risk label, it corresponds to the existence of delay risk in the health status judgment information.

[0137] Health status judgment information and risk status judgment information are determined from the initial health assessment information. The training labels of similar historical enterprise samples are converted into label health status and compared with the health status judgment information for consistency.

[0138] When the health status of the label is consistent with the health status judgment information, the corresponding similar historical enterprise samples are identified as supporting samples; when the training label is a delay risk label and the risk status judgment information is that the delay risk status has been identified, the corresponding similar historical enterprise samples are identified as supporting samples with consistent delay risk identification.

[0139] The supporting samples are associated with the enterprise's unified social credit code, industry category, start date of the operating cycle, end date of the operating cycle, feature matching degree, training label, label health status, status stratification information, health status judgment information, risk status judgment information, subsequent operating status information, and risk status lag relationship to form a supporting sample record, which is then summarized to obtain the supporting samples.

[0140] S4.3 performs a deviation comparison between the training labels of similar historical enterprise samples and the initial health assessment information to identify counterexample samples that have similar current operating characteristics but deviate from the health status.

[0141] Deviation comparison is performed based on similar historical enterprise samples that were not identified as supporting samples, along with their corresponding training labels, state hierarchical information, subsequent operational status information, risk status lag relationship, and feature matching degree.

[0142] Based on the correspondence between training labels and health status judgment information, the training labels of similar historical enterprise samples that were not identified as supporting samples are converted into label health status, and a deviation comparison is performed with the health status judgment information in the initial health assessment information.

[0143] When the health status of the label is inconsistent with the health status judgment information, the corresponding similar historical enterprise samples will be identified as negative examples.

[0144] When the training label is a delay risk label, and the risk status judgment information does not identify the delay risk status, the corresponding similar historical enterprise samples will be identified as negative examples of delay risk identification deviation.

[0145] The risk level is determined in the order of good, normal, weak, and with potential delay. If the risk level corresponding to the labeled health status is higher than the risk level corresponding to the health status assessment information, the negative example is recorded as a risk underestimation negative example; if the risk level corresponding to the labeled health status is lower than the risk level corresponding to the health status assessment information, the negative example is recorded as a risk overestimation negative example.

[0146] For counterexamples of underestimating risk, the corresponding subsequent operating status information and risk status lag relationship are read, and the risk status type, the date of occurrence of the risk status, the time interval, and the number of lag operating cycles are used as subsequent risk delay characteristics.

[0147] The negative example samples are associated with the enterprise's unified social credit code, industry category, start date of the operating cycle, end date of the operating cycle, feature matching degree, training label, label health status, status stratification information, health status judgment information, risk status judgment information, deviation type, subsequent operating status information, and risk status lag relationship to form a negative example sample record, which is then summarized to obtain the negative example sample.

[0148] S4.4: The supporting samples and their corresponding state levels are used as the consistency verification content, and the negative sample, its corresponding state level, and the risk state lag relationship of the negative sample are used as the deviation verification content, forming a sample verification set.

[0149] Read the supporting sample records and combine the supporting samples, feature matching degree, training labels, label health status, status hierarchical information, subsequent operating status information and risk status lag relationship into consistency verification content.

[0150] Read the negative example sample record and combine the negative example sample, feature matching degree, training label, label health status, status stratification information, deviation type, subsequent operating status information and risk status lag relationship into deviation verification content.

[0151] The unified social credit code, industry category, current operating cycle, current operating characteristics, initial health assessment information, operating characteristic comparison information, consistency verification content and deviation verification content of the enterprise to be evaluated are associated to form sample verification records, and the sample verification set is obtained by summarizing them.

[0152] S5: Based on the sample validation set, determine the risk of misjudgment of the initial health assessment information, correct the initial health assessment information, and obtain the business health assessment conclusion of the enterprise to be assessed.

[0153] S5.1: Based on the consistency verification content and deviation verification content in the sample verification set, determine the degree of support and deviation for the initial health assessment information, respectively.

[0154] Based on the supporting samples in the consistency check content and the negative samples in the deviation check content, the number of supporting samples and the number of negative samples are counted. When the number of supporting samples and the number of negative samples are not zero, the average feature matching degree of the supporting samples and the average feature matching degree of the negative samples are calculated respectively, as follows: ; In the formula, To support the average feature matching degree of the samples, For the first Feature matching degree of each supporting sample To support the sample size, To support sample serial numbers, The average feature matching degree of the counterexample samples. For the first Feature matching degree of each counterexample sample The number of counterexample samples. The counterexample sample number; when or When the value is zero, the corresponding average feature matching degree is not included in the comparison.

[0155] When the number of supporting samples and the number of counterexample samples When both are zero, it means that there are no similar historical enterprise samples in the sample verification set that can be used for consistency verification and deviation verification. The degree of support and the degree of deviation are both recorded as uncertain, and the sample verification status is recorded as insufficient samples.

[0156] When the number of supporting samples The number of non-zero and counterexample samples When the value is zero, the degree of support is recorded as strong support, and the degree of deviation is recorded as weak deviation.

[0157] When the number of supporting samples The number of negative examples is zero. When the value is not zero, the degree of support is recorded as weak support, and the degree of deviation is recorded as strong deviation.

[0158] When the number of supporting samples and the number of counterexample samples When none of them are zero, the degree of support and the degree of deviation are determined based on the number of supporting samples, the number of counterexample samples, the average feature matching degree of supporting samples, and the average feature matching degree of counterexample samples.

[0159] when Greater than When all are non-zero, the support level is defined as strong support based on the number of supporting samples, the number of counterexample samples, and the average feature matching degree of the supporting samples, provided that the average feature matching degree of the supporting samples is not lower than the average feature matching degree of the counterexample samples; when... Less than If the average feature matching degree of the supporting samples is lower than that of the average feature matching degree of the counterexample samples, the support level is recorded as weak support; otherwise, it is recorded as general support.

[0160] when Greater than When the average feature matching degree of the counterexample samples is not lower than that of the average feature matching degree of the supporting samples, the degree of deviation is recorded as strong deviation; when Less than If the average feature matching degree of the counterexample samples is lower than that of the average feature matching degree of the supporting samples, the degree of deviation is recorded as weak deviation; otherwise, it is recorded as general deviation.

[0161] S5.2: Based on the state level corresponding to the counterexample sample, determine the distribution of the counterexample sample in the risk delay layer, and in combination with the risk state lag relationship of the counterexample sample, identify the omission of delayed risk in the initial health assessment information.

[0162] Based on the deviation verification content, according to the state stratification information corresponding to the counterexample samples, the counterexample samples are respectively classified into the health and stability layer counterexample set, the risk manifestation layer counterexample set, and the risk delay layer counterexample set, and the number of counterexample samples in each counterexample set is counted.

[0163] When the number of counterexamples in the risk delay layer counterexample set is greater than the number of counterexamples in the health and stability layer counterexample set and the risk manifestation layer counterexample set, the risk delay layer distribution is determined to be a concentrated distribution; when the number of counterexamples in the risk delay layer counterexample set is not zero and does not reach the concentrated distribution, the risk delay layer distribution is determined to be a distribution with existence; when the number of counterexamples in the risk delay layer counterexample set is zero, the risk delay layer distribution is determined to be a distribution without distribution.

[0164] Based on the deviation type and risk status lag relationship of counterexample samples in the risk delay layer counterexample set, situations where delayed risks are omitted are identified. If there are counterexamples of risk underestimation in the risk delay layer counterexample set, and the risk status judgment information in the initial health assessment information does not identify a delayed risk status, it is determined that there is a delay risk omission in the initial health assessment information; if there are no counterexamples of risk underestimation in the risk delay layer counterexample set, or if the risk status judgment information in the initial health assessment information has identified a delayed risk status, it is determined that there is no delay risk omission in the initial health assessment information.

[0165] S5.3: Correlate the degree of support, the degree of deviation, and the omission of delay risk to obtain the risk of misjudgment in the initial health assessment information.

[0166] By correlating the degree of support, the degree of deviation, the distribution of risk delay layers, and the omission of delay risks, a record of misjudgment risk assessment is formed.

[0167] When there is a delay risk omission, and the support level is weak, the deviation level is strong, and the risk delay layer distribution is concentrated, the misjudgment risk is identified as a delay risk omission misjudgment risk, and the misjudgment type is identified as a delay risk omission type; when there is a delay risk omission but the conditions for identifying delay risk omission misjudgment risk are not met, the misjudgment risk is identified as a delay risk suspected misjudgment risk, and the misjudgment type is identified as a delay risk suspected omission type.

[0168] When there is no omission of delay risk and the deviation is strong, the misjudgment risk is identified as a health status deviation misjudgment risk, and the misjudgment type is identified as a health status deviation type; when there is no omission of delay risk and the deviation is moderate, the misjudgment risk is identified as a health status suspected deviation risk, and the misjudgment type is identified as a health status suspected deviation type.

[0169] When there is no delay risk omission and the deviation is weak, the misjudgment risk is determined as low misjudgment risk, and the misjudgment type is determined as low misjudgment type; The low risk of misjudgment is only used as a verification prompt and does not trigger the correction of the health status judgment information and risk status judgment information in the initial health assessment information.

[0170] The initial health assessment information is obtained by associating the enterprise's unified social credit code, industry category, current operating cycle, level of support, degree of deviation, distribution of risk delay layer, omission of delay risk, misjudgment risk, and type of misjudgment with the enterprise to be assessed.

[0171] S5.4: Based on the risk of misjudgment in the initial health assessment information, identify the health status judgments that may be delayed or omitted in the initial health assessment information, correct the initial health assessment information according to the type of misjudgment, obtain the corrected health assessment information, and use it as the conclusion of the business health assessment of the enterprise to be assessed.

[0172] Based on the initial health assessment information, the risk of misjudgment, and the type of misjudgment, the correction methods for the health status assessment information and the risk status assessment information are determined.

[0173] When the misjudgment type is delayed risk omission or suspected delayed risk omission, the health status judgment information in the initial health assessment information is determined as a health status judgment with delayed risk omission. Based on the counterexample samples in the risk delay layer counterexample set whose deviation type is risk underestimation, the corrected risk status type is determined according to the risk status type in the risk status lag relationship. When there are multiple risk status types, the risk status type that occurs most frequently is selected as the corrected risk status type; when the frequency of occurrence is the same, the risk status type with the shortest lag operating cycle is selected as the corrected risk status type.

[0174] When the misjudgment type is delayed risk omission, the health status judgment information is corrected to indicate that delayed risk exists, the risk status judgment information is corrected to indicate that delayed risk status has been identified, and the corrected risk status type is used as the specific risk type of delayed risk status.

[0175] When the misjudgment type is suspected omission of delayed risk, the health status judgment information is corrected to "there is delayed risk concern", the risk status judgment information is corrected to "delay risk concern status is identified", and the corrected risk status type is used as the specific risk type of the delayed risk concern status. Among them, "there is delayed risk concern" is used to indicate that the initial health assessment information has not yet reached the level of delayed risk omission misjudgment risk, but the counterexample sample shows that there is a delayed risk status that needs attention.

[0176] When the misjudgment type is health status deviation, the health status judgment information is corrected based on the health status label that appears most frequently in the counterexample samples, and a corrected risk status judgment information is generated based on the corrected health status judgment information; when the corrected health status judgment information is good or normal, the corrected risk status judgment information is no risk status identified; when the corrected health status judgment information is weak, the corrected risk status judgment information is a manifest risk status identified; when the corrected health status judgment information is a delayed risk, the corrected risk status judgment information is a delayed risk status identified.

[0177] When the misjudgment type is suspected deviation of health status, the health status judgment information and risk status judgment information in the initial health assessment information are retained, and the label health status corresponding to the misjudgment risk and the counterexample sample is used as risk warning information.

[0178] When the misjudgment type is low, the health status judgment information and risk status judgment information in the initial health assessment information are retained, and no status correction is performed.

[0179] The unified social credit code, industry category, current operating cycle, current operating characteristics, initial health assessment information, misjudged risk, misjudgment type, corrected risk status type, corrected health status judgment information, corrected risk status judgment information, and risk warning information of the enterprise to be assessed are linked to obtain the corrected health assessment information, and the corrected health assessment information is used as the conclusion of the enterprise's operational health assessment.

[0180] In summary, this invention, by retrieving support samples and counterexamples based on current operational characteristics and initial health assessment information, can form a sample verification set that combines the ability to verify consistent states and identify deviations. This provides more comprehensive historical sample evidence for the enterprises to be assessed. The sample verification set characterizes similar business models and state differentiation features, marks potential risk boundaries, verifies the initial assessment results, improves the ability to identify complex operational states and risks, and enhances the stability and credibility of the operational health assessment conclusions.

[0181] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A machine learning-based method for assessing enterprise operational health, characterized in that, include: Acquire multi-period operating data and corresponding subsequent operating status information of historical enterprise samples, time-annotate the multi-period operating data and subsequent operating status information, and determine the lag relationship between the original health label and the risk status. Using the lag relationship of risk status as the basis for sample stratification, status stratification information of each historical enterprise sample is generated, and the original health status label is calibrated with delayed risk label to form a set of historical calibration training samples. The enterprise business health assessment model is trained based on a historical calibration training sample set. The current business data of the enterprise to be assessed is characterized into current business characteristics and then input into the enterprise business health assessment model to obtain initial health assessment information. Based on the current operating characteristics and initial health assessment information, support samples that are similar to the current operating characteristics and have the same health status, as well as counterexample samples that are similar to the current operating characteristics but deviate from the health status, are retrieved from the historical calibration training sample set to form a sample verification set; Based on the sample validation set, the risk of misjudgment of the initial health assessment information is determined, and the initial health assessment information is corrected to obtain the conclusion of the business health assessment of the enterprise to be assessed.

2. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The time-annotation of multi-period operating data and subsequent operating status information refers to obtaining multi-period operating data and corresponding subsequent operating status information of historical enterprise samples, and then time-annotating the multi-period operating data and subsequent operating status information according to the chronological relationship between the multi-period operating data and the subsequent operating status information to obtain time-annotated information.

3. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The determination of the lag relationship between the original health status label and the risk status includes: Based on multi-period operating data of historical enterprise samples, the corresponding operating status is determined and initial health status is marked to obtain the original health status label; Based on time-labeled information, the original health status label is sequentially associated with subsequent business status information. The lag of risk status in subsequent business status information relative to the original health status label is identified and associated with the original health status label to obtain the risk status lag relationship.

4. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The generated status hierarchy information for each historical enterprise sample includes: The historical enterprise samples were categorized according to the lag relationship of risk status, resulting in historical enterprise sample groups corresponding to different lag situations. By combining the original health status labels and subsequent operational status information in the historical enterprise sample groups, the historical enterprise samples are divided into the corresponding status levels of the stable health layer, the risk manifestation layer, and the risk delay layer, and then associated with the corresponding historical enterprise samples to generate status stratification information for each historical enterprise sample.

5. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The specific steps for forming the historical calibration training sample set are as follows: Based on the status stratification information, we identify historical enterprise samples that are in the risk delay layer and extract the corresponding original health status labels, subsequent operating status information and risk status lag relationship. Based on the lag relationship of risk status, determine the lag of risk status in subsequent operational status information relative to the original health status label, and calibrate the original health status label corresponding to the historical enterprise sample in the risk lag layer to the delayed risk label according to the lag. Based on the state stratification information, historical enterprise samples in the health and stability layer and the risk manifestation layer are determined, and the original health labels of historical enterprise samples in the health and stability layer and the risk manifestation layer are retained. By associating multi-period operating data, time-stamped information, status stratification information, delay risk labels, and retained original health labels, a set of historical calibration training samples containing training labels is formed.

6. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The enterprise operational health assessment model trained based on a historical calibration training sample set includes: The multi-cycle operating data in the historical calibration training sample set is characterized to obtain historical operating features, which are then associated with the delayed risk label and the retained original health label to form a health training sample. The enterprise operational health assessment model is trained based on the health training samples to obtain the health training output for the health training samples. Based on state hierarchical information, the deviation of the health training output in the health stability layer, risk manifestation layer and risk delay layer is verified in a hierarchical manner to determine the evaluation bias of the enterprise business health assessment model. Based on the assessment bias, identify historical enterprise samples with a tendency to misjudge from the historical calibration training sample set, and use these historical enterprise samples with a tendency to misjudge to continue training the enterprise business health assessment model, thus obtaining the trained enterprise business health assessment model.

7. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The initial health assessment information obtained includes: Obtain the current operating data of the enterprise to be evaluated, perform feature processing on the current operating data, and obtain the current operating characteristics; The current operating characteristics are input into the trained enterprise operating health assessment model, which identifies the health status and risk status of the current operating characteristics, obtains health status judgment information and risk status judgment information, and forms the initial health assessment information of the enterprise to be assessed.

8. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The specific steps for forming the sample verification set are as follows: Align the current business characteristics with the historical business characteristics in the historical calibration training sample set by feature dimensions to obtain business characteristic comparison information. Based on the business characteristic comparison information, retrieve historical enterprise samples that are similar to the current business characteristics in the historical calibration training sample set to obtain similar historical enterprise samples. The training labels of similar historical enterprise samples are compared with the initial health assessment information to identify supporting samples with similar current operating characteristics and consistent health status. By comparing the training labels of similar historical enterprise samples with the initial health assessment information, negative examples with similar current operating characteristics but deviating health status are identified. The supporting samples and their corresponding state levels are used as the consistency verification content, while the negative samples, their corresponding state levels, and the risk state lag relationship of the negative samples are used as the deviation verification content, thus forming a sample verification set.

9. The enterprise operational health assessment method based on machine learning as described in claim 1, characterized in that, The risk of misjudgment in determining the initial health assessment information based on the sample validation set includes: Based on the consistency and deviation verification content in the sample verification set, the degree of support and deviation for the initial health assessment information are determined respectively. Based on the state level corresponding to the counterexample samples, determine the distribution of counterexample samples in the risk delay layer, and combine the risk state lag relationship of the counterexample samples to identify the omission of delayed risks in the initial health assessment information. By correlating the degree of support, the degree of deviation, and the omission of delay risk, the risk of misjudgment in the initial health assessment information is obtained.

10. The enterprise operational health assessment method based on machine learning as described in claim 9, characterized in that, The conclusion of the business health assessment of the enterprise to be assessed refers to determining the health status judgment that was delayed or omitted in the initial health assessment information based on the risk of misjudgment from the initial health assessment information. The initial health assessment information is corrected according to the type of misjudgment to obtain the corrected health assessment information, which is then used as the conclusion of the business health assessment of the enterprise to be assessed.