Method for predicting enterprise state based on residual network and multiple attention mechanisms
By combining the prediction model of residual network and multi-head attention mechanism, the problem of low accuracy of enterprise status prediction in the prior art is solved, and higher accuracy and timeliness are achieved, and more accurate risk assessment and decision-making support is provided.
Patent Information
- Application Number
- CN202510212722.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
When predicting the status of an enterprise, the accuracy of the existing technology is low, making it difficult to identify future abnormalities in the enterprise in advance, resulting in delayed alarms and unable to meet the needs of early warning.
A prediction model based on residual network and multi-head attention mechanism is adopted, and the business information of the enterprise, public information and social related information are obtained, and segmentation and conversion are performed according to the preset time period, enterprise information characteristics are extracted, and state prediction is performed using the time series multi-head attention mechanism.
It significantly improves the accuracy and timeliness of enterprise status prediction, can more accurately reflect the development trend of the enterprise, provide strong data support and more accurate risk assessment tools.
Smart Images

Figure CN120047230A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of predicting enterprise status. Specifically, it relates to a method, device, computer-readable storage medium, and electronic device for predicting enterprise status based on a residual network and a multi-head attention mechanism. Background Art
[0002] Currently, when conducting business, the operating status of an enterprise is usually determined through manual review. This can only judge the current operating status of the enterprise and it is difficult to identify the probability of the enterprise being abnormal in a future period. For some of our businesses (such as forward, swap, etc.), when the enterprise is already in an abnormal state, there will be a certain lag in issuing an alarm, which cannot meet the function of early warning.
[0003] In the prior art, whether using traditional statistical solutions or artificial intelligence solutions, only existing indicators of the enterprise, indicators of capital risk resistance ability, stock analysis, etc. are collected and rated, sorted, or trained for a model. However, these data do not fully consider the impact of time on an enterprise, nor do they consider the impact on the development trend of the enterprise, and do not correlate and analyze the data of the enterprise at different times. Therefore, the obtained result is only the current state of the enterprise rather than the state of the enterprise over a period of time, and it cannot play the role of early warning when an accident occurs. Summary of the Invention
[0004] The main objective of this application is to provide a method, device, computer-readable storage medium, and electronic device for predicting enterprise status based on a residual network and a multi-head attention mechanism, so as to at least solve the problem of low prediction accuracy in existing solutions for predicting enterprise status.
[0005] To achieve the above objective, according to one aspect of this application, a method for predicting enterprise status based on a residual network and a multi-head attention mechanism is provided, including: obtaining enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes operating information, social public information, and social association information; performing segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and using a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features; using an enterprise status prediction model based on a temporal multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise.
[0006] Optionally, an enterprise status prediction model based on a temporal multi-head attention mechanism is adopted to predict the enterprise status of the enterprise according to the enterprise information features, and a prediction result of the enterprise is obtained, including: performing matrix segmentation processing on the enterprise information features by using the temporal multi-head attention mechanism to obtain a plurality of segmented feature matrices, where the segmented feature matrices include a query matrix, a key matrix, and a value matrix; calculating attention scores according to the plurality of segmented feature matrices to obtain target attention scores, and using the enterprise status prediction model to perform prediction processing on the enterprise status of the enterprise according to the target attention scores to obtain the prediction result of the enterprise.
[0007] Optionally, calculating the attention scores according to the plurality of segmented feature matrices includes: using an attention calculation formula: to calculate the attention score Attention(Q, K, V), where Q is a query matrix vector, K is a key matrix vector, V is a value matrix vector, and d k is the dimension.
[0008] Optionally, after predicting the enterprise status of the enterprise according to the enterprise information features to obtain the prediction result of the enterprise, the method further includes: obtaining the actual enterprise status corresponding to the time period of the prediction result of the enterprise; comparing the prediction result with the actual enterprise status information to generate a comparison result, and updating and iterating the enterprise status prediction model according to the comparison result.
[0009] Optionally, before predicting the enterprise status of the enterprise according to the enterprise information features by using an enterprise status prediction model based on a temporal multi-head attention mechanism, the method further includes: obtaining external environment data and enterprise characteristic data, where the external environment data includes industry dynamic data and economic indicator data, and the enterprise characteristic data includes an enterprise operation cycle and an enterprise industry characteristic; optimizing the prediction interval of the enterprise status prediction model by combining the external environment data and the enterprise characteristic data to obtain the optimized enterprise status prediction model.
[0010] Optionally, before segmenting and converting the enterprise information according to a preset time period, the method further includes: performing data cleaning processing on the enterprise information, where the data cleaning processing includes: removing data with a data missing percentage greater than a preset percentage in the enterprise information, and averaging the outliers in the enterprise information by using a 3σ outlier detection method.
[0011] Optionally, before performing feature extraction processing on the enterprise information matrix using a deep residual network, the method further includes: constructing the deep residual network, where the deep residual network is a network formed by stacking convolutional modules with skip connections layer by layer.
[0012] According to another aspect of the present application, there is provided an apparatus for predicting an enterprise state based on a residual network and a multi-head attention mechanism, including: a first acquisition unit configured to acquire enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information; a processing unit configured to perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and perform feature extraction processing on the enterprise information matrix using a deep residual network to obtain enterprise information features; a prediction processing unit configured to use an enterprise state prediction model based on a temporal multi-head attention mechanism to perform prediction processing on the enterprise state of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise.
[0013] According to still another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods for predicting an enterprise state based on a residual network and a multi-head attention mechanism.
[0014] According to yet another aspect of the present application, there is provided an electronic device, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods for predicting an enterprise state based on a residual network and a multi-head attention mechanism.
[0015] Applying the technical solution of the present application, enterprise information is obtained according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information; the enterprise information is segmented and transformed according to a preset time period to obtain an enterprise information matrix, and a deep residual network is used to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features; a corporate state prediction model based on a temporal multi-head attention mechanism is used to perform prediction processing on the corporate state of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise. Through the combination of the deep residual network and the multi-head attention mechanism, the accuracy and timeliness of corporate state prediction are significantly improved. The deep residual network can effectively extract complex features in enterprise information, while the multi-head attention mechanism can capture key temporal information in the changes of corporate state, enabling the prediction model to more accurately reflect the development trend of the enterprise. It not only provides strong data support for enterprise decision-making, but also provides a more accurate risk assessment tool for industries such as finance and investment, with broad application prospects and significant economic benefits, and solves the problem of low prediction accuracy in existing solutions for predicting corporate state. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0017] Figure 1 The hardware structure block diagram of a mobile terminal showing a method for predicting corporate state based on a residual network and a multi-head attention mechanism provided in an embodiment of the present application is shown;
[0018] Figure 2 The flowchart showing a method for predicting corporate state based on a residual network and a multi-head attention mechanism provided in an embodiment of the present application is shown;
[0019] Figure 3 The flowchart showing a specific method for predicting corporate state based on a residual network and a multi-head attention mechanism provided in an embodiment of the present application is shown;
[0020] Figure 4 The comparison schematic diagram of a common convolution block and a residual convolution block provided in an embodiment of the present application is shown;
[0021] Figure 5 The schematic diagram of a multi-head attention mechanism provided in an embodiment of the present application is shown;
[0022] Figure 6 The structure block diagram of a device for predicting corporate state based on a residual network and a multi-head attention mechanism provided in an embodiment of the present application is shown.
[0023] Among them, the above-mentioned drawings include the following reference numerals:
[0024] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners
[0025] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] As introduced in the background art, the existing solutions for predicting the state of an enterprise have relatively low prediction accuracy. To solve the problem of relatively low prediction accuracy in the existing solutions for predicting the state of an enterprise, the embodiments of the present application provide a method, device, computer-readable storage medium and electronic device for predicting the state of an enterprise based on a residual network and a multi-head attention mechanism.
[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention.
[0030] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of a method for predicting the state of an enterprise based on a residual network and a multi-head attention mechanism according to an embodiment of the present invention. As Figure 1As shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0031] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method for predicting enterprise status based on the residual network and the multi-head attention mechanism in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0032] In this embodiment, a method for predicting enterprise status based on a residual network and a multi-head attention mechanism running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And, although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] Figure 2It is a flowchart of a method for predicting enterprise status based on a residual network and a multi-head attention mechanism according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0034] Step S201, obtain enterprise information according to the unified social credit code of the enterprise, where the above-mentioned enterprise information includes business information, social public information, and social association information;
[0035] Specifically, according to the unified social credit code of the enterprise, obtain the business information, social public information, and social association information of the enterprise. The business information over the years includes the number of land transfers and assignments, and the amount of pledges and mortgages; the social public information includes the number of transfers and the number of patents; the social association information includes the number of administrative supervision times and the number of administrative penalty times; the government affairs big data includes the amount of tax paid, the number of social security participants, the number of housing management records, etc. And classify each feature using numbers.
[0036] Step S202, perform segmentation and transformation processing on the above-mentioned enterprise information according to a preset time period to obtain an enterprise information matrix, and use a deep residual network to perform feature extraction processing on the above-mentioned enterprise information matrix to obtain enterprise information features;
[0037] Specifically, the preset time period can be set to one quarter, and these data are segmented by quarter. The final result is that the data for each quarter form a matrix, and the data of the enterprise over the years will form a three-dimensional matrix, where the length is the number of features, the width is 3 (3 for quarters), and the height is the number of such matrices that the enterprise can be divided into by quarter. This matrix is used as the overall input of the training data. As for the labels of the training data, they need to be manually labeled at the beginning.
[0038] Step S203, use an enterprise status prediction model based on a time-series multi-head attention mechanism to predict the enterprise status of the above-mentioned enterprise according to the above-mentioned enterprise information features to obtain a prediction result of the above-mentioned enterprise, where the above-mentioned prediction result includes the development trend of the above-mentioned enterprise.
[0039] Specifically, the multi-head temporal attention mechanism is used to associate the information of key data in time, enabling the model to consider the development trend of the enterprise. In the temporal data prediction task, there is also an association relationship between each time step to be predicted and one or more time steps in the historical data. This is also the starting point of temporal data prediction, to grasp the change law of the data from the historical data, so as to map the predicted data according to the historical data. The temporal multi-head attention mechanism is used to process the extracted features to improve the model's ability to capture trends in long-term prediction and make the model more interpretable. In this way, through the correlation matrix in the temporal multi-head attention mechanism, the relationship of the enterprise to be predicted in the time dimension can be grasped, avoiding the failure of the prediction model due to too long historical data or too long time steps to be predicted.
[0040] Through this embodiment, enterprise information is obtained according to the unified social credit code of the enterprise. Among them, the above-mentioned enterprise information includes business information, social public information, and social association information; the above-mentioned enterprise information is segmented and transformed according to a preset time period to obtain an enterprise information matrix, and a deep residual network is used to perform feature extraction processing on the above-mentioned enterprise information matrix to obtain enterprise information features; an enterprise state prediction model based on a temporal multi-head attention mechanism is used to perform prediction processing on the enterprise state of the above-mentioned enterprise according to the above-mentioned enterprise information features to obtain the prediction result of the above-mentioned enterprise. Among them, the above-mentioned prediction result includes the development trend of the above-mentioned enterprise. Through the combination of the deep residual network and the multi-head attention mechanism, the accuracy and timeliness of enterprise state prediction are significantly improved. The deep residual network can effectively extract complex features in enterprise information, while the multi-head attention mechanism can capture key temporal information in the change of enterprise state, enabling the prediction model to more accurately reflect the development trend of the enterprise. It not only provides strong data support for enterprise decision-making, but also provides a more accurate risk assessment tool for industries such as finance and investment, with broad application prospects and significant economic benefits, solving the problem of low prediction accuracy in existing solutions for predicting enterprise state.
[0041] In the specific implementation process, an enterprise state prediction model based on a temporal multi-head attention mechanism is used to perform prediction processing on the enterprise state of the above-mentioned enterprise according to the above-mentioned enterprise information features to obtain the prediction result of the above-mentioned enterprise, including: using the above-mentioned temporal multi-head attention mechanism to perform matrix segmentation processing on the above-mentioned enterprise information features to obtain multiple segmented feature matrices, where the above-mentioned segmented feature matrices include a query matrix, a key matrix, and a value matrix; calculating attention scores according to the multiple above-mentioned segmented feature matrices to obtain target attention scores, and using the above-mentioned enterprise state prediction model to perform prediction processing on the enterprise state of the above-mentioned enterprise according to the above-mentioned target attention scores to obtain the above-mentioned prediction result of the above-mentioned enterprise.
[0042] Through the multi - head attention mechanism, the model can analyze enterprise information from different perspectives, improving the comprehensiveness and accuracy of predictions. It has important application value especially in fields such as financial analysis and risk control.
[0043] More specifically, calculating attention scores based on multiple above - mentioned segmentation feature matrices includes: using the attention calculation formula: Calculating the above - mentioned attention score Attention(Q, K, V), where Q is the query matrix vector, K is the key matrix vector, V is the value matrix vector, and d k is the dimension. This calculation method ensures that the model can focus on the most relevant information. For scenarios dealing with a large amount of enterprise data, such as big data analysis platforms, it can significantly improve processing efficiency and prediction accuracy.
[0044] Furthermore, after predicting the enterprise status of the above - mentioned enterprise according to the above - mentioned enterprise information characteristics and obtaining the prediction result of the above - mentioned enterprise, the above - mentioned method further includes: obtaining the actual enterprise status of the above - mentioned enterprise corresponding to the time period of the above - mentioned prediction result; comparing the above - mentioned prediction result with the above - mentioned actual enterprise status information to generate a comparison result, and updating and iterating the above - mentioned enterprise status prediction model according to the above - mentioned comparison result. This closed - loop model optimization mechanism enables the prediction model to continuously learn and adapt to new data, maintaining the continuous improvement of prediction ability. For institutions that need to monitor enterprise status for a long time, such as credit rating companies, it can provide more stable and reliable prediction services.
[0045] In addition, this embodiment can also establish a decision - feedback mechanism, combining the model prediction result with the actual business decision to form a closed - loop. For example, adjusting the parameters of the enterprise credit assessment model according to the prediction result, or optimizing the business process to improve the timeliness and effectiveness of risk warning. Specifically, it includes the following steps:
[0046] Step 1: Model output explanation When predicting the enterprise status, this model will output a series of prediction values and potential risk assessments. These outputs are based on the analysis of enterprise historical data and characteristics, and can predict the operating status, financial health, or potential risk level of the enterprise in the future for a period of time. For example, predicting whether an enterprise has a default risk, whether the operating status is stable, or whether it is in a leading position in the industry, etc.
[0047] Step 2: Decision - making analysis Combining the prediction result, decision - makers in enterprises or financial institutions can analyze and adjust the current business decision. For example, if the prediction shows that a certain enterprise has a high default risk in the next few months, measures can be taken in advance, such as increasing credit guarantees, adjusting loan conditions, or recalling loans in advance, to avoid potential financial risks.
[0048] Step 3: Result Application The predicted results and risk assessments can be directly applied to business processes. In aspects such as enterprise cooperation, loan approval, and investment decisions, based on the predicted results of the model, more cautious strategies can be formulated, such as increasing loan interest rates, requiring more collateral, and increasing the review frequency, etc., to ensure the safety of funds and the soundness of operations.
[0049] Step 4: Business Process Optimization Based on the predicted results of the model, the business process can be optimized. For example, for the identified high-risk enterprises, the internal credit assessment process can be optimized, introducing more stringent review criteria or adding additional review steps to reduce potential losses. At the same time, for low-risk enterprises, the process can be simplified to improve the efficiency of business processing.
[0050] Step 5: Model Iteration and Optimization An important link in the decision feedback mechanism is to feed back the actual business results (such as the actual default situation of enterprises, changes in operating status, etc.) to the prediction model for model iteration and optimization. By continuously collecting and analyzing the actual business results, the parameters of the model can be adjusted continuously, optimizing feature selection and weight allocation, so that the model can more accurately predict the enterprise status. This mechanism ensures the continuous learning and evolution of the model and can adapt to the ever-changing market environment and enterprise behavior patterns.
[0051] Such a decision feedback mechanism can also prompt enterprises or financial institutions to adjust their policies and strategies. For example, if the model prediction indicates that certain industries or business areas are of high risk, the company may adjust its business expansion strategy and pay more attention to risk control; on the contrary, if certain areas are predicted to be of low risk, the company may increase investment or credit support.
[0052] Furthermore, before predicting the enterprise status of the above-mentioned enterprise according to the above-mentioned enterprise information characteristics by using an enterprise status prediction model based on a temporal multi-head attention mechanism, the above-mentioned method further includes: obtaining external environment data and enterprise characteristic data, where the above-mentioned external environment data includes industry dynamic data and economic indicator data, and the above-mentioned enterprise characteristic data includes the enterprise operation cycle and the enterprise industry characteristics; combining the above-mentioned external environment data and the above-mentioned enterprise characteristic data to optimize the prediction interval of the above-mentioned enterprise status prediction model, and obtaining the optimized above-mentioned enterprise status prediction model.
[0053] The external environmental data of this method includes industry dynamics, economic indicators, etc. By constructing a multi-source data framework that includes internal and external information, the development environment of an enterprise can be evaluated more comprehensively, thereby improving the accuracy of prediction. The prediction interval is optimized based on the operating cycle of the enterprise, industry characteristics, etc. For example, for enterprises with obvious seasonality, the window size of the attention mechanism can be adjusted to make it more focused on the time periods related to seasonal changes, thereby improving the pertinence and accuracy of prediction. Considering both the external environment and enterprise characteristics, the prediction model can more accurately predict the development status of an enterprise under specific conditions, which can provide a more accurate basis for investment decisions for investors.
[0054] Specifically, before segmenting and transforming the above enterprise information according to a preset time period, the above method further includes: performing data cleaning on the above enterprise information, where the data cleaning includes: removing data with data loss greater than a preset percentage in the above enterprise information, and averaging the outliers in the above enterprise information using the 3σ outlier detection method. Among them, the preset percentage can be set to 50%.
[0055] There are mainly two types of data cleaning in this method. One is to directly remove features with data loss exceeding 50%, and the other is to average the outliers using the 3σ outlier detection method. Data cleaning is a key step to ensure the quality of the input data of the prediction model, which can avoid prediction biases caused by data quality problems and can provide more reliable data support for enterprise management that relies on high-quality data for decision-making.
[0056] Furthermore, before performing feature extraction on the above enterprise information matrix using a deep residual network, the above method further includes: constructing the above deep residual network, where the above deep residual network is a network formed by stacking convolutional modules with skip connections layer by layer.
[0057] The construction method of the deep residual network in this method can effectively solve the problem of gradient disappearance in the training of deep networks, improve the training efficiency and prediction performance of the model, and can provide more powerful analysis capabilities for scenarios that need to process complex enterprise information, such as enterprise intelligent analysis systems.
[0058] In order to enable those skilled in the art to understand the technical solution of the present application more clearly, the implementation process of the method for predicting enterprise status based on a residual network and a multi-head attention mechanism of the present application will be described in detail below with specific embodiments.
[0059] This embodiment relates to a specific method for predicting enterprise status based on a residual network and a multi-head attention mechanism. First, data cleaning technology is used to remove missing items and unreasonable items in enterprise data. Then, each item of data is separated by quarter to form a time series. Then, the processed data is extracted with features through a residual convolutional neural network technology, which can effectively solve the problem of only considering several indicators and the inability to associate between data. Then, the features are input into the multi-head attention mechanism, and through the attention matrix, point-level modeling is performed on all time steps, thereby enhancing the model's learning ability for multi-time-step prediction tasks. The temporal attention mechanism can effectively improve the current patent's lack of attention to the enterprise's development trend over time to enhance the effectiveness of trend prediction, as Figure 3 shown, and specifically includes the following contents:
[0060] Step S1: Data preparation. According to the unified social credit code of the enterprise, obtain the enterprise's business information, social public information, and social association information. The business information over the years is shown in Table 1, including the number of land transfers and assignments, and the amount of pledges and mortgages; the social public information includes the number of transfers and the number of patents; the social association information includes the number of administrative supervision times and the number of administrative penalty times; the government affairs big data includes the amount of tax paid, the number of social security participants, the number of housing management records, etc. And each feature is classified using numbers. For each feature, we perform data cleaning. There are mainly two types of data cleaning. One is to directly remove features with more than 50% data missing, and the other is to use the 3σ outlier detection method to average the outliers. Next, we divide these data by quarter. The final result is that the data for each quarter forms a matrix, and the data of the enterprise over the years will form a three-dimensional matrix, where the length is the number of features, the width is the set number of divided data (quarter is 3), and the height is the number of such matrices that the enterprise can be divided into by quarter. This matrix is used as the overall input of the training data, and the input data is shown in Table 1. As for the labels of the training data, they need to be manually labeled at the beginning. There are two types of labels, classified as high-quality enterprises set to 0 and risky enterprises set to 1.
[0061] Table 1
[0062]
[0063] Table 1 is an example table of input data. Each row represents a month, and the data is for a quarter. The number of patents is divided into intervals, 0-100 is 0, and 100-500 is 1. The input data is such a matrix.
[0064] Step S2: Feature extraction by the residual network. The data processed in Step S1 is input into a deep residual network (ResNet) for feature extraction. ResNet is a network formed by stacking convolutional modules with skip connections (usually called ResBlocks) layer by layer. So first, ResBlock is introduced. Figure 4 It is a schematic diagram of the structures of two ordinary convolutional layers (the left is an ordinary convolutional block, and the right is a residual convolutional block. For the convenience of description, it is called "ConvBlock") and ResBlock. Figure 4 The ConvBlock in the left figure contains two convolutional layers and a relu activation function between them. We assume that the last red relu does not belong to ConvBlock, and whether it is added is determined by the subsequent network structure. We can represent Figure 4 the ConvBlock in it with the function F(x) (excluding the last red relu). On the basis of ConvBlock, ResBlock adds a skip connection, which can effectively avoid the problem of performance degradation. This skip connection directly sends the input X to the original output F(x), as shown in Figure 4 the right figure of the residual convolutional block in it, and adds them together. So ResBlock is represented by the function F(x)+x. By repeatedly stacking ResBlocks, we can build a ResNet with a specified depth. For the original convolutional neural network, as the depth increases, the extracted features become more abstract, and even the gradient may disappear, and the initial input will also be gradually forgotten. The residual module can effectively improve this problem, fit more appropriate features, and improve the feature extraction ability of the network. Therefore, the second part is to input the data generated in the first part into the residual network for feature extraction, and finally input the extracted features into the next step for temporal association to output the final result.
[0065] Step S3: Use the multi-head temporal attention mechanism to associate the information of key data in time, so that the model can consider the trend of enterprise development. In the temporal data prediction task, there is also an association relationship between each time step to be predicted and one or more time steps in the historical data. This is also the starting point of temporal data prediction, to grasp the change law of the data from the historical data, so as to map the predicted data according to the historical data. The present invention uses the temporal multi-head attention mechanism to process the extracted features to improve the trend capture ability of the model in long-term prediction and make the model more interpretable. The multi-head attention mechanism is as shown in Figure 5As shown, the input time series data matrix has a size of L*N, where L represents the length of the data in the time dimension and N represents the number of time series variables. First, the time series matrix is mapped to a high-dimensional space through three fully connected layers with the same shape to obtain the query matrix Q, the key matrix K, and the value matrix V. Wq and Bq, Wk and Bk, Wv and Bv are the weight and bias matrices of the three fully connected layers respectively. Then, matrix splitting is performed, and the matrices are split into multiple matrices for attention calculation respectively. The calculation formula is where Q is the query matrix vector, K is the key matrix vector, V is the value matrix vector, and d k is the dimension. Finally, the calculated scores are fused to output the final result.
[0066] In the embodiment of the present application, the enterprise's historical data is processed into a matrix from the time dimension, which is convenient for operation and improves the correlation between data. The ResNet network is used for feature extraction to improve the feature extraction ability. The time series multi-head attention mechanism is used to process the extracted features to enhance the model's trend capture ability in long-term prediction and make the model more interpretable. Using this model, the development trend of the enterprise can be predicted, which has more reference value in handling long-term business. And through the correlation matrix in the time series multi-head attention mechanism, the relationship of the enterprise to be predicted in the time dimension is grasped, avoiding the failure of the prediction model due to too long historical data or too long time steps to be predicted.
[0067] The embodiment of the present application also provides a device for predicting the enterprise status based on the residual network and the multi-head attention mechanism. It should be noted that the device for predicting the enterprise status based on the residual network and the multi-head attention mechanism in the embodiment of the present application can be used to execute the method for predicting the enterprise status based on the residual network and the multi-head attention mechanism provided by the embodiment of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0068] The following introduces the device for predicting the enterprise status based on the residual network and the multi-head attention mechanism provided by the embodiment of the present application.
[0069] Figure 6 is a schematic diagram of the device for predicting the enterprise status based on the residual network and the multi-head attention mechanism according to the embodiment of the present application. As Figure 6 shown, the device includes:
[0070] The first acquisition unit 61 is configured to acquire enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information;
[0071] The processing unit 62 is configured to perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and perform feature extraction processing on the enterprise information matrix by using a deep residual network to obtain enterprise information features;
[0072] The prediction processing unit 63 is configured to use an enterprise status prediction model based on a temporal multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise.
[0073] In this embodiment, the first acquisition unit is configured to acquire enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information; the processing unit is configured to perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and perform feature extraction processing on the enterprise information matrix by using a deep residual network to obtain enterprise information features; the prediction processing unit is configured to use an enterprise status prediction model based on a temporal multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise. By combining the deep residual network and the multi-head attention mechanism, the accuracy and timeliness of enterprise status prediction are significantly improved. The deep residual network can effectively extract complex features in enterprise information, while the multi-head attention mechanism can capture key temporal information in the change of enterprise status, enabling the prediction model to more accurately reflect the development trend of the enterprise. It not only provides strong data support for enterprise decision-making, but also provides a more accurate risk assessment tool for industries such as finance and investment, with broad application prospects and significant economic benefits, and solves the problem of low prediction accuracy in existing solutions for predicting enterprise status.
[0074] As an optional solution, the prediction processing unit includes a matrix segmentation module and a prediction module. The matrix segmentation module is configured to perform matrix segmentation processing on the enterprise information features by using the temporal multi-head attention mechanism to obtain a plurality of segmented feature matrices, where the segmented feature matrices include a query matrix, a key matrix, and a value matrix; the prediction module is configured to calculate attention scores according to the plurality of segmented feature matrices to obtain a target attention score, and use the enterprise status prediction model to perform prediction processing on the enterprise status of the enterprise according to the target attention score to obtain the prediction result of the enterprise.
[0075] An alternative solution is that the prediction module includes a calculation sub-module, which is used to adopt the attention calculation formula: Calculate the above attention score Attention(Q, K, V), where Q is the query matrix vector, K is the key matrix vector, V is the value matrix vector, and d k is the dimension.
[0076] An alternative solution is that the device further includes a second acquisition unit and a comparison processing unit. The second acquisition unit is used to acquire the actual enterprise status of the enterprise corresponding to the prediction result of the enterprise in the corresponding time period after predicting the enterprise status of the enterprise according to the above enterprise information characteristics; the comparison processing unit is used to compare the prediction result with the actual enterprise status information, generate a comparison result, and update and iterate the enterprise status prediction model according to the comparison result.
[0077] An alternative solution is that the device further includes a third acquisition unit and an optimization processing unit. The third acquisition unit is used to acquire external environment data and enterprise characteristic data before predicting the enterprise status of the enterprise according to the above enterprise information characteristics by using the enterprise status prediction model based on the temporal multi-head attention mechanism, where the external environment data includes industry dynamic data and economic indicator data, and the enterprise characteristic data includes the enterprise operation cycle and the enterprise industry characteristics; the optimization processing unit is used to optimize the prediction interval of the enterprise status prediction model by combining the external environment data and the enterprise characteristic data to obtain the optimized enterprise status prediction model.
[0078] An alternative solution is that the device further includes a data cleaning unit, which is used to clean the enterprise information before performing segmentation and conversion processing on the enterprise information according to a preset time period, where the data cleaning processing includes: removing data with data loss greater than a preset percentage in the enterprise information, and averaging the outliers in the enterprise information by using the 3σ outlier detection method.
[0079] An alternative solution is that the device further includes a construction unit, which is used to construct the deep residual network before performing feature extraction processing on the enterprise information matrix by using the deep residual network, where the deep residual network is a network formed by stacking convolutional modules with skip connections layer by layer.
[0080] The above device for predicting enterprise status based on residual network and multi-head attention mechanism includes a processor and a memory. The above first acquisition unit, processing unit, prediction processing unit, etc. are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above program units stored in the memory. The above modules are all located in the same processor; alternatively, the above modules are respectively located in different processors in any combination form.
[0081] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of low prediction accuracy in the existing solutions for predicting enterprise status can be solved.
[0082] The memory may include non-permanent memory in computer-readable media, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory includes at least one storage chip.
[0083] An embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program. Wherein, when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the method for predicting enterprise status based on residual network and multi-head attention mechanism.
[0084] Specifically, the method for predicting enterprise status based on residual network and multi-head attention mechanism includes:
[0085] Step S201, obtain enterprise information according to the unified social credit code of the enterprise. Wherein, the above enterprise information includes business information, social public information, and social association information;
[0086] Step S202, perform segmentation and conversion processing on the above enterprise information according to a preset time period to obtain an enterprise information matrix, and use a deep residual network to perform feature extraction processing on the above enterprise information matrix to obtain enterprise information features;
[0087] Step S203, use an enterprise status prediction model based on a time-series multi-head attention mechanism to perform prediction processing on the enterprise status of the above enterprise according to the above enterprise information features to obtain the prediction result of the above enterprise. Wherein, the above prediction result includes the development trend of the above enterprise.
[0088] An embodiment of the present invention provides a processor. The above processor is used to run a program. Wherein, when the above program runs, it executes the method for predicting enterprise status based on residual network and multi-head attention mechanism.
[0089] Specifically, the method for predicting enterprise status based on residual network and multi-head attention mechanism includes:
[0090] Step S201, obtain enterprise information according to the unified social credit code of the enterprise, wherein the enterprise information includes business information, social public information and social association information;
[0091] Step S202, perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and use a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features;
[0092] Step S203, use an enterprise status prediction model based on a time-series multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, wherein the prediction result includes the development trend of the enterprise.
[0093] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps:
[0094] Step S201, obtain enterprise information according to the unified social credit code of the enterprise, wherein the enterprise information includes business information, social public information and social association information;
[0095] Step S202, perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and use a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features;
[0096] Step S203, use an enterprise status prediction model based on a time-series multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, wherein the prediction result includes the development trend of the enterprise.
[0097] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0098] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:
[0099] Step S201, obtain enterprise information according to the unified social credit code of the enterprise, wherein the enterprise information includes business information, social public information and social association information;
[0100] Step S202, perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and use a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features;
[0101] Step S203: Use an enterprise status prediction model based on a temporal multi-head attention mechanism to perform a prediction process on the enterprise status of the above-mentioned enterprise according to the above-mentioned enterprise information characteristics, and obtain a prediction result of the above-mentioned enterprise, where the above-mentioned prediction result includes the development trend of the above-mentioned enterprise.
[0102] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 a process or multiple processes and / or blocks Figure 1 steps of the functions specified in a block or multiple blocks.
[0107] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0108] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0109] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0110] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.
[0111] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0112] 1) A method for predicting enterprise status based on residual network and multi-head attention mechanism of the present application includes: obtaining enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information; performing segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and using a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features; using an enterprise status prediction model based on a time-series multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise. By combining the deep residual network and the multi-head attention mechanism, the accuracy and timeliness of enterprise status prediction are significantly improved. The deep residual network can effectively extract complex features in enterprise information, while the multi-head attention mechanism can capture key time-series information in the changes of enterprise status, enabling the prediction model to more accurately reflect the development trend of the enterprise. It not only provides strong data support for enterprise decision-making, but also provides a more accurate risk assessment tool for industries such as finance and investment, has broad application prospects and significant economic benefits, and solves the problem of low prediction accuracy in existing solutions for predicting enterprise status.
[0113] 2) A device for predicting enterprise status based on residual network and multi-head attention mechanism of the present application includes: a first acquisition unit for obtaining enterprise information according to the unified social credit code of the enterprise, where the enterprise information includes business information, social public information, and social association information; a processing unit for performing segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and using a deep residual network to perform feature extraction processing on the enterprise information matrix to obtain enterprise information features; a prediction processing unit for using an enterprise status prediction model based on a time-series multi-head attention mechanism to perform prediction processing on the enterprise status of the enterprise according to the enterprise information features to obtain a prediction result of the enterprise, where the prediction result includes the development trend of the enterprise. By combining the deep residual network and the multi-head attention mechanism, the accuracy and timeliness of enterprise status prediction are significantly improved. The deep residual network can effectively extract complex features in enterprise information, while the multi-head attention mechanism can capture key time-series information in the changes of enterprise status, enabling the prediction model to more accurately reflect the development trend of the enterprise. It not only provides strong data support for enterprise decision-making, but also provides a more accurate risk assessment tool for industries such as finance and investment, has broad application prospects and significant economic benefits, and solves the problem of low prediction accuracy in existing solutions for predicting enterprise status.
[0114] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for predicting enterprise status based on residual network and multi-head attention mechanism, characterized in that: include: Acquire enterprise information according to the enterprise's unified social credit code, wherein the enterprise information includes business information, social public information and social related information; Segmenting and converting the enterprise information according to a preset time period to obtain an enterprise information matrix, and using a deep residual network to extract features from the enterprise information matrix to obtain enterprise information features; An enterprise status prediction model based on a temporal multi-head attention mechanism is adopted to predict the enterprise status of the enterprise according to the enterprise information characteristics to obtain a prediction result of the enterprise, wherein the prediction result includes the development trend of the enterprise.
2. The method according to claim 1, characterized in that The enterprise status prediction model based on the time-series multi-head attention mechanism is adopted to predict the enterprise status of the enterprise according to the enterprise information characteristics, and the prediction result of the enterprise is obtained, including: The time-series multi-head attention mechanism is used to perform matrix segmentation processing on the enterprise information features to obtain multiple segmentation feature matrices, wherein the segmentation feature matrices include a query matrix, a key matrix, and a value matrix; An attention score is calculated based on the multiple segmentation feature matrices to obtain a target attention score, and the enterprise status prediction model is used to predict the enterprise status of the enterprise based on the target attention score to obtain the prediction result of the enterprise.
3. The method according to claim 2, characterized in that Calculating an attention score according to the plurality of segmentation feature matrices includes: Using the attention calculation formula: Calculate the attention score Attention(Q,K,V), where Q is the query matrix vector, K is the key matrix vector, V is the value matrix vector, d k For dimension.
4. The method according to claim 1, characterized in that: After predicting the enterprise status of the enterprise according to the enterprise information feature to obtain the prediction result of the enterprise, the method further includes: Obtaining the actual enterprise status of the enterprise in the time period corresponding to the forecast result; The prediction result and the actual enterprise status information are compared to generate a comparison result, and the enterprise status prediction model is updated and iterated according to the comparison result.
5. The method according to claim 1, characterized in that Before using the enterprise status prediction model based on the time-series multi-head attention mechanism to predict the enterprise status of the enterprise according to the enterprise information features, the method further includes: Acquire external environment data and enterprise characteristic data, wherein the external environment data includes industry dynamic data and economic indicator data, and the enterprise characteristic data includes enterprise operation cycle and enterprise industry characteristics; The prediction interval of the enterprise status prediction model is optimized in combination with the external environment data and the enterprise characteristic data to obtain the optimized enterprise status prediction model.
6. The method according to claim 1, characterized in that Before segmenting and converting the enterprise information according to the preset time period, the method further includes: The enterprise information is subjected to data cleaning processing, wherein the data cleaning processing includes: removing data in the enterprise information with data missing greater than a preset percentage, and averaging outliers in the enterprise information using a 3σ outlier detection method.
7. The method according to claim 1, characterized in that Before using the deep residual network to perform feature extraction processing on the enterprise information matrix, the method further includes: The deep residual network is constructed, wherein the deep residual network is a network formed by stacking layers of convolutional modules with skip connections.
8. A device for predicting enterprise status based on residual network and multi-head attention mechanism, characterized in that: include: A first acquisition unit is used to acquire enterprise information according to the unified social credit code of the enterprise, wherein the enterprise information includes business information, social public information and social related information; A processing unit, configured to perform segmentation and conversion processing on the enterprise information according to a preset time period to obtain an enterprise information matrix, and perform feature extraction processing on the enterprise information matrix using a deep residual network to obtain enterprise information features; A prediction processing unit is used to adopt an enterprise status prediction model based on a temporal multi-head attention mechanism to predict the enterprise status of the enterprise according to the enterprise information characteristics to obtain a prediction result of the enterprise, wherein the prediction result includes the development trend of the enterprise.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the method for predicting enterprise status based on residual network and multi-head attention mechanism as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of claims 1 to 7 for predicting enterprise status based on a residual network and a multi-head attention mechanism.