Enterprise multi-dimensional data feature fusion method combined with AI technology

By combining AI technology with KS test and risk deviation consistency analysis, the problems of data redundancy and feature selection deviation in enterprise multi-dimensional data fusion are solved, and more accurate enterprise risk assessment is achieved.

CN120724400BActive Publication Date: 2025-12-30SHANSHOUFU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511248507.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-12-30
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing technologies suffer from data redundancy and feature selection deviations from normal characteristics in multi-dimensional data fusion for enterprises, leading to inaccurate risk assessments, especially since the differences in data attributes between enterprises in different industries are not fully considered.

Method used

Using AI technology, the statistical significance of attributes and risk values ​​is determined by the KS test method. Combining the consistency and importance of risk bias, attribute data are iteratively selected and weighted to form a set of feasible attribute data.

Benefits of technology

It improves the representativeness and accuracy of data features, reduces data redundancy, and enhances the effectiveness and accuracy of enterprise risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724400B_ABST
    Figure CN120724400B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data fusion, in particular to a multi-dimensional data feature fusion method for enterprises combined with AI technology. The method takes the attribute data of any dimension as the attribute to be analyzed for the enterprises in the same industry, determines the statistical significance of the attribute to be analyzed and the risk value through a K-S test method, compares the attribute to be analyzed and the risk value to determine the risk deviation consistency of the attribute to be analyzed, combines the statistical significance and the risk deviation consistency to determine the importance of risk assessment, adds the attribute data to a candidate set according to the size order of the importance, determines the recognition degree of the attribute data added to the candidate set, selects the candidate set corresponding to the maximum recognition degree as the final candidate set, and takes the attribute data in the final candidate set as the feasibility attribute data to perform weighted fusion on the feasibility attribute data according to the recognition degree. The application only performs feature fusion on the selected attribute, avoids data redundancy, and improves the data fusion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data fusion technology, specifically to a method for fusing multi-dimensional data features of enterprises by combining AI technology. Background Technology

[0002] Enterprise data includes financial statements, supply chain data, customer relationship management data, as well as external industry reports and public opinion reports. A single data source can hardly fully reflect the true situation of an enterprise. Therefore, obtaining multidimensional data about an enterprise and using data fusion to identify enterprise risks can reduce the bias of human experience.

[0003] However, the fusion process involves a large amount of enterprise data unrelated to risk. This irrelevant data can interfere with the feature fusion results, increasing the difficulty of multi-dimensional data analysis. Furthermore, due to potential risk correlations between multiple data points, data redundancy arises. This redundancy also increases the difficulty of analyzing multi-dimensional enterprise data, leading to fused data features that fail to accurately reflect enterprise risk. Therefore, selecting appropriate multi-dimensional data features is a crucial step in achieving feature fusion.

[0004] Currently, the selection of key data for enterprises typically relies on experience and judgment within their respective industries. For example, the service industry focuses more on user and service data, while the manufacturing industry focuses more on data related to cash flow and debt collection. Therefore, different types of enterprises focus on different data attributes, and the data used to assess enterprise risk also differs. Traditional methods mainly use clustering or association analysis to uniformly select data attributes for all enterprises, without considering the specific characteristics of the industry or the enterprise itself. For instance, industries such as telecommunications, electronics, and technology have more patent disputes and significant cross-licensing agreements compared to other industries. These unique characteristics cause the selection of enterprise data attributes to deviate from normal feature selection, and these specific attributes are more likely to reflect the potential risks of the enterprise. Summary of the Invention

[0005] To address the technical problem that unifying the selection of enterprise data attributes through clustering or association analysis leads to deviations from normal feature selection, this invention aims to provide a method for fusing multi-dimensional enterprise data features using AI technology. The specific technical solution adopted is as follows:

[0006] In a first aspect, embodiments of the present invention provide a method for fusing multi-dimensional data features of an enterprise using AI technology, the method comprising:

[0007] Obtain multidimensional attribute data of enterprises;

[0008] Companies are categorized according to their industry and risk values ​​are labeled for different attribute data.

[0009] For companies in the same industry, using attribute data of any dimension as the attribute to be analyzed, the statistical significance of the attribute to be analyzed and the risk value is determined by the KS test method; the attribute to be analyzed and the risk value are compared to determine the consistency of the risk deviation of the attribute to be analyzed; and the importance of the risk assessment of the attribute to be analyzed is determined by combining the statistical significance and the consistency of the risk deviation.

[0010] Attribute data is iteratively selected according to importance, and added to the candidate set. The distinctiveness of the attribute data added to the candidate set is determined. The candidate set with the highest distinctiveness is selected as the final candidate set. The attribute data in the final candidate set is used as the industry's feasible attribute data, and the feasible attribute data is weighted and fused according to distinctiveness.

[0011] Secondly, a multi-dimensional data feature fusion system for enterprises that combines AI technology is provided, the system comprising the following modules:

[0012] The enterprise data acquisition module is used to acquire multidimensional attribute data of an enterprise;

[0013] The risk labeling module is used to classify companies according to their industry and label risk values ​​for different attribute data.

[0014] The risk assessment module is used to analyze enterprises in the same industry, using attribute data of any dimension as the attribute to be analyzed. The KS test method is used to determine the statistical significance of the attribute to be analyzed and the risk value; the attribute to be analyzed and the risk value are compared to determine the consistency of the risk deviation of the attribute to be analyzed; and the importance of the risk assessment of the attribute to be analyzed is determined by combining the statistical significance and the consistency of the risk deviation.

[0015] The data fusion module iteratively selects attribute data in order of importance, adds the attribute data to the candidate set, determines the identifiability of the attribute data added to the candidate set, selects the candidate set with the highest identifiability as the final candidate set, and uses the attribute data in the final candidate set as the industry's feasible attribute data, and performs weighted fusion of the feasible attribute data based on identifiability.

[0016] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the various possible implementations of the first aspect.

[0017] Fourthly, embodiments of the present invention provide a computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform the method described in the first aspect or any possible implementation thereof.

[0018] Fifthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the various possible implementations of the first aspect.

[0019] The embodiments of the present invention have at least the following beneficial effects:

[0020] Because enterprise data has high dimensionality, not all attribute data can be used for enterprise risk assessment. Therefore, the importance of an attribute is determined by analyzing the consistency between the distribution of each attribute data and the risk distribution, and the correspondence between each enterprise's risk value and attribute data. The combination of multi-dimensional attributes is considered to assess the performance of these combinations on enterprise risk. The distinctiveness of the selected attribute or the current candidate set is judged by analyzing each change in the set. This approach avoids analyzing all combinations of multi-dimensional data to select enterprise risk attributes, as the high dimensionality of enterprise data leads to an exponential increase in combinations, making it difficult to judge the importance of attribute combinations and resulting in incomplete attribute assessments. This approach achieves attribute selection from multi-dimensional enterprise data. Feature fusion is performed only on the selected feasible attribute data to increase the value of the data while avoiding data redundancy. This makes the data features constructed from the selected attributes more representative. Attached Figure Description

[0021] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a method for fusing multi-dimensional data features of an enterprise using AI technology, as provided in one embodiment of the present invention;

[0023] Figure 2 This is a system block diagram of an enterprise multi-dimensional data feature fusion system combining AI technology, provided as an embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram of the structure of a computer device provided in one embodiment of the present invention. Detailed Implementation

[0025] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a multi-dimensional data feature fusion method for enterprises that combines AI technology, based on the present invention.

[0026] In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form.

[0027] In the description of the embodiments of the present invention, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present invention, "multiple" means two or more.

[0028] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0030] The embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0031] The following description, in conjunction with the accompanying drawings, details a specific scheme for a multi-dimensional data feature fusion method for enterprises that combines AI technology, provided by this invention.

[0032] Please see Figure 1 The diagram illustrates a flowchart of a multi-dimensional data feature fusion method for enterprises that combines AI technology, according to an embodiment of the present invention. The method includes the following steps:

[0033] Step S100: Obtain the enterprise's multidimensional attribute data.

[0034] Obtain relevant attribute data of enterprises from public channels, including:

[0035] Obtain data from the business registration information platform regarding frequent changes in registered capital and administrative penalty records;

[0036] Data such as taxpayer credit rating, tax arrears announcements, and tax audit results can be obtained from the tax information disclosure platform.

[0037] Obtain data on the company's debt situation and cash flow from the company's publicly disclosed financial reports;

[0038] Obtain relevant data such as the number of patents, the number of patents granted, and patent litigation from patent and intellectual property disputes;

[0039] Obtain relevant data such as user reviews and public opinion from social media platforms or enterprise feedback platforms;

[0040] Obtain multi-dimensional attribute data of enterprises through multiple public platforms.

[0041] Furthermore, the attribute data collected by enterprises is preprocessed to quantify the data. For data that cannot be quantified, such as user reviews, a risk discrimination model can be constructed based on the data characteristics of that attribute data. In a preferred embodiment of this invention, a CNN network can be used. The attribute data that cannot be quantified is input into the network model, which outputs the risk discrimination result of that attribute data. This risk discrimination result is used as the data value of the attribute data. The risk discrimination result is a value between 0 and 1. Since it needs to be distinguished from enterprise risk discrimination, only single-dimensional attribute data is input here, rather than the enterprise's multi-dimensional attribute data. Examples include capital change frequency, administrative trigger count, and severity. After quantifying the attribute data, the z-score method is used to standardize the values ​​of all attribute data to eliminate the problem of bias in the data feature analysis process caused by different units and data ranges of attribute data from different dimensions. The standardized data is used as the data value of the enterprise's attribute data. The enterprise's multi-dimensional attribute data consists of multiple attribute data, with each dimension corresponding to one attribute data.

[0042] Step S200: Classify enterprises according to their industry and label risk values ​​for different attribute data.

[0043] In this embodiment of the invention, enterprises are categorized according to their industry, such as internet, finance, energy, and information technology. In other embodiments, enterprises may be categorized in other ways, such as by enterprise size or revenue. Different enterprise categorization methods will affect comparisons between enterprises, thus influencing the final risk assessment results. In this embodiment of the invention, subsequent steps are analyzed based on the enterprise's industry.

[0044] First, risk assessment and analysis are conducted on the attribute data of different enterprises, and risk values ​​for different attribute data are obtained through manual annotation. Specifically, the risk assessment reports of the enterprises are analyzed to achieve the desired risk assessment results.

[0045] Step S300: For companies in the same industry, use attribute data of any dimension as the attribute to be analyzed, and determine the statistical significance of the attribute to be analyzed and the risk value using the KS test method; compare the attribute to be analyzed and the risk value to determine the consistency of the risk deviation of the attribute to be analyzed; combine the statistical significance and the consistency of the risk deviation to determine the importance of the risk assessment of the attribute to be analyzed.

[0046] To avoid discrepancies in the data attributes that different industries focus on—for example, the service industry needs to focus more on customer and user data, while manufacturing suppliers need to focus more on financial and debt data—we categorize companies within the same industry for easier subsequent analysis. All analyses in the following steps will focus on companies within the same industry.

[0047] In this embodiment of the invention, attribute data of any dimension is set as the attribute to be analyzed.

[0048] Taking the attribute to be analyzed as an example, obtain the data value corresponding to the attribute to be analyzed. If the attribute to be analyzed has a high correlation with the enterprise's risk, it means that the attribute to be analyzed can be used to better assess the enterprise's risk; conversely, if the data value corresponding to the attribute to be analyzed has no correlation with the enterprise's risk, it means that the attribute to be analyzed is poor at assessing the enterprise's risk and cannot be used to better assess the enterprise's risk, and is considered an interfering attribute.

[0049] First, the statistical significance of the attribute to be analyzed and the risk value is determined using the KS test:

[0050] First, two datasets are established: a dataset of attributes to be analyzed for all companies in the same industry, and a dataset of risk values ​​for each company. The p-values ​​of the attributes to be analyzed and the risk values ​​are determined using the KS test, which serves as statistical significance. Statistical significance indicates whether the attributes to be analyzed and the risk values ​​of companies originate from the same distribution function, and is used to determine the similarity between the distribution of the attribute and the distribution of the company's risk values. The p-value reflects the probability distribution between the attributes to be analyzed and the risk values ​​of companies; the smaller the p-value, the weaker the correlation, and vice versa.

[0051] It's important to note that the above method, which splits data from all companies within the same industry into two datasets to validate data distribution, neglects the correlation between the analyzed attribute data characteristics and risk values. Since these belong to the same industry, analyzing a single company would overlook risk changes across the entire industry. Therefore, calculating the difference between the risk value of the same industry and the data value of the attribute being analyzed reflects the bias in judging the risk value of the company based on the attribute being analyzed, thus obtaining the risk assessment value of the attribute being analyzed. If these risk deviations remain largely consistent, or if the company's risk deviations are strongly correlated with its risk value, then this attribute characteristic can be used for risk assessment of the company.

[0052] By comparing the attribute to be analyzed with the risk value, the consistency of the risk deviation of the attribute to be analyzed is determined. Specifically:

[0053] Step 1: Determine the risk assessment value of the attribute to be analyzed based on the difference between the data value and the risk value. Specifically: calculate the difference between the data value and the risk value of the attribute to be analyzed, and use it as the risk assessment value of the attribute to be analyzed.

[0054] In this embodiment of the invention, for the i-th enterprise, taking the k-th attribute data as the attribute to be analyzed as an example, the risk assessment value... The calculation formula is: ;in, Let i be the risk value of the i-th enterprise; Let be the data value of the k-th attribute feature of the i-th enterprise.

[0055] Step 2: Based on the dispersion of the risk assessment values ​​of the attribute to be analyzed among all enterprises in the same industry, determine the dispersion assessment value of the attribute to be analyzed. Specifically: calculate the standard deviation of the risk assessment values ​​of the attribute to be analyzed among all enterprises in the same industry, and use it as the dispersion assessment value of the attribute to be analyzed; perform negative correlation normalization mapping on the dispersion assessment value to obtain the consistency of risk deviation of the attribute to be analyzed.

[0056] In this embodiment of the invention, the risk deviation consistency of the attribute to be analyzed is... The calculation formula is: Where norm is the normalization function; The risk assessment value of the attribute to be analyzed; Risk assessment values ​​for the attributes to be analyzed for all enterprises.

[0057] Differences in data distribution can also affect the correlation between enterprise risk deviation and risk value. If the data values ​​of attribute data and risk values ​​come from datasets with different data distributions, then a low correlation is normal and does not reflect the true risk assessment.

[0058] Therefore, further analysis is conducted to determine the data correlation of the attribute to be analyzed after excluding data distribution interference. Specifically, this involves: determining the rank correlation coefficient between the risk assessment value and the data value of the attribute to be analyzed; and calculating the normalized value of the difference between the rank correlation coefficient of the attribute to be analyzed and the statistical significance, which serves as the data correlation of the attribute to be analyzed after excluding data distribution interference. It should be noted that obtaining the rank correlation coefficient is a well-known technique among those skilled in the art and will not be elaborated upon here. It should also be noted that the rank correlation coefficient between the risk assessment value and the data value of the attribute to be analyzed is calculated using the risk assessment values ​​and data values ​​of the attribute to be analyzed from all enterprises in the same industry at the same sampling time.

[0059] In some embodiments, the correlation of data of the attribute to be analyzed after excluding data distribution interference. The calculation formula is: ;in, The rank correlation coefficient between the risk assessment value and the risk value; For statistical significance.

[0060] If the interference of data distribution is not excluded, the rank correlation coefficient may be too small due to certain deviations in data distribution, leading to misjudgment of the correlation between the actual enterprise risk deviation and the risk value.

[0061] When analyzing the risk deviation of attribute data, strong consistency or a strong correlation with risk values ​​indicates that the risk deviation of the attribute data has a strong regularity with the enterprise. Therefore, the maximum value between risk deviation consistency and data correlation is selected as the regularity characteristic reflecting the attribute data. Specifically, the maximum value between risk deviation consistency and data correlation is used as an important judgment value. If the attribute data has a large regularity characteristic and its distribution is the same as the risk value data distribution, then the risk assessment importance of the attribute characteristic is higher.

[0062] The importance of the risk assessment of the attribute to be analyzed is determined by combining data relevance, risk deviation consistency, and statistical significance. Specifically: the risk deviation consistency and data relevance are compared, and the maximum value between them is taken as the importance judgment value; the importance of the risk assessment of the attribute to be analyzed is obtained by combining the importance judgment value and the statistical significance, wherein both the importance judgment value and the statistical significance are positively correlated with the importance of the risk assessment. More specifically: the product of the importance judgment value and the statistical significance is calculated as the importance of the risk assessment of the attribute to be analyzed.

[0063] In this embodiment of the invention, the formula for calculating the importance of the risk assessment of the attribute to be analyzed is as follows: Where max is the function for finding the maximum value; This is an important judgment value.

[0064] The above analysis was performed on each attribute data point to determine the importance of risk assessment for each attribute. This avoids the situation where significant anomalies in the attribute values ​​are considered normal within the industry.

[0065] Traditional attribute importance assessments, such as those using clustering to verify consistency between clustering results and data labels, often overlook data distribution characteristics, such as whether the distribution is concentrated or divergent. For example, clustering can still proceed even when the distribution is too concentrated, but because the data is concentrated, data of the same type will be forcibly divided into multiple categories, leading to inaccurate analysis. By judging the importance of each attribute, the problem of low accuracy caused by simply separating data based on data analysis is avoided.

[0066] Step S400: Iteratively select attribute data according to importance, add the attribute data to the candidate set, and determine the identifiability of the attribute data added to the candidate set; select the candidate set with the highest identifiability as the final candidate set; use the attribute data in the final candidate set as the industry's feasible attribute data, and perform weighted fusion of the feasible attribute data according to the identifiability.

[0067] Because analyzing single-attribute data may have limitations, there may be situations where multiple dimensions of data from different companies reflect the company's risk, making the connection between single-attribute data and risk values ​​weaker, and thus the importance of single-attribute data to company risk less apparent.

[0068] On the other hand, different attribute data may have a strong correlation with risk values. Different attributes may also have correlations and changes, jointly reflecting corporate risk. This means that, from the perspective of individual attribute data, each attribute is highly important to the risk value. However, choosing to use these attributes as a whole to reflect corporate risk does not significantly improve the situation. As a result, there is redundancy among the attributes, which leads to the selected multi-dimensional data of the company not being representative, thus affecting the data fusion results.

[0069] For example, when evaluating the fast-moving consumer goods (FMCG) industry, inventory turnover is often considered a sign of good business performance. However, cash flow must also be considered. Otherwise, even if the turnover rate is good, a lack of cash flow can still lead to high operational risks for the company.

[0070] After obtaining the risk assessment importance of each attribute data in step S300, attribute data is iteratively selected and added to the candidate set according to its importance. It should be noted that the set of attribute data from companies in the same industry that are not added to the candidate set is called the remaining set. That is, the initial state of the remaining set is the entire set of all attribute data, while the initial state of the candidate set is an empty set. The execution steps of iteratively selecting attribute data and adding it to the candidate set according to its importance continuously update both the candidate set and the remaining set.

[0071] When adding an attribute from the remaining set to the candidate set, calculate the importance of the risk assessment before and after adding the attribute. If the set is empty, the importance is 0. If the set contains only one element, the importance of that attribute is used as the importance of the set. It should be noted that the importance of the risk assessment before and after adding the attribute is the average importance of the attribute data within the set.

[0072] Calculate the difference in importance of the corresponding risk assessments before and after adding attribute data to the candidate set, and use this as the rate of change before and after adjusting the candidate set.

[0073] Rate of change of the candidate set before and after adjustment The calculation formula is: ;in, The importance of risk assessment before adding attribute data to the candidate set; The importance of risk assessment after adding attribute data to the candidate set.

[0074] If adding this attribute to the candidate set results in a significant improvement in the set's importance, it indicates that the combination of this attribute and the candidate set reflects the enterprise's risk. Conversely, if the improvement in the set's importance is limited after adjustment, it suggests that the attribute is redundant, and related attributes or attribute sets reflecting this attribute's characteristics already exist in the candidate set. This is how the candidate set is determined. The distinguishability, more specifically: The distinguishability of adding attribute data to the candidate set is determined by combining the rate of change and the data values ​​of the attribute data added to the candidate set. In this embodiment of the invention, the change amount and the real-time data values ​​of the attribute data added to the candidate set are calculated as the distinguishability of adding attribute data to the candidate set.

[0075] In some embodiments, attribute data is added to the candidate set for distinctiveness. The calculation formula is: ;in, The rate of change of the candidate set before and after adjustment; This represents the real-time data values ​​of the attribute data currently added to the candidate set. This distinctiveness also reflects the enterprise attribute characteristics added to the candidate set.

[0076] Attribute data is iteratively selected and added to the candidate set in order of importance until the importance of the attribute in the risk assessment reaches a preset threshold. At this point, adding attribute data to the candidate set stops. The identifiability of each attribute data addition is determined, and the candidate set with the highest identifiability is selected as the final candidate set. In this embodiment, the preset threshold is 0.8; in other embodiments, this value can be adjusted by the implementer according to actual circumstances. If the importance is less than 0.8, it indicates that the currently selected attribute set is insufficient to represent the enterprise risk. Attribute data is then repeatedly selected from the remaining set and added to the candidate set. Finally, after determining the final candidate set, the candidate set with the highest identifiability is selected as the final candidate set.

[0077] The attribute data in the final candidate set is used as the data attributes for analyzing enterprise risk in that industry. More specifically, the attribute data in the final candidate set is recorded as feasibility attribute data. This analysis is performed for each industry, thus obtaining the enterprise data attributes that should be selected for analyzing enterprise risk in each industry, which are then used as feasibility attribute data. This completes the selection of enterprise attribute data for different industries.

[0078] Finally, for each enterprise, the feasibility attribute data is weighted and fused based on the recognizability. Specifically, the recognizability of each feasibility attribute data added to the candidate set is used as the weight of the corresponding feasibility attribute data, and the multidimensional feasibility attribute data corresponding to the enterprise is weighted and fused.

[0079] Please see Figure 2 , Figure 2 This invention provides a system block diagram of an enterprise multi-dimensional data feature fusion system that combines AI technology. The system includes the following modules:

[0080] The enterprise data acquisition module is used to acquire multidimensional attribute data of an enterprise;

[0081] The risk labeling module is used to classify companies according to their industry and label risk values ​​for different attribute data.

[0082] The risk assessment module is used to analyze enterprises in the same industry, using attribute data of any dimension as the attribute to be analyzed. The KS test method is used to determine the statistical significance of the attribute to be analyzed and the risk value; the attribute to be analyzed and the risk value are compared to determine the consistency of the risk deviation of the attribute to be analyzed; and the importance of the risk assessment of the attribute to be analyzed is determined by combining the statistical significance and the consistency of the risk deviation.

[0083] The data fusion module iteratively selects attribute data in order of importance, adds the attribute data to the candidate set, determines the identifiability of the attribute data added to the candidate set, selects the candidate set with the highest identifiability as the final candidate set, and uses the attribute data in the final candidate set as the industry's feasible attribute data, and performs weighted fusion of the feasible attribute data based on identifiability.

[0084] Alternatively, the transmission medium may be a wired link, such as, but not limited to, coaxial cable, fiber optic cable and digital subscriber line, or a wireless link, such as, but not limited to, wireless Fidelity (WIFI), Bluetooth and mobile device networks.

[0085] It should be noted that the device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above.

[0086] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. For example, as shown... Figure 3 As shown, the computer device 600 includes: a memory 610, a processor 620, and a computer program 630 stored in the memory 610 and running on the processor 620, wherein when the processor 620 executes the computer program 630, the computer device can execute any of the enterprise multi-dimensional data feature fusion methods combined with AI technology described above.

[0087] Furthermore, embodiments of the present invention also protect an apparatus that may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform an enterprise multi-dimensional data feature fusion method combining AI technology provided by embodiments of the present invention.

[0088] In this embodiment of the invention, the device can be divided into functional modules according to the above method example. For example, each module can correspond to a separate function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.

[0089] When each module is divided according to its function, the device may also include a signal uploading module, a determination module, and an adjustment module. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here.

[0090] It should be understood that the apparatus provided in this embodiment of the invention is used to execute the above-described method for fusing multi-dimensional data features of an enterprise incorporating AI technology, and thus can achieve the same effect as the above-described implementation method.

[0091] When using integrated units, the device may include a processing module and a storage module. When applied to a device, the processing module can be used to control and manage the device's operations. The storage module can be used to support the device in executing program code, etc. The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits as described in this disclosure. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of Digital Signal Processing (DSP) and a microprocessor, etc., and the storage module may be a memory.

[0092] In addition, the device provided in the embodiments of the present invention may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute the enterprise multi-dimensional data feature fusion method combined with AI technology provided in the above embodiments.

[0093] This invention also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the aforementioned method steps to implement the enterprise multi-dimensional data feature fusion method combining AI technology provided in the above embodiments.

[0094] This invention also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to achieve the enterprise multi-dimensional data feature fusion method combining AI technology provided in the above embodiments.

[0095] In this invention, the apparatus, computer-readable storage medium, computer program product, or chip provided in the embodiments are all used to execute the corresponding methods described above. Therefore, the beneficial effects they achieve can be referred to the beneficial effects in the corresponding methods described above, and will not be repeated here. Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In the embodiments provided by this invention, it should be understood that the disclosed apparatus and method can be implemented in other ways.

[0096] The device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0097] It should also be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0098] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0099] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0100] The above content is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. An enterprise multi-dimensional data feature fusion method combined with an AI technology, characterized in that, The method comprises the following steps: Obtain multi-dimensional attribute data of enterprises; Classify the enterprises according to the industries in which the enterprises are located, and mark risk values for different attribute data; For enterprises in the same industry, take attribute data in any dimension as attribute to be analyzed, determine statistical significance of the attribute to be analyzed and the risk value by a K-S test method, compare the attribute to be analyzed and the risk value to determine risk deviation consistency of the attribute to be analyzed, and determine importance of risk evaluation of the attribute to be analyzed in combination with the statistical significance and the risk deviation consistency; The method for obtaining the risk deviation consistency comprises the following steps: Determine a risk evaluation value of the attribute to be analyzed according to a difference between a data value of the attribute to be analyzed and the risk value, determine a dispersion evaluation value of the attribute to be analyzed according to dispersion of the risk evaluation values of the attribute to be analyzed of all enterprises in the same industry, and perform negative correlation normalization mapping on the dispersion evaluation value to obtain the risk deviation consistency of the attribute to be analyzed; Select attribute data in an order of importance, add the attribute data to a candidate set, determine a recognition degree of the attribute data added to the candidate set, select a candidate set corresponding to a maximum recognition degree as a final candidate set, take attribute data in the final candidate set as feasible attribute data of the industry, and perform weighted fusion on the feasible attribute data according to the recognition degree; 2. The enterprise multi-dimensional data feature fusion method combined with an AI technology according to claim 1, characterized in that, The method for obtaining the recognition degree comprises the following steps: Take a set composed of attribute data not added to the candidate set in the enterprises in the same industry as a remaining set, calculate importance of risk evaluation corresponding to the remaining set before and after adding attribute data in the remaining set to the candidate set each time, calculate a difference value of the importance of risk evaluation corresponding to the remaining set before and after adding the attribute data to the candidate set as a change rate before and after adjustment of the candidate set, and determine the recognition degree of the attribute data added to the candidate set in combination with the change rate and a data value of the attribute data added to the candidate set. 3.The enterprise multi-dimensional data feature fusion method combined with AI technology according to claim 1, characterized in that, The method for determining the risk evaluation value of the attribute to be analyzed according to the difference between the data value of the attribute to be analyzed and the risk value comprises the following steps: Calculate a difference value between the data value of the attribute to be analyzed and the risk value as the risk evaluation value of the attribute to be analyzed.

4. The enterprise multi-dimensional data feature fusion method combined with AI technology according to claim 1, characterized in that, The method for determining the dispersion evaluation value of the attribute to be analyzed according to dispersion of the risk evaluation values of the attribute to be analyzed of all enterprises in the same industry comprises the following steps: Calculate a standard deviation of the risk evaluation values of the attribute to be analyzed of all enterprises in the same industry as the dispersion evaluation value of the attribute to be analyzed. The method for determining the importance of risk evaluation of the attribute to be analyzed in combination with the statistical significance and the risk deviation consistency comprises the following steps: Determine a rank correlation coefficient of the risk evaluation value and the data value of the attribute to be analyzed; 5. The enterprise multi-dimensional data feature fusion method combined with AI technology according to claim 4, characterized in that, Calculate a normalized value of a difference between the rank correlation coefficient of the attribute to be analyzed and the statistical significance as a data correlation of the attribute to be analyzed after interference of data distribution is excluded; Determine the importance of risk evaluation of the attribute to be analyzed in combination with the data correlation, the risk deviation consistency and the statistical significance. The method for determining the importance of risk evaluation of the attribute to be analyzed in combination with the data correlation, the risk deviation consistency and the statistical significance comprises the following steps: Compare the risk deviation consistency and the data correlation, obtain a maximum value in the risk deviation consistency and the data correlation as an important judgment value; Combine the important judgment value and the statistical significance to obtain the importance of the risk assessment of the attribute to be analyzed, wherein the important judgment value and the statistical significance are positively correlated with the importance of the risk assessment.

6. The enterprise multi-dimensional data feature fusion method combined with an AI technology according to claim 5, characterized in that, The combining the important judgment value and the statistical significance to obtain the importance of the risk assessment of the attribute to be analyzed comprises: Calculating a product value of the important judgment value and the statistical significance as the importance of the risk assessment of the attribute to be analyzed.

7. The enterprise multi-dimensional data feature fusion method combined with an AI technology according to claim 1, characterized in that, The selecting the candidate set corresponding to the maximum recognition degree as the final candidate set comprises: According to the size order of the importance, iteratively selecting attribute data, adding the attribute data to the candidate set, until the importance of the risk assessment of the attribute after the attribute data is added to the candidate set reaches a preset threshold, stopping adding attribute data to the candidate set, determining the recognition degree of each attribute data added to the candidate set, and selecting the candidate set corresponding to the maximum recognition degree as the final candidate set.

8. The enterprise multi-dimensional data feature fusion method combined with an AI technology according to claim 7, characterized in that, The weighted fusion of the feasibility attribute data according to the recognition degree comprises: Taking the recognition degree of each feasibility attribute data added to the candidate set as the weight of the corresponding feasibility attribute data, and performing weighted fusion on the feasibility attribute data.

Citation Information

Patent Citations

  • A financial risk assessment method, device, equipment and storage medium

    CN119762235A

  • Multi-source data fusion enterprise finance and tax integrated risk management and control platform

    CN120107004A