An enterprise data element governance method and device based on big data
By classifying and weighting analysis of enterprise public data based on big data, the enterprise portraits are generated, which solves the problems of high labor costs and poor targeting in the existing technology, and achieves efficient and accurate enterprise portrait generation.
Patent Information
- Application Number
- CN202510126411.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-01-27
AI Technical Summary
When generating corporate portraits in the prior art, there are problems such as high labor costs and poor targeting, and it is difficult to generate reference labels in front of huge corporate public data information.
Through a big data-based method, the public data information of the target enterprise is determined, the weight coefficient is determined according to the correlation, and the neural network model is used to generate data element labels, establish mapping relationships, and generate a portrait of the target enterprise.
It realizes efficient and accurate generation of corporate portraits, reduces labor costs, and improves the targeted labels and the reference value of corporate portraits.
Smart Images

Figure CN119557739B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data mining, and more particularly, to a method and apparatus for enterprise data element governance based on big data. Background Art
[0002] With the development of big data, "digital transformation" has become the main theme of the development of all industries. Enterprises, as the basic units of social and economic activities, play an important role in promoting the development of the data economy.
[0003] Currently, the main difficulty in the digital transformation of enterprises lies in the lack of understanding of their own development status and market competitiveness, resulting in the inability to accurately formulate development strategies that meet their own needs.
[0004] Therefore, it is becoming increasingly important to construct enterprise portraits.
[0005] Enterprise portraits can use multi-dimensional views of data to objectively and truly reflect the characteristics of each enterprise.
[0006] Currently, the way to generate enterprise portraits is usually to analyze the public data of enterprises to generate a series of enterprise tags.
[0007] In the face of a vast amount of public enterprise information, the method of manually tagging by data annotators often consumes a large amount of labor costs.
[0008] Currently, there are also some technologies that analyze public data through machine learning to generate tags. However, this method has poor pertinence. In the face of a vast amount of public enterprise data information, the workload of data analysis is very large, and the types and quantities of the corresponding generated tags are chaotic, which is not conducive to generating valuable enterprise portraits for reference.
[0009] Therefore, those skilled in the art urgently need to find a new technical solution to solve the above problems. Summary of the Invention
[0010] To overcome the problems existing in the related technologies, the present disclosure provides a method and apparatus for enterprise data element governance based on big data.
[0011] According to the first aspect of the embodiments of the present disclosure, a method for enterprise data element governance based on big data is provided. The method includes:
[0012] Determine the target enterprise that needs to conduct data element governance, and collect the public data information of the target enterprise;
[0013] Divide the public data information into several data classifications according to data relevance;
[0014] Determine the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise;
[0015] Determine the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, so as to determine the target data label of the target enterprise according to the mapping relationship;
[0016] Generate a target enterprise portrait according to the target data label and the weight coefficient to complete the data element governance of the target enterprise.
[0017] Optionally, the target enterprise that needs to perform data element governance is determined, and the public data information of the target enterprise is collected, including:
[0018] Determine the target enterprise that needs to perform data element governance;
[0019] Collect the public data information of the target enterprise from public channels.
[0020] Optionally, the public data information is divided into several data classifications according to data relevance, including:
[0021] For each piece of public data information , calculate the first correlation coefficient between the public data information and other public data information respectively, and obtain the first correlation coefficient between every two pieces of public data information, where is the expected value;
[0022] Divide the public data information with the first correlation coefficient greater than the preset coefficient threshold into one data classification;
[0023] Obtain several data classifications.
[0024] Optionally, the step of determining the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise includes:
[0025] Determine the weight coefficient corresponding to each data classification according to the correlation between the enterprise type and the data classification , where is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value among all the second correlation coefficients between the data classifications and the enterprise type, is the minimum value among all the second correlation coefficients between the data classifications and the enterprise type;
[0026] The method for obtaining the second correlation coefficient is:
[0027] For each data classification, obtain each piece of public data information in the data classification The third correlation coefficient with the enterprise type y ;
[0028] Calculate the average value of all the third correlation coefficients in the data classification to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
[0029] Optionally, the determining the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, so as to determine the target data label of the target enterprise according to the mapping relationship includes:
[0030] According to the historical mapping relationship between the historical data classification and the historical data element label, use each piece of historical public data information in the historical data classification as the input, and use the historical data element label corresponding to the historical data classification as the output to train the neural network model to obtain a trained data element label generation model;
[0031] Use each piece of public data information in the data classification as the input of the data element label generation model, and obtain the data element label corresponding to each data classification of the target enterprise according to the output of the data element label generation model.
[0032] Optionally, the method further includes:
[0033] Visually display the target enterprise portrait.
[0034] According to the second aspect of the disclosed embodiments of the present invention, there is provided a big data-based enterprise data element governance device, and the device includes:
[0035] A data collection module, which determines a target enterprise that needs to perform data element governance and collects the public data information of the target enterprise;
[0036] A data classification module, connected to the data collection module, divides the public data information into several data classifications according to data relevance;
[0037] A weight coefficient determination module, connected to the data classification module, determines the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise;
[0038] A data label determination module, connected to the weight coefficient determination module, determines the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, so as to determine the target data label of the target enterprise according to the mapping relationship;
[0039] An enterprise profile generation module, connected to the data label determination module, generates a target enterprise profile according to the target data label and the weight coefficient to complete the governance of the data elements of the target enterprise.
[0040] Optionally, the data classification module includes:
[0041] A first correlation coefficient determination unit, for each piece of public data information , calculates the first correlation coefficient between the public data information and other public data information respectively, obtains the first correlation coefficients between every two public data information, where is the expected value;
[0042] A data classification unit, connected to the first correlation coefficient determination unit, classifies the public data information with the first correlation coefficient greater than the preset coefficient threshold into one data classification;
[0043] A data classification acquisition unit, connected to the data classification unit, obtains several data classifications.
[0044] Optionally, the weight coefficient determination module includes:
[0045] Determines the weight coefficient corresponding to each data classification according to the correlation between the enterprise type and the data classification , where is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value among all the second correlation coefficients between all data classifications and the enterprise type, is the minimum value among all the second correlation coefficients between all data classifications and the enterprise type;
[0046] The method for obtaining the second correlation coefficient is:
[0047] For each data classification, obtains the third correlation coefficient between each piece of public data information in the data classification and the enterprise type y;
[0048] Calculates the average value of all the third correlation coefficients in the data classification to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
[0049] Optionally, the data label determination module includes:
[0050] A model training unit trains a neural network model by using each piece of historical public data information in a historical data classification as an input and the historical data element label corresponding to the historical data classification as an output according to the historical mapping relationship between the historical data classification and the historical data element label, and obtains a trained data element label generation model;
[0051] A data label acquisition unit is connected to the model training unit, uses each piece of public data information in the data classification as an input to the data element label generation model, and obtains the data element label corresponding to each data classification of the target enterprise according to the output of the data element label generation model.
[0052] In summary, the present invention discloses an enterprise data element governance method and device based on big data. The method includes: determining a target enterprise that needs to perform data element governance, and collecting public data information of the target enterprise; dividing the public data information into several data classifications according to data relevance; determining a weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise; determining the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, so as to determine the target data label of the target enterprise according to the mapping relationship; generating a target enterprise portrait according to the target data label and the weight coefficient, so as to complete the data element governance of the target enterprise.
[0053] It is possible to classify data, perform weight analysis using multi-dimensional data, conveniently label each type of data, generate an enterprise portrait according to the label and the corresponding weight coefficient, and provide convenience for subsequent data processing.
[0054] Other features and advantages of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure.
[0056] In the drawings:
[0057] Figure 1 is a flowchart showing a method for enterprise data element governance based on big data according to an exemplary embodiment;
[0058] Figure 2 is according to Figure 1 shows a flowchart of a data classification division method;
[0059] Figure 3 is according to Figure 1Flow schematic diagram of a data tag acquisition method shown;
[0060] Figure 4 It is a structural block diagram of an enterprise data element governance device based on big data shown according to an exemplary embodiment;
[0061] Figure 5 It is according to Figure 4 Structural block diagram of a data classification module shown;
[0062] Figure 6 It is according to Figure 4 Structural block diagram of a data tag determination module shown. Detailed implementation manners
[0063] The following will describe in detail the specific implementation manners disclosed in the present invention with reference to the accompanying drawings.
[0064] It should be understood that the specific implementation manners described herein are only for explaining and illustrating the present disclosure, and are not used to limit the present disclosure.
[0065] Figure 1 It is a flow schematic diagram of an enterprise data element governance method based on big data shown according to an exemplary embodiment, as Figure 1 shown, and the method includes:
[0066] In step 101, determine the target enterprise that needs to conduct data element governance, and collect the public data information of the target enterprise.
[0067] Exemplarily, determine the target enterprise that needs to conduct data element governance; collect the public data information of the target enterprise from public channels.
[0068] The public channels include: enterprise official websites, patent databases, brand registration information, recruitment websites, business query platforms, etc. Obtain the public information of the enterprise from the above channels to conduct multi-dimensional analysis of the public data information of the enterprise.
[0069] The public data information obtained from public channels may include: enterprise operating income, number of social security participants, industry where the enterprise is located, types and amounts of data assets owned, recruitment data, land use situation, bidding, patents, software copyrights, news, major authoritative certification lists, etc.
[0070] It can be understood that the manner of obtaining public data information can be: cloud computing service technology, artificial intelligence technology, manual collection by data collectors or provided by a third party, etc.
[0071] In addition to the data obtained from public channels, information data from the target enterprise's internal management system or data provided by third parties can also be integrated to increase the data source channels, so as to more comprehensively analyze the data element tags of the target enterprise.
[0072] It should be noted that the enterprises carrying out data element governance usually include: (1) Data resource enterprises: Enterprises that own data resources, take data as their core assets, or are engaged in the integrated utilization of industrial chain data, and whose core business activities are data acquisition, accumulation, management, and analysis.
[0073] (2) Data technology enterprises: Enterprises engaged in data computing, storage, mining, and analysis, with technical service capabilities such as data aggregation, storage, governance, transmission, management, security, and development of algorithm models, and can use technical means or data tools to improve data quality, realize data elementization, and provide basic conditions for empowering data applications.
[0074] (3) Data application enterprises: Focusing on key industrial fields such as intelligent manufacturing, low-altitude economy, intelligent connected vehicles, digital finance, medical and health, digital culture, commerce, and logistics, they use enterprise or industry-precipitated data and public data to empower the entire process of enterprise R & D and production and the deep transformation of the industry, improve the life scene experience, and promote business model innovation.
[0075] (4) Data service enterprises: Centering on market demand, by providing professional services such as data asset accounting, quality assessment, compliance consulting, and security assessment, they help enterprises better understand and utilize data, and then improve the decision-making quality, operation efficiency, market competitiveness, and innovation and development level of enterprises.
[0076] (5) Data security enterprises: Mainly provide network security, information security, and data security guarantees, and provide services such as classification and grading of enterprise sensitive data, data desensitization, deep access control, security inspection, prevention of data leakage, and traceability tracking to achieve the trustworthy and secure circulation of data.
[0077] (6) Data infrastructure enterprises: Build and operate network infrastructure (such as 5G, industrial Internet, Internet of Things, etc.), computing power infrastructure (such as data centers, intelligent computing centers, supercomputing centers, cloud computing, etc.), data circulation facilities (such as data spaces, blockchain platforms, privacy computing platforms, data components, data networks, data markets, etc.), and data security infrastructure (such as data compliance platforms, etc.).
[0078] In step 102, the public data information is divided into several data classifications according to data relevance.
[0079] Exemplarily, the public data information of the target enterprise usually contains multiple dimensions, and the data of different dimensions are divided into their respective data classifications.
[0080] When classifying, according to the correlation between public data information, data with high correlation is classified into one data category, and corresponding data labels are assigned to the target enterprise according to each data category, which is beneficial to improving the accuracy of data labels.
[0081] Facing the huge and miscellaneous public data information obtained, if directly processing tags according to each piece of public data, not only is the workload huge, but the generated tags also lack strong pertinence, and it is difficult to draw an accurate enterprise portrait based on the tags in the follow-up.
[0082] First, classify the public data information according to the correlation, and generate labels for each data category.
[0083] It can be understood that before dividing the public data into each data category, it is also necessary to perform data cleaning and preprocessing operations on the public data information, including: removing duplicate data records; processing missing data by interpolation, mean filling or deletion; unifying the formats of all public data information; identifying and processing outliers through statistical methods or machine learning algorithms.
[0084] Through the above noise reduction processing, valid data is retained, which is convenient for the subsequent data analysis process.
[0085] Specifically, Figure 2 is based on Figure 1 a schematic flowchart of a data classification and division method shown, as Figure 2 shown, this step 102 includes:
[0086] In step 1021, for each piece of public data information, calculate the first correlation coefficient between this public data information and other public data information respectively.
[0087] Among them, for each piece of public data information , calculate this public data information and other public data information the first correlation coefficient between them respectively, obtain the first correlation coefficient between pairwise public data information, is the expected value.
[0088] Exemplarily, organize the information obtained from public channels to obtain each piece of public data information, and the organization process can be through natural language processing technology, sentiment analysis technology, entity recognition technology, etc. It can be understood that each piece of public data information can be vectorized to facilitate the calculation of the similarity between vectors.
[0089] In step 1022, divide the public data information with the first correlation coefficient greater than the preset coefficient threshold into one data category.
[0090] In step 1023, a number of data classifications are obtained.
[0091] Exemplarily, calculate the first correlation coefficient between pairwise public data, traverse all the first correlation coefficients, and divide the public data information with the first correlation coefficient greater than a preset threshold into one data classification until all the public data information belongs to the corresponding data classification.
[0092] In step 103, according to the enterprise type of the target enterprise, determine the weight coefficient corresponding to each data classification.
[0093] Exemplarily, according to the enterprise type of the target enterprise (such as: educational enterprise, legal service enterprise, technology enterprise, cultural and entertainment enterprise, etc.), determine the weight coefficients of each data classification. Different types of enterprises may have different focuses in different dimensions.
[0094] For example, for a technology enterprise, technological innovation is the core competitiveness of the enterprise, and the weight coefficients corresponding to the data classifications in the dimensions of patent information, honors obtained in scientific and technological innovation competitions, product innovation, etc. in the data classification should be relatively high.
[0095] Specifically, determining the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise includes:
[0096] Determine the weight coefficient corresponding to each data classification according to the correlation between the enterprise type and the data classification , where is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value among all the second correlation coefficients between all data classifications and the enterprise type, is the minimum value among all the second correlation coefficients between all data classifications and the enterprise type.
[0097] Exemplarily, regard the enterprise type as a piece of public data information of the enterprise, calculate the second correlation coefficient between the enterprise type and each data classification, and determine the weight coefficient corresponding to each data classification according to the second correlation coefficient.
[0098] Among them, the method for obtaining the second correlation coefficient is:
[0099] For each data classification, obtain each piece of public data information in this data classification and the third correlation coefficient with the enterprise type y ;
[0100] Calculate the mean value of all the third correlation coefficients in the data classification to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
[0101] In step 104, based on the historical data classification and historical data element labels, determine the mapping relationship between the data classification and the data element labels, and based on this mapping relationship, determine the target data labels of the target enterprise.
[0102] Specifically, Figure 3 is a schematic flowchart of a method for obtaining data labels shown in Figure 1 As shown in Figure 3 This step 104 includes:
[0103] In step 1041, based on the historical mapping relationship between the historical data classification and the historical data element labels, use each historical public data information in the historical data classification as input, and use the historical data element labels corresponding to this historical data classification as output to train the neural network model and obtain a trained data element label generation model.
[0104] In step 1042, use each public data information in the data classification as input to the data element label generation model, and based on the output of the data element label generation model, obtain the data element labels corresponding to each data classification of the target enterprise.
[0105] Exemplarily, the historical data element labels can be labels manually annotated by data annotators for each historical public data. Divide the historical public data into multiple historical data classifications through the correlation coefficient formula in the disclosed embodiments of the present invention, and sort out the historical data element labels corresponding to each historical data classification.
[0106] Train the neural network model through the historical data classification and historical data element labels to obtain a data element label generation model, so as to output the data element labels of the target enterprise.
[0107] When performing model training, text preprocessing and feature screening can be carried out through technologies such as pandas and scikit-learn, machine learning or deep learning can be used for model training, scikit-learn technology can be used for model evaluation, and grid search tuning or Bayesian optimization random tuning technology can be used for model parameter tuning.
[0108] In step 105, generate a target enterprise portrait based on the target data labels and weight coefficients to complete the data element governance of the target enterprise.
[0109] Exemplarily, after generating the target enterprise portrait, it can be combined with an intelligent recommendation system to provide accurate enterprise matching services for investors, partners, etc. based on the enterprise portrait, or combined with a risk warning system to establish a risk warning mechanism to timely discover potential problems and take response measures.
[0110] In addition, the target enterprise portrait can also be visually displayed.
[0111] Exemplarily, develop a visualization tool. For example, display the performance of the target enterprise in each dimension through charts to help the management make more informed decisions and help the enterprise better understand and apply the label results.
[0112] Figure 4 It is a structural block diagram of an enterprise data element governance device based on big data shown according to an exemplary embodiment. As Figure 4 shown, the device 400 includes:
[0113] A data collection module 410 determines a target enterprise that needs to conduct data element governance and collects public data information of the target enterprise;
[0114] A data classification module 420 is connected to the data collection module 410 and divides the public data information into several data classifications according to data relevance;
[0115] A weight coefficient determination module 430 is connected to the data classification module 420 and determines the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise;
[0116] A data label determination module 440 is connected to the weight coefficient determination module 430, determines the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, and determines the target data label of the target enterprise according to the mapping relationship;
[0117] An enterprise portrait generation module 450 is connected to the data label determination module 440, generates a target enterprise portrait according to the target data label and the weight coefficient to complete the data element governance of the target enterprise.
[0118] Figure 5 It is according to Figure 4 shown, a structural block diagram of a data classification module. As Figure 5 shown, the data classification module 420 includes:
[0119] A first correlation coefficient determination unit 421 calculates the first correlation coefficient between each piece of public data information and other public data information respectively, obtains the first correlation coefficients between pairs of public data information, where is the expected value; is the expected value;
[0120] The data classification unit 422, connected to the first correlation coefficient determination unit 421, classifies the public data information with the first correlation coefficient greater than the preset coefficient threshold into one data classification;
[0121] The data classification acquisition unit 423, connected to the data classification unit 422, acquires a number of data classifications.
[0122] Optionally, the weight coefficient determination module 430 includes:
[0123] Determine the weight coefficient corresponding to each data classification according to the correlation between the industry type and the data classification , where is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value among all the second correlation coefficients between all data classifications and the enterprise type, is the minimum value among all the second correlation coefficients between all data classifications and the enterprise type;
[0124] The method for obtaining the second correlation coefficient is:
[0125] For each data classification, obtain each public data information in this data classification and the third correlation coefficient between the enterprise type y ;
[0126] Calculate the mean value of all the third correlation coefficients in the data classification to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
[0127] Figure 6 is according to Figure 4 shows a structural block diagram of a data label determination module, as Figure 6 shown, the data label determination module 440 includes:
[0128] The model training unit 441, according to the historical mapping relationship between the historical data classification and the historical data element label, uses each historical public data information in the historical data classification as input and the historical data element label corresponding to the historical data classification as output to train the neural network model and obtain a trained data element label generation model;
[0129] The data label acquisition unit 442, connected to the model training unit 441, uses each public data information in the data classification as the input of the data element label generation model, and obtains the data element label corresponding to each data classification of the target enterprise according to the output of the data element label generation model.
[0130] In summary, the present invention discloses a method and apparatus for enterprise data element governance based on big data. The method includes: determining a target enterprise that needs to conduct data element governance, and collecting public data information of the target enterprise; dividing the public data information into several data classifications according to data relevance; determining a weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise; determining the mapping relationship between the data classification and the data element label according to the historical data classification and the historical data element label, so as to determine the target data label of the target enterprise according to the mapping relationship; generating a target enterprise portrait according to the target data label and the weight coefficient, so as to complete the data element governance of the target enterprise.
[0131] It is possible to classify data, perform weight analysis using multi-dimensional data, conveniently label each type of data, and generate an enterprise portrait according to the label and the corresponding weight coefficient, providing convenience for subsequent data processing.
[0132] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0133] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.
[0134] Furthermore, any combination can be made between various different embodiments of the present disclosure, as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
Claims
1. A method for enterprise data element governance based on big data, characterized in that: The method comprises: Identify target enterprises that require data element governance and collect public data information of the target enterprises; Dividing the public data information into several data categories according to data relevance; Determine a weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise; Determining a mapping relationship between the data classification and the data element label based on the historical data classification and the historical data element label, and determining a target data label for the target enterprise based on the mapping relationship; Generate a target enterprise portrait based on the target data label and weight coefficient to complete the data element governance of the target enterprise; The public data information is divided into several data categories according to data relevance, including: For each piece of public data information , respectively calculate the public data information and other public data information The first correlation coefficient between , obtain the first correlation coefficient between the public data information, where, is the expected value; Classify the public data information whose first correlation coefficient is greater than a preset coefficient threshold into a data category; Get several data categories; Determining the weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise includes: Determine the weight coefficient corresponding to each data classification based on the correlation between the enterprise type and the data classification ,in, is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value of the second correlation coefficient between all data categories and enterprise types, The minimum value of the second correlation coefficient between all data categories and enterprise types; The method for obtaining the second correlation coefficient is: For each data category, obtain each public data information in the data category The third correlation coefficient between the enterprise type y ; The average of all third correlation coefficients in the data classification is calculated to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
2. The enterprise data element governance method based on big data according to claim 1 is characterized in that: The determination of target enterprises requiring data element governance and the collection of public data information of the target enterprises include: Identify target enterprises that require data element governance; Collect public data information of the target enterprise from public channels.
3. The enterprise data element governance method based on big data according to claim 1 is characterized in that: Determining a mapping relationship between data classification and data element label based on historical data classification and historical data element label, and determining a target data label for the target enterprise based on the mapping relationship, includes: According to the historical mapping relationship between historical data classification and historical data element labels, each piece of historical public data information in the historical data classification is used as input, and the historical data element label corresponding to the historical data classification is used as output to train the neural network model to obtain a trained data element label generation model; Each piece of public data information in the data classification is used as the input of the data element label generation model, and the data element label corresponding to each data classification of the target enterprise is obtained according to the output of the data element label generation model.
4. The enterprise data element governance method based on big data according to claim 1 is characterized in that: The method further comprises: Provide a visual display of the target enterprise portrait.
5. An enterprise data element management device based on big data, characterized in that: The device comprises: The data collection module determines the target enterprises that need data element governance and collects the public data information of the target enterprises; A data classification module, connected to the data acquisition module, divides the public data information into several data categories according to data relevance; A weight coefficient determination module, connected to the data classification module, determines a weight coefficient corresponding to each data classification according to the enterprise type of the target enterprise; a data label determination module, connected to the weight coefficient determination module, determining a mapping relationship between data classification and data element label based on historical data classification and historical data element label, and determining a target data label for the target enterprise based on the mapping relationship; An enterprise portrait generation module, connected to the data tag determination module, generates a target enterprise portrait based on the target data tag and weight coefficient to complete the data element governance of the target enterprise; The data classification module includes: The first correlation coefficient determination unit, for each piece of public data information , respectively calculate the public data information and other public data information The first correlation coefficient between , obtain the first correlation coefficient between the public data information, where, is the expected value; a data classification unit, connected to the first correlation coefficient determination unit, for classifying public data information having a first correlation coefficient greater than a preset coefficient threshold into a data classification; a data classification acquisition unit, connected to the data classification unit, to acquire a plurality of data classifications; The weight coefficient determination module includes: Determine the weight coefficient corresponding to each data classification based on the correlation between the enterprise type and the data classification ,in, is the second correlation coefficient between the i-th data classification and the enterprise type, is the maximum value of the second correlation coefficient between all data categories and enterprise types, The minimum value of the second correlation coefficient between all data categories and enterprise types; The method for obtaining the second correlation coefficient is: For each data category, obtain each public data information in the data category The third correlation coefficient between the enterprise type y ; The average of all third correlation coefficients in the data classification is calculated to obtain the second correlation coefficient between the i-th data classification and the enterprise type.
6. The enterprise data element management device based on big data according to claim 5 is characterized in that: The data label determination module includes: a model training unit, which trains a neural network model based on a historical mapping relationship between historical data classifications and historical data element labels, taking each piece of historical public data information in the historical data classification as input and the historical data element label corresponding to the historical data classification as output, and obtains a trained data element label generation model; A data label acquisition unit is connected to the model training unit, and uses each public data information in the data classification as the input of the data element label generation model, and obtains the data element label corresponding to each data classification of the target enterprise based on the output of the data element label generation model.
Citation Information
Patent Citations
Enterprise environment social governance portrait construction method and system based on big data
CN115829400A