Enterprise subject type identification method and device, equipment and medium
By using basic power data and enterprise identification information, combined with the target probability identification model, predicting and determining the enterprise entity type, the problems of large workload, long cycle and poor results in the existing technology are solved, and efficient and accurate enterprise entity type identification is achieved.
Patent Information
- Application Number
- CN202510557540.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
There are problems such as large workload, long cycle and poor results in determining the type of entity in the existing enterprise.
By obtaining the basic power data and enterprise identification information in the target area, using the pre-trained target probability identification model, predicting the probability data of the enterprise setting the subject type, and determining the entity type of the enterprise based on the probability data.
Improve the accuracy and efficiency of enterprise entity type determination, and reduce workload and cycle.
Smart Images

Figure CN120493059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to an enterprise entity type identification device, equipment and medium. Background Art
[0002] The classification of private economic market operators is an important means for power companies to provide power big data services to the government and economic development.
[0003] Currently, the technology for classifying private economic market operators is primarily based on screening electricity usage records and external data classification. This is done by selecting the enterprise type corresponding to the beginning of the unified social credit code, as published by the National Organization Unified Social Credit Code Data Service Center, and excluding state-controlled and foreign-controlled enterprises. Alternatively, external data linkage can be used to fill in the gaps in the classification process. Data management units acquire information on private economic market operators through enterprise data procurement, self-collection, or purchase from authoritative institutions. After acquisition, the information is then matched and linked with power system archival data to classify the enterprises. However, current methods for classifying entities suffer from high workload, long cycles, and poor results. Summary of the Invention
[0004] The present invention provides an enterprise entity type identification device, equipment and medium to solve the problems of heavy workload, long cycle and poor effect in the existing enterprise entity type determination process.
[0005] According to one aspect of the present invention, a method for identifying an enterprise entity type is provided, comprising:
[0006] Obtain basic power data and enterprise identification information of each enterprise in the target area;
[0007] Inputting the basic power data and the enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set;
[0008] The target entity type corresponding to the target enterprise is determined according to the probability data.
[0009] According to another aspect of the present invention, there is provided a device for identifying an enterprise entity type, comprising:
[0010] The data acquisition module is used to obtain the basic power data and enterprise identification information of each enterprise in the target area;
[0011] a probability recognition module, configured to input the basic power data and the enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set;
[0012] The subject type determination module is used to determine the target subject type corresponding to the target enterprise according to the probability data.
[0013] According to another aspect of the present invention, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the enterprise entity type identification method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the enterprise entity type identification method described in any embodiment of the present invention when executed.
[0018] The technical solution of the embodiment of the present invention obtains the basic power data and enterprise identification information of each enterprise in the target area; inputs the basic power data and the enterprise identification information into the target probability recognition model to obtain the probability data of the target enterprise being the set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set; and determines the target subject type corresponding to the target enterprise based on the probability data. This technical solution obtains the probability of an enterprise being the set subject type through model prediction, and then determines the enterprise subject type based on the probability data. This can solve the problems of large workload, long cycle, and poor effect in the existing process of determining the enterprise subject type.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a flow chart of a method for identifying an enterprise entity type according to the first embodiment of the present invention;
[0022] Figure 2 This is a flow chart of a method for identifying enterprise entity types provided according to the second embodiment of the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a device for identifying enterprise entity types according to a third embodiment of the present invention;
[0024] Figure 4 It is a structural diagram of an electronic device provided according to the fourth embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first", "second", "initial" and "target" in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] Example 1
[0028] Figure 1This is a flow chart of a method for identifying an enterprise entity type according to a first embodiment of the present invention. This embodiment is applicable to situations where the entity type of an enterprise is to be identified. The method can be executed by an enterprise entity type identification device. The enterprise entity type identification device can be implemented in the form of hardware and / or software. The enterprise entity type identification device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:
[0029] S110: Obtain basic power data and enterprise identification information of each enterprise in the target area.
[0030] The target area can be understood as the area where enterprise type identification is required. In this embodiment, the target area can be individual regions determined based on provinces, or designated regions determined based on other divisions, and can be determined based on actual needs. In this embodiment, the target area can include individual enterprises. Basic power data can include electricity consumption data, load data, capacity data, and distribution data. The basic power data in this embodiment can be time-limited. For example, basic power data for each enterprise in the target area can be obtained for a month, or it can be for other time periods. This can be set based on actual needs and is not limited in this embodiment. Distribution data can specifically include regional distribution data and industry distribution data. Regional distribution data can refer to the distribution of enterprises in different regions within a province. Industry distribution data can refer to the distribution of the industries to which an enterprise's business belongs. Electricity consumption data can refer to the total amount of electricity consumed or generated by the enterprise's power system during a set time period. The set time period can include daily, monthly, and annual time periods. In this embodiment, electricity consumption data can be understood as the enterprise's monthly electricity consumption data. Load data can refer to the electricity demand at a specific point or area in the enterprise's power system at a specific time. Capacity data may refer to the capacity data of an enterprise, specifically representing the scale of the enterprise. For example, enterprises may be categorized as large, medium, or small based on their capacity data, or may be further categorized. Enterprise identification information may be understood as identifying information for an enterprise. In this embodiment, enterprise identification information may refer to the enterprise's unified social credit code information.
[0031] In this embodiment, the power data, load data, capacity data, distribution data and other data of each enterprise included in the target area and the corresponding unified social credit code information of each enterprise can be obtained.
[0032] S120: Inputting the basic power data and enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type.
[0033] The target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is trained using the initial probability recognition model and the power sample dataset. In this embodiment, the target enterprise can be understood as any enterprise in the target area, each of which has its corresponding enterprise identification information. The target probability recognition model in this embodiment can be a pre-trained recognition model that is used to identify the probability data of the enterprise being the target entity type. The set entity type can be a private economic market operating entity. The probability data can be understood as the predicted probability of the enterprise being a private economic market operating entity type.
[0034] This example investigates relevant information about private economic market operators to identify their primary characteristics. It also provides an objective definition of these entities, which can include registered non-state-owned enterprises, non-state-controlled enterprises, and non-foreign-owned market operators. In this example, the power infrastructure data and enterprise identification information corresponding to each enterprise can be input into a target probability identification model to determine the probability that each enterprise is a private economic market operator.
[0035] In this embodiment, optionally, the basic power data includes electricity data, load data, capacity data and distribution data; the basic power data and enterprise identification information are input into the target probability recognition model to obtain probability data that the target enterprise is the set subject type, including: inputting the electricity data, load data, capacity data and distribution data and the enterprise identification information into the target probability recognition model to obtain probability data that the target enterprise is the set subject type.
[0036] In this embodiment, the electricity data, load data, capacity data and distribution data of each enterprise in a set time period and the identification information corresponding to each enterprise can be input into a pre-trained target probability recognition model to obtain the probability data of each enterprise being an operating entity in the private economic market.
[0037] Through such a setting in this embodiment, the probability of each enterprise being a set subject type can be predicted based on the basic power data and probability identification model of each enterprise in the target area, so as to obtain the subject type situation of each enterprise, thereby improving the accuracy of determining the enterprise subject type.
[0038] S130. Determine the target entity type corresponding to the target enterprise based on the probability data.
[0039] The target entity type may refer to a private economic market operating entity or a non-private economic market operating entity. In this embodiment, the probability data of each enterprise being a set entity type, obtained through model prediction, can be compared with a set threshold value. Based on the comparison result, the target entity type corresponding to each enterprise can be determined as a private economic market operating entity or a non-private economic market operating entity.
[0040] The technical solution of the embodiment of the present invention obtains the basic power data and enterprise identification information of each enterprise in the target area; inputs the basic power data and enterprise identification information into the target probability recognition model to obtain probability data that the target enterprise is a set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set; and the target subject type corresponding to the target enterprise is determined based on the probability data. This technical solution obtains the probability of an enterprise being a set subject type through model prediction, and then determines the enterprise subject type based on the probability data. This can solve the problems of large workload, long cycle time, and poor results in the existing process of determining the subject type of an enterprise.
[0041] Example 2
[0042] Figure 2 This is a flow chart of a method for identifying the type of enterprise subject provided in accordance with the second embodiment of the present invention. This embodiment is optimized based on the above embodiment. The specific optimization is: determining the target subject type corresponding to the target enterprise according to the probability data, including: judging whether the probability data is greater than the probability threshold; if the probability data is greater than the probability threshold, the enterprise subject type corresponding to the target enterprise is a private enterprise subject type; if the probability data is less than or equal to the probability threshold, the enterprise subject type corresponding to the target enterprise is a non-private enterprise subject type. Figure 2 As shown, the method includes:
[0043] S210: Obtain basic power data and enterprise identification information of each enterprise in the target area.
[0044] S220: Input the basic power data and enterprise identification information into the target probability recognition model to obtain probability data that the target enterprise is the set subject type.
[0045] Among them, the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set.
[0046] S230: Determine whether the probability data is greater than a probability threshold.
[0047] The probability threshold may be a pre-set threshold for the probability data. For example, the probability threshold in this embodiment may be 0.5 or 0.8, and may also be set based on actual needs. In this embodiment, the probability data predicted by the judgment model as the target enterprise being a set entity type is determined to be greater than the pre-set probability threshold.
[0048] S240. If the probability data is greater than the probability threshold, the enterprise entity type corresponding to the target enterprise is a private enterprise entity type.
[0049] The private enterprise entity type may refer to the type of private economic market operating entity. In this embodiment, if the probability data of the target enterprise being the set entity type obtained through model prediction is greater than a pre-set probability threshold, then the enterprise entity type corresponding to the enterprise can be considered to be a private enterprise entity type.
[0050] S250. If the probability data is less than or equal to the probability threshold, the enterprise entity type corresponding to the target enterprise is a non-private enterprise entity type.
[0051] The non-private enterprise entity type may refer to a type of entity that is not a private economic market operator. In this embodiment, if the probability data of the target enterprise being the set entity type obtained through model prediction is less than or equal to a pre-set probability threshold, the enterprise entity type corresponding to the enterprise can be considered to be a non-private enterprise entity type.
[0052] For example, in this embodiment, Pi can be the probability data of a certain enterprise as the set subject type, and the probability threshold is 0.5. Then, the enterprise with Pi>0.5 is considered to be a private economic market operating entity, and the sample with Pi<0.5 is a non-private economic market operating entity, which can be marked as positive and negative respectively. At the same time, all enterprises are sorted in descending order according to Pi. The higher the enterprise is, the greater the probability of being a private economic market operating entity from the perspective of power data. A certain number of enterprises can be selected in batches according to actual needs. In this embodiment, the private entity identification results of the enterprise and the corresponding enterprise information can also be listed as a result list table of model prediction, and then statistical analysis is performed based on the generated result list table combined with dimensional data such as electricity, load, and capacity to generate a result analysis report related to the electricity consumption of private economic market operating entities in different industries and regions.
[0053] In this embodiment, optionally, the step of obtaining a target probability recognition model by training an initial probability recognition model and a power sample data set includes: obtaining a power sample data set; extracting a set proportion of data from the power sample data set as a training sample data set; screening the training sample data set to obtain a labeled sample data set; and training the initial probability recognition model based on the labeled sample data set and the training sample data set to obtain a target probability recognition model.
[0054] The power sample dataset can be a full dataset consisting of sample data obtained by processing basic power data from various provinces and regions within a historical period obtained from Taichung, the power grid headquarters data. The set ratio can be a pre-set ratio. For example, the set ratio can be between 20% and 80%, which can be set based on actual needs. In this embodiment, a stratified sampling approach can be used to extract a set ratio of data from the power sample dataset as the training sample dataset. The set ratio in this embodiment can be determined based on the actual sample data. Specifically, since the sample data volume of all different enterprises in the full dataset varies, it is understandable that different enterprises may include enterprises of different sizes and businesses. Therefore, in this embodiment, different ratios can be selected for sample data from different industries or sample data containing different sample sizes. For example, for enterprises in the manufacturing industry, a set ratio of 50% or 60% can be selected. If the sample data volume of a particular industry is too small, 70% of the data can be selected for training. The specific set ratio needs to be flexibly adjusted based on the specific scenario.
[0055] Sample screening can be considered as a process that involves screening out anomalous data samples and labeling a sample dataset. A labeled sample dataset can refer to a sample dataset that has been confidence-labeled. In this embodiment, strong rules independent of features can be defined. Based on these strong rules, the training sample dataset is filtered to obtain positive samples. These positive samples are then confidence-labeled to obtain a labeled sample dataset. In this embodiment, only a subset that meets the strictly defined rules can be labeled, and the labeled sample dataset can account for approximately 20%.
[0056] In this embodiment, a full dataset consisting of sample data is obtained by processing the basic power data for each province and region obtained from Taichung in the power grid headquarters data during a historical period. Stratified sampling is used to extract a set proportion of data from the full dataset according to a pre-set ratio as the training sample dataset; wherein the proportion of the training sample dataset is set to θ∈[20%,80%]. The training sample dataset is then filtered for abnormal samples and the labeled sample dataset is screened to obtain a labeled sample dataset. The labeled sample dataset in this embodiment can be sample data corresponding to private economic market operators.
[0057] In this embodiment, each sample in the training sample data set can be input into the initial probability recognition model to obtain the first probability data corresponding to the sample, and the labeled sample data set can be input into the initial probability recognition model to obtain the corresponding second probability data. Then, the loss function is determined based on the first probability data and the second probability data, and the parameters of the initial probability recognition model are adjusted according to the loss function. The sample data and the labeled sample data are repeatedly input into the initial probability recognition model, and the model is iterated multiple times until a target probability recognition model that meets the requirements is obtained.
[0058] Furthermore, in this embodiment, the trained target probability recognition model can continue to be used to predict all labeled sample data sets, and the average of the labeled probabilities given by the model is used as the probability (P+) that the labeled samples are labeled; then the trained model is used to predict whether the training sample data set is labeled. For example, if the probability of the i-th sample being labeled is Pi, then the probability that the sample is actually a positive sample can be determined as: P = Pi / P+. In summary, the probability P that all samples are positive samples (private economic market operators) can be predicted, thereby obtaining the proportion of labeled samples in the full amount of training samples.
[0059] Through such a setting in this embodiment, the full sample data set can be divided into training sample data sets with different proportions, and the labeled sample data set can be obtained by screening, so that the initial probability recognition model can be trained by the training sample data set and the labeled sample data set to obtain the target probability recognition model, thereby improving the accuracy and reliability of the target probability model recognition.
[0060] In this embodiment, optionally, obtaining a power sample data set includes: obtaining original power basic data of different regions in a historical time period from a data center; performing data cleaning processing on the original power basic data to obtain processed power basic data; performing feature extraction on the processed power basic data to obtain various feature data; and constructing a power sample data set based on various feature data.
[0061] The data center can refer to the National Electric Power Network Headquarters data center, which contains various electricity data for each province and region. The historical period can refer to the existing basic electricity data for each region over the past period. In this embodiment, the historical period can be one month, one year, or two years, and can be set based on actual needs. Data cleaning can include steps such as initial deduplication, outlier annotation, null value processing, secondary deduplication, and outlier filling. Feature extraction can be the process of extracting corresponding feature types from the processed basic electricity data. In this embodiment, based on the production characteristics of private economic market operators and combining available data resources, five core feature categories were selected: monthly enterprise electricity consumption, enterprise type, regional distribution, industry distribution, and load characteristics. Within each feature category, one or more specific features can be further refined, and the specific settings can be based on actual needs. Furthermore, after extracting each feature data, this embodiment also conducts data tracing to clarify the information sources required to construct these features. This process ensures that the final required data table and its field information fully meet the requirements of feature construction.
[0062] In this embodiment, the original basic power data of different regions in the historical time period are subjected to data cleaning processing steps such as initial deduplication, outlier annotation, null value processing (filling or deletion), secondary deduplication and outlier filling to ensure that the data can be directly used to identify, calculate and establish feature standards; then, the basic power data after data cleaning is subjected to feature extraction of enterprise monthly power consumption, enterprise type, regional distribution, industry distribution and load characteristic data to obtain various feature data, and then each feature data is constructed to obtain a power sample data set. Furthermore, in this embodiment, the corresponding enterprise monthly power consumption, enterprise type, regional distribution, industry distribution and load characteristic data extracted according to timestamp information can be used as a group of sample data sets, and each group of sample data sets can be extracted according to different timestamp information, and each group of sample data sets can be aggregated to form a power sample data set.
[0063] Furthermore, in this embodiment, to traverse and verify the data quality provided by each province in the data, after obtaining the original basic power data for different regions in the data from Taichung over a historical period, data verification is performed on the data. Specific data verification steps include total statistics for key fields, null value statistics, outlier statistics, and duplicate value statistics. Based on this data verification, provincial data tables with qualified data quality are selected for further data cleaning.
[0064] Through such a setting in this embodiment, the original basic power data can be subjected to various data processing to obtain a power sample data set, thereby improving the data quality of the sample and enhancing the authenticity and reliability of the source of the sample data.
[0065] In this embodiment, optionally, the training sample data set is screened to obtain a labeled sample data set, including: performing sample filtering on the training sample data set to obtain a filtered training sample data set; and screening the filtered training sample data set according to set rules to obtain a labeled sample data set.
[0066] Sample filtering may refer to filtering out abnormal samples. A filtered training sample dataset may be understood as a dataset that does not contain abnormal samples. A set rule may refer to a pre-set rule. The set rule in this embodiment may be used to filter annotated sample datasets. In this embodiment, a strong rule unrelated to feature data may be set as a rule for filtering annotated sample datasets.
[0067] In this embodiment, the training sample data set can be filtered for abnormal samples using an anomaly detection algorithm to obtain a filtered training sample data set; the filtered training sample data set is then screened according to set rules to obtain positive sample data that meets the set rules, and the screened positive sample data is confidence-labeled to obtain a labeled sample data set.
[0068] In this embodiment, through such a setting, the training sample data set can be screened and processed to obtain a labeled sample data set with confidence labels, so as to facilitate the training of the initial probability recognition model.
[0069] In this embodiment, optionally, the training sample data set is subjected to sample filtering to obtain a filtered training sample data set, including: using an anomaly detection algorithm and a clustering algorithm to detect abnormal samples and sparse area samples in the training sample set respectively; removing the abnormal samples and sparse area samples to obtain a filtered training sample data set.
[0070] Among them, the anomaly detection algorithm can refer to any algorithm that can detect anomalies in sample data. The clustering algorithm can be any algorithm that can cluster data. Exemplarily, the anomaly detection algorithm in this embodiment can be the Isolation Forest (iForest) algorithm. The clustering algorithm can be a density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN). Abnormal sample data can refer to samples detected by the anomaly detection algorithm. Sparse area samples can be understood as samples corresponding to sparse areas obtained after clustering samples using a clustering algorithm.
[0071] In this embodiment, the Isolation Forest algorithm (n_estimators = 100, contamination = 0.01) can be used to detect outliers, and the training sample data set is filtered for abnormal samples to obtain sample data after abnormal sample filtering. Then, the sample data after abnormal sample filtering is clustered using the DBSCAN clustering algorithm (eps = 0.5, min_samples = 5), thereby removing sparse area samples and obtaining the final filtered training sample data set.
[0072] In this embodiment, through such a setting, abnormal samples and sparse area samples can be removed from the training sample data set to obtain a filtered training sample data set, thereby improving the availability and effectiveness of the sample data.
[0073] This embodiment can use the data middle platform of the big data center, and rely on the basic power data such as electricity consumption, load and capacity of all industrial and commercial users provided by the provincial companies in the headquarters distributed in the secondary middle platform to carry out modeling prediction and analysis, and analyze and predict the distribution of private economic market operators from individuals to tens of millions, so as to facilitate subsequent work based on the distribution of private economic market operators.
[0074] The technical solution of the embodiment of the present invention obtains the basic power data and enterprise identification information of each enterprise in the target area; inputs the basic power data and enterprise identification information into the target probability recognition model to obtain the probability data of the target enterprise being the set subject type; wherein, the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set; judges whether the probability data is greater than the probability threshold; if the probability data is greater than the probability threshold, the enterprise subject type corresponding to the target enterprise is a private enterprise subject type; if the probability data is less than or equal to the probability threshold, the enterprise subject type corresponding to the target enterprise is a non-private enterprise subject type. This technical solution obtains the probability of an enterprise being the set subject type through model prediction, and then determines the enterprise subject type based on the probability data, which can solve the problems of large workload, long cycle and poor effect in the existing process of determining the enterprise subject type.
[0075] Example 3
[0076] Figure 3 Schematic diagram of a device for identifying enterprise entity types according to the third embodiment of the present invention. Figure 3 As shown, the device includes:
[0077] The data acquisition module 310 is used to obtain basic power data and enterprise identification information of each enterprise in the target area;
[0078] Probabilistic identification module 320 is used to input the basic power data and enterprise identification information into the target probability identification model to obtain probability data that the target enterprise is the set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability identification model is obtained by training the initial probability identification model and the power sample data set;
[0079] The subject type determination module 330 is used to determine the target subject type corresponding to the target enterprise based on the probability data.
[0080] Optionally, the subject type determination module 330 is specifically configured to:
[0081] Determine whether the probability data is greater than the probability threshold;
[0082] If the probability data is greater than the probability threshold, the target enterprise's corresponding corporate entity type is a private enterprise entity type;
[0083] If the probability data is less than or equal to the probability threshold, the corporate entity type corresponding to the target enterprise is a non-private enterprise entity type.
[0084] Optionally, the basic power data includes power data, load data, capacity data, and distribution data;
[0085] The probability identification module 320 is specifically configured to:
[0086] The electricity data, load data, capacity data, distribution data and enterprise identification information are input into the target probability recognition model to obtain the probability data of the target enterprise being the set subject type.
[0087] Optionally, the enterprise entity type identification device may further include a target probability identification model obtained by training the following modules;
[0088] A sample data acquisition module is used to acquire a power sample data set;
[0089] A training sample extraction module is used to extract a set proportion of data from the power sample data set as a training sample data set;
[0090] The screening module is used to screen the training sample data set to obtain the labeled sample data set;
[0091] The model training module is used to train the initial probability recognition model based on the labeled sample data set and the training sample data set to obtain the target probability recognition model.
[0092] Optional, sample data acquisition module, specifically used to:
[0093] Obtain the original basic power data of different regions in the historical period from the data center;
[0094] Perform data cleaning on the original basic power data to obtain processed basic power data;
[0095] Perform feature extraction on the processed basic power data to obtain various feature data;
[0096] A power sample data set is constructed based on various feature data.
[0097] Optional screening modules include:
[0098] A sample filtering unit is used to filter the training sample data set to obtain a filtered training sample data set;
[0099] The sample screening unit is used to screen the filtered training sample data set according to the set rules to obtain the labeled sample data set.
[0100] Optional sample filtering unit, specifically used for:
[0101] Anomaly detection algorithm and clustering algorithm are used to detect abnormal samples and sparse area samples in the training sample set respectively;
[0102] Abnormal samples and sparse area samples are removed to obtain a filtered training sample dataset.
[0103] An enterprise entity type identification device provided by an embodiment of the present invention can execute an enterprise entity type identification method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.
[0104] Example 4
[0105] Figure 4 1 is a schematic diagram of the structure of an electronic device provided according to embodiment four of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0106] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0107] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0108] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the method for identifying the type of an enterprise entity.
[0109] In some embodiments, the enterprise principal type identification method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the enterprise principal type identification method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the enterprise principal type identification method in any other suitable manner (e.g., via firmware).
[0110] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0114] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0115] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0116] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0117] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for identifying the type of an enterprise entity, characterized in that: include: Obtain basic power data and enterprise identification information of each enterprise in the target area; Inputting the basic power data and the enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set; The target entity type corresponding to the target enterprise is determined according to the probability data.
2. The method according to claim 1, characterized in that Determining the target entity type corresponding to the target enterprise according to the probability data includes: Determining whether the probability data is greater than a probability threshold; If the probability data is greater than the probability threshold, the enterprise entity type corresponding to the target enterprise is a private enterprise entity type; If the probability data is less than or equal to the probability threshold, the enterprise entity type corresponding to the target enterprise is a non-private enterprise entity type.
3. The method according to claim 1, characterized in that The basic power data includes power data, load data, capacity data and distribution data; Inputting the basic power data and the enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type includes: The electricity data, load data, capacity data and distribution data and the enterprise identification information are input into a target probability recognition model to obtain probability data of the target enterprise being a set subject type.
4. The method according to claim 1, wherein The step of obtaining a target probability recognition model by training the initial probability recognition model and the power sample data set includes: Obtain a sample data set of electricity; Extracting a set proportion of data from the power sample data set as a training sample data set; Screening the training sample data set to obtain a labeled sample data set; The initial probability recognition model is trained based on the labeled sample data set and the training sample data set to obtain a target probability recognition model.
5. The method according to claim 4, characterized in that Get the power sample dataset, including: Obtain the original basic power data of different regions in the historical period from the data center; Performing data cleaning on the original basic power data to obtain processed basic power data; Performing feature extraction on the processed basic power data to obtain various feature data; A power sample data set is constructed based on the various characteristic data.
6. The method according to claim 4, characterized in that The training sample data set is screened to obtain a labeled sample data set, including: Performing sample filtering on the training sample data set to obtain a filtered training sample data set; The filtered training sample data set is screened according to the set rules to obtain the labeled sample data set.
7. The method according to claim 6, characterized in that Performing sample filtering on the training sample data set to obtain a filtered training sample data set includes: Anomaly detection algorithm and clustering algorithm are respectively used to detect abnormal samples and sparse area samples in the training sample set; The abnormal samples and the sparse area samples are removed to obtain a filtered training sample data set.
8. A device for identifying the type of an enterprise entity, characterized in that: include: The data acquisition module is used to obtain the basic power data and enterprise identification information of each enterprise in the target area; a probability recognition module, configured to input the basic power data and the enterprise identification information into a target probability recognition model to obtain probability data that the target enterprise is a set subject type; wherein the target enterprise is the enterprise corresponding to the enterprise identification information, and the target probability recognition model is obtained by training the initial probability recognition model and the power sample data set; The subject type determination module is used to determine the target subject type corresponding to the target enterprise according to the probability data.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the enterprise entity type identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the enterprise entity type identification method according to any one of claims 1 to 7 when executed.