Methods, devices and terminal devices for environmental protection enterprise user portraits

By building a target user classification model and labeling system, building an attribute probability model based on the electricity consumption big data and expert evaluation of environmental protection enterprises, the problem of failure to effectively integrate electricity consumption data and build a full life cycle evaluation system in the existing technology, and the accurate and comprehensive evaluation of user portraits of environmental protection enterprises is achieved.

CN115034289BActive Publication Date: 2025-05-27国网河北省电力有限公司营销服务中心 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210470919.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-05-27
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

The existing technology has failed to effectively integrate electricity consumption data of environmental protection enterprises and has not built a comprehensive and full life cycle evaluation system, resulting in deviations in the portrait research of environmental protection enterprises.

Method used

By building a target user classification model and a target user labeling system, building an attribute probability model based on electricity consumption big data and expert evaluation, we will obtain user portraits of environmental protection enterprises.

Benefits of technology

It has achieved all-round and full life cycle evaluation of environmental protection enterprises, improved the accuracy and reliability of the portrait, and avoided evaluation deviations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115034289B_ABST
    Figure CN115034289B_ABST
Patent Text Reader

Abstract

This application is applicable to the field of environmental protection monitoring technology, and provides a method, device and terminal device for the user portrait of environmental protection enterprises. The method includes: constructing a target user classification model, which is constructed based on the large amount of electricity consumption data of target users; constructing a target user label system, which is constructed at least based on expert evaluation, data collection and research; based on the target user classification model and the target user label system, constructing an attribute probability model to obtain a target user portrait. This application provides an evaluation system for the whole life cycle of environmental protection enterprises in all aspects by constructing a user portrait of environmental protection enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of environmental protection monitoring, and particularly relates to a method, device and terminal device for user portraits of environmental protection enterprises. Background Art

[0002] In real life, there are some phenomena that the environmental supervision system fails to maximize its role, and various systems and corresponding phenomena are not connected in series to form a visual display. Based on this, a method for constructing an environmental protection enterprise portrait to achieve precise monitoring of environmental protection enterprise emissions has emerged.

[0003] However, existing research on environmental protection enterprise portraits has not effectively integrated the electricity consumption data of enterprise users, and has not constructed an all-round and full-life cycle evaluation system for environmental protection enterprises, resulting in deviations in its research pre-evaluation. Summary of the Invention

[0004] To overcome the problems in the related art, embodiments of this application provide a method, device and terminal device for user portraits of environmental protection enterprises. By constructing a user portrait of an environmental protection enterprise, an all-round and full-life cycle evaluation system for environmental protection enterprises is provided.

[0005] This application is implemented through the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a method for user portraits of environmental protection enterprises, including: constructing a target user classification model, the target user classification model being constructed based on the electricity consumption big data of target users; constructing a target user label system, the target user label system being constructed based on at least expert evaluation, data collection and research; and constructing an attribute probability model based on the target user classification model and the target user label system, and obtaining the target user portrait by calculating the posterior probability.

[0007] In a possible implementation manner of the first aspect, the target user classification model includes constructing a target user electricity consumption big data and load clustering model, the target user electricity consumption data and load clustering model being constructed based on the K-means clustering model, including:

[0008] Randomly select K target user electricity consumption data as the number of categories for initial clustering, where K is a positive integer.

[0009] Calculate the sum of the squared distances from each target user electricity consumption data to the K initial clustering centers, and its objective function expression is:

[0010]

[0011] In the formula, K is the number of categories for initial clustering, c m is the clustering center of the m-th category of initial clustering, and x nis the power consumption data of the target user in the m-th type of initial cluster, d(c m , x n ) is the distance from the power consumption data of the target user in this cluster group to the initial cluster center, u mn represents the component of the power consumption data of the n-th target user belonging to the m-th type of initial cluster. Then J(C) is the sum of the squares of the distances from the power consumption data of each target user to the K initial cluster centers.

[0012] Based on the objective function, update the cluster center of the initial cluster. The expression is:

[0013]

[0014] In the formula, n ∈ [1, …, N], N is the number of samples of the power consumption data of the target user. When u mn = 1, it means that the power consumption data of the n-th target user belongs to the m-th type of initial cluster. When u mn = 0, it means that the power consumption data of the n-th target user does not belong to the m-th type of initial cluster.

[0015] Judge whether the objective function is less than the first threshold. If the objective function is greater than the first threshold, continue to adjust the sum of the squares of the distances from the power consumption data of the target user to the K initial cluster centers until it is less than or equal to the first threshold. If the objective function is less than or equal to the first threshold, stop the iteration and output the number of clusters and each cluster group matrix.

[0016] In a possible implementation manner of the first aspect, constructing the target user classification model further includes constructing a clustering evaluation model based on the CH index. The CH index clustering evaluation model is used to evaluate the K-means clustering model. The CH index clustering evaluation model includes calculating the sum of squared errors between cluster groups, the sum of squared errors within cluster groups, and the variance ratio.

[0017] Calculate the sum of squared errors between cluster groups. Its expression is:

[0018] Calculate the sum of squared errors within cluster groups. Its expression is:

[0019] Calculate the variance ratio. Its expression is: In the above expressions, N is the number of samples of the power consumption data of the target user, n i is the number of power consumption data of the n-th target user in the i-th class, m i is the cluster center of the i-th cluster group, m is the average value of all cluster group centers, c i is the i-th cluster group matrix, and x is the power consumption data sample in c i .

[0020] If the variance ratio is less than or equal to the second threshold, it is determined that the result of the K-means clustering model is qualified, and the construction of the target user classification model is completed. If the variance ratio is greater than the second threshold, it is determined that the result of the K-means clustering model is unqualified, and the K-means clustering model is reconstructed until the variance ratio is less than or equal to the second threshold.

[0021] In a possible implementation manner of the first aspect, the target user label system is composed of a combination of labels at one or more levels, where the labels at multiple levels have a subordination relationship.

[0022] In a possible implementation manner of the first aspect, the target user label system includes at least one first-level label; the first-level label is at least one of the qualification level, low-carbon emission reduction, and credit behavior labels. The first-level label includes at least one second-level label, and the second-level label is subordinate to the first-level label, and one second-level label is subordinate to only one first-level label. The qualification level includes at least the second-level labels: basic information, report certificates, production manufacturing, and personnel and plant. Low-carbon emission reduction includes at least the second-level labels: low-carbon economy, low-carbon consumption, and low-carbon environment. Credit behavior includes at least the second-level labels: credit information, statement information, and environmental protection information of the target user enterprise.

[0023] In a possible implementation manner of the first aspect, the second-level label includes at least one third-level label, and the third-level label is subordinate to the second-level label, and one third-level label is subordinate to only one second-level label. Basic information includes at least the third-level labels: unit type, registered capital, net output value, registered location, and total funds. Report certificates include at least the third-level label: type and quantity of certificates. Production manufacturing includes at least the third-level labels: number of test equipment, purchase unit price of test equipment, number of production equipment, and purchase unit price of production equipment. Personnel and plant includes at least the third-level labels: total number of employees, number of management personnel, number of technical personnel, number of production personnel, location of the factory area, and total area of the factory area. Low-carbon economy includes at least the third-level labels: output value, output value growth rate, enterprise profit, per capita output value, carbon productivity, proportion of low-carbon technology investment, and proportion of low-carbon environmental protection investment. Low-carbon energy consumption includes at least the third-level labels: carbon emissions, energy conversion rate, non-coal energy ratio, energy carbon emission coefficient, and energy intensity. Low-carbon environment includes at least the third-level labels: pollutant emission intensity, pollutant control coefficient, green coverage rate of the factory area, effective utilization rate of waste, carbon emissions per unit output value, and pollution emissions per unit output value. Credit information includes at least the third-level labels: current liabilities and repaid debts. Statement information includes at least the third-level labels: information subject statement, reporting agency description, and credit reference center annotation. Enterprise environmental protection information includes at least the third-level labels: enterprise illegal information, punishment and execution information, environmental protection approval information, environmental protection certification information, and clean production audit information.

[0024] In a possible implementation of the first aspect, based on the target user classification model and the target user label system, an attribute probability model is constructed to obtain the target user portrait, including: constructing an attribute probability model based on each clustering group matrix in the target user classification model and the third-level labels in the target user label system, where the attribute probability model is constructed based on the naive Bayes method, including:

[0025] Set the third-level label attribute set corresponding to the k-th clustering group as a k , then a k =(a k,1 , a k,2 , …, a k,i , …, a k,T ), i ∈ [1, …, T], T is the number of third-level labels, k ∈ [1, …, K], K is the number of clustering categories;

[0026] According to the situation of the i-th third-level label of all users in the k-th clustering group, the i-th third-level label a k,i of the k-th clustering group can be divided into q categories, that is, the feature set b k,i =(b k,i,1 , b k,i,2 , …, b k,i,j , …, b k,i,q )(j = 1, …, q;

[0027] Using the naive Bayes formula, calculate the posterior probability, and the expression is:

[0028]

[0029] In the formula, the category j corresponding to the largest probability among P(b k,i,1 |a k,i ), P(b k,i,2 |a k,i ), ……, and P(b k,i,q |a k,i ) is used as the division result of the i-th third-level label a k,i of the k-th category of target user electricity consumption data. Integrate the division results of T third-level labels to form the complete portrait of the k-th clustering group.

[0030] In the second aspect, an environmental protection enterprise user portrait device provided by an embodiment of the present application includes: a classification model construction module for constructing a target user classification model, where the target user classification model is based on the electricity consumption big data of the target user; a label system construction module for constructing a target user label system; the target user label system is at least based on expert evaluation, data collection, and investigation and research; an attribute probability construction module for constructing an attribute probability model based on the target user classification model and the target user label system to obtain the target user portrait.

[0031] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for the environmental protection enterprise user portrait as described in any item of the first aspect.

[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method for the environmental protection enterprise user portrait as described in any item of the first aspect.

[0033] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, it causes the terminal device to execute the method for the environmental protection enterprise user portrait as described in any item of the first aspect above.

[0034] It can be understood that the beneficial effects of the above second aspect to fifth aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0035] In the embodiment of the present application, a target user classification model and a target user system label are constructed, and then an attribute probability model is constructed based on the target user classification model and the target user system label. By calculating the posterior probability, an environmental protection enterprise user portrait is obtained. Compared with the prior art, the present application provides an evaluation system for the whole life cycle of environmental protection enterprises through constructing the environmental protection enterprise user portrait.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a schematic flowchart of the method for the environmental protection enterprise user portrait provided by an embodiment of the present application;

[0039] Figure 2 is a schematic framework diagram of the target user label system provided by an embodiment of the present application;

[0040] Figure 3 is a schematic structural diagram of the environmental protection enterprise user portrait device provided by an embodiment of the present application;

[0041] Figure 4It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0042] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0043] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0044] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0045] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detecting [the described condition or event]" or "in response to detecting [the described condition or event]" according to the context.

[0046] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0047] Referring to "one embodiment" or "some embodiments" described in the specification of the present application means that specific features, structures, or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0048] With the increasing urgency of accurately detecting the emissions of environmental protection enterprises, some ecological environment bureaus and local power companies are exploring the "power big data + environmental protection" cooperation model. Relying on the power big data platform, they achieve the sharing of power big data and monitoring data, and combine the long-term on-site inspection experience of the environmental protection department and the analysis of illegal data samples to optimize the environmental supervision model. Among them, Sun Yi et al. published the paper "Comprehensive Clustering Method for User Load Characteristics and Adjustable Potential Oriented to Power Big Data" in the 18th issue of Volume 41 of the Chinese Society for Electrical Engineering in 2021, achieving accurate clustering considering the multi-dimensional influencing factors of user electricity consumption behavior; Cheng Han et al. published the paper "Evaluation of Agricultural Land Suitability Based on RS, AHP, and MEA: A Case Study in Jilin Province, China" in the 4th issue of Volume 11 of Agriculture in 2021, using the analytic hierarchy process to solve the problem of multi-criteria for supplier optimization. However, existing research has not effectively integrated the electricity consumption data and environmental protection data of enterprise users, and has not constructed an all-round and full-life cycle evaluation system for environmental protection enterprises, which will lead to deviations in their evaluation. With the large access of massive environmental protection and power data, how to achieve an accurate portrait of environmental protection enterprise users requires studying the electricity consumption types of environmental protection enterprises and establishing a corresponding label system.

[0049] Based on the above problems, in the embodiments of the present application, a method for depicting environmental protection enterprises is provided, that is: by constructing a target user classification model and a target user system label, and then based on the target user classification model and the target user label system, constructing an attribute probability model, and obtaining the portrait of environmental protection enterprise users by calculating the posterior probability.

[0050] As Figure 1 shown, the schematic flow chart of the method for depicting environmental protection enterprise users provided by an embodiment of the present application is referred to Figure 1 , and the details of this method are as follows:

[0051] In step 10, a target user classification model is constructed.

[0052] The target user classification model is constructed based on the electricity big data of the target user. The electricity big data not only includes power data such as the electricity consumption period, electricity consumption duration, and electricity consumption of the target user, but also combines with objective factors such as the climate environment and geographical conditions of the local area, so as to more objectively and comprehensively monitor the electricity consumption of the target user, and the production capacity of the target user can also be inferred based on the electricity big data. Moreover, by comparing the emission information of local polluting enterprises with the production capacity of the target enterprise, it helps the ecological environment departments at all levels to obtain the latest pollution prevention and control information more timely.

[0053] In a possible implementation, the target user classification model is constructed based on the big data of the target user's electricity consumption and the load clustering model. Various methods can be used to establish the clustering model, such as the K-means clustering algorithm, the hierarchical clustering algorithm, the DBSCAN algorithm, etc. In this application, the K-means clustering model is taken as an example for specific illustration, but the construction of other clustering model methods will not limit the solution of this application.

[0054] In a possible implementation, a K-means clustering model is constructed based on the electricity consumption data of the target user. The specific method includes:

[0055] Randomly select K target user electricity consumption data as the number of categories for initial clustering, where K is a positive integer;

[0056] Calculate the sum of the squares of the distances from each target user electricity consumption data to the K initial clustering centers respectively, and classify the target user electricity consumption data sample with the smallest sum of the squares of the distances into this clustering group. The specific objective function expression is as follows:

[0057]

[0058] In the formula, K is the number of categories for initial clustering, c m is the clustering center of the m-th initial clustering, x n is the target user electricity consumption data in the m-th initial clustering, d(c m , x n ) is the distance from the target user electricity consumption data in this clustering group to this initial clustering center, u mn represents the component of the n-th target user electricity consumption data belonging to the m-th initial clustering, then J(C) is the sum of the squares of the distances from each target user electricity consumption data to the K initial clustering centers.

[0059] Based on the objective function, update the clustering center of the initial clustering. The expression is:

[0060]

[0061] In the formula, n ∈ [1,..., N], N is the number of target user electricity consumption data samples. When u mn = 1, it means that the n-th target user electricity consumption data belongs to the m-th initial clustering. When u mn = 0, it means that the n-th target user electricity consumption data does not belong to the m-th initial clustering.

[0062] Judge whether the objective function is less than the first threshold. If the objective function is greater than the first threshold, continue to adjust the sum of the squares of the distances from the target user electricity consumption data to the K initial clustering centers until it is less than or equal to the first threshold; if the objective function is less than or equal to the first threshold, stop the iteration and output the number of clusters and the matrix of each clustering group.

[0063] In a possible implementation, to better cluster the power consumption data of target users and ensure the accuracy of the user portraits of environmental protection enterprises, this application also introduces the CH index (Calinski-Harabasz Index) to evaluate the accuracy of the above K-means clustering model.

[0064] Exemplarily, the CH index clustering evaluation model includes the sum of squared errors between clusters SSB, the sum of squared errors within clusters SSW, and the variance ratio VRC k , and the specific method is as follows:

[0065] Calculate the sum of squared errors between clusters SSB, and the expression is:

[0066]

[0067] In the formula, n i is the number of power consumption data of target users in the i-th category, m i is the clustering center of the n-th i-th clustering group, and m is the average value of all clustering group centers.

[0068] Calculate the sum of squared errors within clusters SSW, and the expression is:

[0069]

[0070] In the formula, c i is the matrix of the i-th clustering group, and x is the power consumption data sample of the target user in c i .

[0071] It should be noted that the calculation order of the sum of squared errors between clusters SSB and the sum of squared errors within clusters SSW is not in a specific order.

[0072] According to the calculated SSB and SSW, substitute them into the variance ratio VRC k formula, and the expression is:

[0073]

[0074] In the formula, when the value of VRC k is smaller, it indicates that the within-class distance of a single clustering group is smaller, and the between-class distance of the clustering group is larger, indicating a better clustering effect. Therefore, according to actual needs and experience summary, a second threshold is set.

[0075] If the variance ratio value VRC k is less than or equal to the second threshold, it is determined that the result of the K-means clustering model is qualified, and the target user classification model construction is completed.

[0076] If the variance ratio value VRC kIf it is greater than the second threshold, it is determined that the result of the K-means clustering model is unqualified, and the K-means clustering model needs to be reconstructed until the variance ratio VRC k is less than or equal to the second threshold.

[0077] In this application, by first establishing a K-means clustering model based on the power consumption big data of the target users to cluster the target users, initially outlining the prototype of the user portrait of environmental protection enterprises, and then introducing a CH index clustering evaluation model to evaluate the outlined prototype of the user portrait and re-clustering the unqualified clustering models, the accuracy of the user portrait of environmental protection enterprises can be well guaranteed.

[0078] If only the power consumption big data of the target users is taken as the sole consideration and the actual situation of the target users themselves is ignored, then the accurate monitoring and supervision of environmental protection enterprises cannot be achieved, let alone the establishment of a scientific evaluation system for environmental protection enterprises. Therefore, in order to accurately evaluate environmental protection enterprises in all aspects and throughout the life cycle and establish an accurate user portrait, it is also necessary to collect the actual situation of the target users themselves and establish a perfect label system.

[0079] In step 20, a target user label system is constructed.

[0080] In a possible implementation manner, to construct a target user label system, at least means such as expert evaluation, data collection, and investigation and research on the target users are required to be able to comprehensively understand the target users. Among them, the target user label system can be composed of one or more levels of label combinations, and there is a subordination relationship between multiple levels of labels. Refer to Figure 2 , the framework schematic diagram of the target user label system provided by an embodiment of this application.

[0081] Exemplarily, in Figure 2 , this application divides the target user label system into 3 levels, and there is a subordination relationship between the labels of different levels. At the same time, each level is further divided into several labels.

[0082] Exemplarily, the first level of the target user label system includes at least 3 first-level labels, namely the qualification level label, the low-carbon emission reduction label, and the credit behavior label.

[0083] Exemplarily, the second level of the target user label system includes at least 10 second-level labels, and different second-level labels are uniquely subordinate to one first-level label.

[0084] Optionally, the qualification level includes at least second-level labels: basic information, report certificates, production manufacturing, and personnel and workshops.

[0085] Optionally, low-carbon emission reduction at least includes secondary labels: low-carbon economy, low-carbon consumption, and low-carbon environment.

[0086] Optionally, credit reporting behavior at least includes secondary labels: credit information, statement information, and environmental protection information of the target user enterprise.

[0087] Exemplarily, the third level of the target user label system contains at least 44 tertiary labels, and different tertiary labels are uniquely affiliated with one secondary label.

[0088] Optionally, basic information at least includes tertiary labels: unit type, registered capital, net output value, place of registration, and total funds.

[0089] Optionally, report certificates at least include tertiary labels: certificate type and quantity.

[0090] Optionally, production and manufacturing at least include tertiary labels: number of test equipment, purchase unit price of test equipment, number of production equipment, and purchase unit price of production equipment.

[0091] Optionally, personnel and factory buildings at least include tertiary labels: total number of employees, number of management personnel, number of technical personnel, number of production personnel, location of the factory area, and total area of the factory area.

[0092] Optionally, low-carbon economy at least includes tertiary labels: output value, output value growth rate, enterprise profit, per capita output value, carbon productivity, proportion of low-carbon technology investment, and proportion of low-carbon environmental protection investment.

[0093] Optionally, low-carbon energy consumption at least includes tertiary labels: carbon emissions, energy conversion rate, non-coal energy ratio, energy carbon emission coefficient, and energy intensity.

[0094] Optionally, low-carbon environment at least includes tertiary labels: pollutant emission intensity, pollutant control coefficient, green coverage rate of the factory area, effective utilization rate of waste, carbon emissions per unit output value, and pollution emissions per unit output value.

[0095] Optionally, credit information at least includes tertiary labels: current liabilities and repaid debts.

[0096] Optionally, statement information at least includes tertiary labels: information subject statement, reporting agency description, and credit reference center annotation.

[0097] Optionally, enterprise environmental protection information at least includes tertiary labels: enterprise illegal information, penalty and execution information, environmental protection approval information, environmental protection certification information, and clean production audit information.

[0098] It should be specially noted that Figure 2The target users shown in [figure], as well as the three levels, three first-level tags, ten second-level tags, and forty-four third-level tags in the system are only for better elaborating the technical solution of this application and do not limit this application.

[0099] It should also be specifically noted that this application does not limit the sequence of steps 10 and 20. Step 10 can be executed first, or step 20 can be executed first.

[0100] By establishing a target user tag system, this application can further understand the basic situations related to the operations, pollutant emissions, and enterprise electrical equipment of target users, providing a reliable and accurate basis for establishing a full-life-cycle evaluation system for target users and accurately depicting the user portraits of environmental protection enterprises.

[0101] In step 30, an attribute probability model is constructed based on the target user classification model and the target user system tags.

[0102] In a possible implementation manner, an attribute probability model is constructed based on each clustering group matrix in the target user classification model and the third-level tags in the target user tag system.

[0103] Exemplarily, this application uses the Naive Bayes method to construct an attribute probability model, and the specific construction process is as follows:

[0104] First, set the third-level tag attribute set corresponding to the k-th clustering group as a k , then a k = (a k,1 , a k,2 , …, a k,i , …, a k,T ), i ∈ [1, …, T], T is the number of third-level tags, k ∈ [1, …, K], and K is the number of clustering categories;

[0105] Then, according to the situation of the i-th third-level tag of all users in the k-th clustering group, the i-th third-level tag a k,i of the k-th clustering group can be divided into q categories, that is, the feature set b k,i = (b k,i,1 , b k,i,2 , …, b k,i,j , …, b k,i,q ) (j = 1, …, q);

[0106] Finally, use the Naive Bayes formula to calculate the posterior probability, and the expression is:

[0107]

[0108] In the formula, P(b k,i,1 | a k,i ), P(bk,i,2 |a k,i )..., and P(b k,i,q |a k,i ) The category j corresponding to the maximum probability among them is used as the i-th third-level label a of the power consumption data of the k-th type of target user k,i The partitioning result. Integrate the partitioning results of T third-level labels to form a complete portrait of the k-th clustering group.

[0109] In this step, by using the target user classification model and the target user label system to construct an attribute probability model based on the Naive Bayes method, it can very accurately outline the user portrait of environmental protection enterprises, enabling the ecological environment department, law enforcement officers, and environmental protection enterprises to obtain information on the pollution prevention and control level of crimes in a timely manner, and constructing an all-round and full-life-cycle evaluation system for environmental protection enterprises.

[0110] Corresponding to the method for the user portrait of environmental protection enterprises described in the above embodiments Figure 3 The structural block diagram of the environmental protection enterprise user portrait device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0111] See Figure 3 , the environmental protection enterprise user portrait device in the embodiment of the present application may include: a classification model construction module 201, a label system construction module 202, and an attribute probability construction module 203.

[0112] The classification model construction module 201 is used to construct a target user classification model, and the target user classification model is constructed based on the power consumption big data of the target user.

[0113] In some embodiments, the classification model construction module 201 is further used to construct a target user power consumption big data and load clustering model, and the target user power consumption data and load clustering model is constructed based on the K-means clustering model, including:

[0114] Randomly select K target user power consumption data as the number of categories for initial clustering, where K is a positive integer.

[0115] Calculate the sum of the squares of the distances from each target user power consumption data to the K initial clustering centers respectively, and its objective function expression is:

[0116]

[0117] In the formula, K is the number of categories for initial clustering, c m is the clustering center of the m-th type of initial clustering, x n is the target user power consumption data in the m-th type of initial clustering, d(c m , x n) is the distance from the electricity consumption data of the target user in this clustering group to the initial clustering center, u mn represents the component of the electricity consumption data of the nth target user belonging to the mth initial clustering, then J(C) is the sum of the squares of the distances from the electricity consumption data of each target user to the K initial clustering centers.

[0118] Based on the objective function, update the clustering center of the initial clustering, and the expression is:

[0119]

[0120] In the formula, n ∈ [1,…, N], N is the number of samples of the electricity consumption data of the target user. When u mn = 1 indicates that the electricity consumption data of the nth target user belongs to the mth initial clustering. When u mn = 0 indicates that the electricity consumption data of the nth target user does not belong to the mth initial clustering.

[0121] Determine whether the objective function is less than the first threshold. If the objective function is greater than the first threshold, continue to adjust the sum of the squares of the distances from the electricity consumption data of the target user to the K initial clustering centers until it is less than or equal to the first threshold; if the objective function is less than or equal to the first threshold, stop the iteration and output the number of clusters and the matrix of each clustering group.

[0122] In some embodiments, the classification model construction module 201 is further configured to build a CH-index clustering evaluation model, and the CH-index clustering evaluation model is used to evaluate the K-means clustering model; the CH-index clustering evaluation model includes calculating the sum of squared errors between clustering groups, the sum of squared errors within clustering groups, and the variance ratio.

[0123] Calculate the sum of squared errors between clustering groups, and its expression is:

[0124] Calculate the sum of squared errors within clustering groups, and its expression is:

[0125] Calculate the variance ratio, and its expression is: In the above expressions, N is the number of samples of the electricity consumption data of the target user, n i is the number of electricity consumption data of the nth target user in the i-th class, m i is the clustering center of the i-th clustering group, m is the average value of all clustering group centers, c i is the matrix of the i-th clustering group, and x is the sample of the electricity consumption data of the target user in c i in.

[0126] If the variance ratio is less than or equal to the second threshold, it is determined that the result of the K-means clustering model is qualified, and the construction of the target user classification model is completed. If the variance ratio is greater than the second threshold, it is determined that the result of the K-means clustering model is unqualified, and the K-means clustering model is reconstructed until the variance ratio is less than or equal to the second threshold.

[0127] The label system construction module 202 is used to construct a target user label system, and the target user label system is constructed at least based on expert evaluation, data collection, and research.

[0128] In some embodiments, the label system construction module 202 is further used to construct a target user label system that includes at least one first-level label; the first-level label is at least one of the qualification level, low-carbon emission reduction, and credit behavior labels. The first-level label includes at least one second-level label, the second-level label belongs to the first-level label, and one second-level label belongs to only one first-level label. The qualification level includes at least the second-level labels: basic information, report certificates, production manufacturing, and personnel and plant. Low-carbon emission reduction includes at least the second-level labels: low-carbon economy, low-carbon consumption, and low-carbon environment. Credit behavior includes at least the second-level labels: credit information, statement information, and environmental protection information of the target user enterprise.

[0129] In some embodiments, the label system construction module 202 is further used to construct a target user label system in which the second-level label includes at least one third-level label, the third-level label belongs to the second-level label, and one third-level label belongs to only one second-level label. Basic information includes at least the third-level labels: unit type, registered capital, net output value, registered location, and total funds. Report certificates include at least the third-level label: type and quantity of certificates. Production manufacturing includes at least the third-level labels: number of test equipment, purchase unit price of test equipment, number of production equipment, and purchase unit price of production equipment. Personnel and plant includes at least the third-level labels: total number of employees, number of management personnel, number of technical personnel, number of production personnel, location of the factory area, and total area of the factory area. Low-carbon economy includes at least the third-level labels: output value, output value growth rate, enterprise profit, per capita output value, carbon productivity, proportion of low-carbon technology investment, and proportion of low-carbon environmental protection investment. Low-carbon energy consumption includes at least the third-level labels: carbon emissions, energy conversion rate, non-coal energy ratio, energy carbon emission coefficient, and energy intensity. Low-carbon environment includes at least the third-level labels: pollutant emission intensity, pollutant control coefficient, green coverage rate of the factory area, effective utilization rate of waste, carbon emissions per unit output value, and pollution emissions per unit output value. Credit information includes at least the third-level labels: current liabilities and repaid debts. Statement information includes at least the third-level labels: information subject statement, reporting agency description, and credit reference center annotation. Enterprise environmental protection information includes at least the third-level labels: enterprise illegal information, penalty and execution information, environmental protection approval information, environmental protection certification information, and clean production audit information.

[0130] The attribute probability construction module 203 is used to construct an attribute probability model based on the target user classification model and the target user label system, and obtain a target user profile, including: constructing an attribute probability model based on each clustering group matrix in the target user classification model and the third-level labels in the target user label system, where the attribute probability model is constructed based on the naive Bayes method.

[0131] In some embodiments, the attribute probability construction module 203 is further used for the construction of the naive Bayes method, including:

[0132] Set the third-level label attribute set corresponding to the k-th clustering group as a k , then a k =(a k,1 , a k,2 , …, a k,i , …, a k,T ), i ∈ [1, …, T], T is the number of third-level labels, k ∈ [1, …, K], K is the number of clustering categories;

[0133] According to the situation of the i-th third-level label of all users in the k-th clustering group, the i-th third-level label a k,i of the k-th clustering group can be divided into q categories, that is, the feature set b k,i =(b k,i,1 , b k,i,2 , …, b k,i,j , …, b k,i,q )(j = 1, …, q;

[0134] Adopt the naive Bayes formula to calculate the posterior probability, and the expression is:

[0135]

[0136] In the formula, the category j corresponding to the largest probability among P(b k,i,1 |a k,i ), P(b k,i,2 |a k,i ), ……, and P(b k,i,q |a k,i ) is used as the division result of the i-th third-level label a k,i of the k-th category of target user electricity consumption data. Integrate the division results of T third-level labels to form a complete profile of the k-th clustering group.

[0137] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0138] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0139] An embodiment of this application also provides a terminal device. Refer to Figure 4 , the terminal device 300 may include: at least one processor 310, a memory 320, and a computer program stored in the memory 320 and executable on the at least one processor 310. When the processor 310 executes the computer program, it implements the steps in any of the foregoing method embodiments, such as Figure 1 the steps 10 to 30 in the illustrated embodiment. Alternatively, when the processor 310 executes the computer program, it implements the functions of each module / unit in the foregoing device embodiments, such as Figure 3 the functions of the illustrated modules 201 to 203.

[0140] Exemplarily, the computer program can be divided into one or more modules / units. One or more modules / units are stored in the memory 320 and executed by the processor 310 to complete this application. The one or more modules / units can be a series of computer program segments capable of performing specific functions, and these program segments are used to describe the execution process of the computer program in the terminal device 300.

[0141] Those skilled in the art can understand that Figure 4 this is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or combine some components, or different components, such as input / output devices, network access devices, buses, etc.

[0142] The processor 310 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0143] The memory 320 may be an internal storage unit of the terminal device or an external storage device of the terminal device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 320 is used to store the computer program and other programs and data required by the terminal device. The memory 320 may also be used to temporarily store data that has been output or is to be output.

[0144] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0145] The method for providing an environmental protection enterprise user portrait according to the embodiments of the present application can be applied to terminal devices such as computers, wearable devices, in-vehicle devices, tablet computers, laptop computers, netbooks, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, mobile phones, etc. The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.

[0146] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in each of the above-mentioned embodiments of the environmental protection enterprise user portrait method can be implemented.

[0147] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal is enabled to execute the steps in each of the above-mentioned embodiments of the environmental protection enterprise user portrait method.

[0148] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.

[0149] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0150] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0151] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0152] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application and should all be included in the protection scope of the present application.

Claims

1. A method for user profiling of environmental protection enterprises, characterized in that, it includes: Constructing a target user classification model, which is constructed based on the large electricity consumption data of target users; Constructing a target user label system, which is constructed based on at least expert evaluation, data collection and research; Based on the target user classification model and the target user label system, constructing an attribute probability model, and obtaining a target user profile by calculating the posterior probability; Among them, the constructing an attribute probability model based on the target user classification model and the target user label system to obtain a target user profile includes: Based on each cluster group matrix in the target user classification model and the third-level labels in the target user label system, constructing the attribute probability model; The attribute probability model is constructed based on the naive Bayes method and includes: Set the three - level label attribute set corresponding to the k - th clustering group as a k , then a k =(a k,1 , a k,2 , …, a k,i , …, a k,T ), i ∈ [1, …, T], T is the number of three - level labels, k ∈ [1, …, K], K is the number of clustering categories; According to the situation of the i-th third-level tag of all users in the k-th clustering group, the i-th third-level tag a of the k-th clustering group can be k,i divided into q categories, that is, the feature set b k,i =(b k,i,1 , b k,i,2 , …, b k,i,j , …, b k,i,q )(j = 1, …, q); Using the naive Bayes formula to calculate the posterior probability, and the expression is: wherein, P(b k,i,1 |a k,i ), P(b k,i,2 |a k,i ), ……, and the category j corresponding to the maximum probability among P(b k,i,q |a k,i ) is used as the division result of the i-th three-level label a k,i of the power consumption data of the target users in the k-th category; the division results of T three-level labels are integrated to form a complete portrait of the k-th clustering group.

2. The method for user profiling of environmental protection enterprises according to claim 1, characterized in that, The target user classification model includes constructing a target user electricity consumption data and load clustering model, and the target user electricity consumption data and load clustering model is constructed based on the K-means clustering model, and includes: Randomly selecting K target user electricity consumption data as the number of categories for initial clustering, where K is a positive integer; Calculating the sum of the squares of the distances from each target user electricity consumption data to the K initial cluster centers respectively, and its objective function expression is: where K is the number of categories in the initial clustering, and c m is the clustering center of the m-th initial clustering, and x n is the electricity consumption data of the target user in the m-th initial clustering. d(c m , x n ) is the distance from the electricity consumption data of the target user in this clustering group to this initial clustering center. u mn represents the component of the electricity consumption data of the n-th target user belonging to the m-th initial clustering. Then J(C) is the sum of the squares of the distances from each of the electricity consumption data of the target users to the K initial clustering centers; Based on the objective function, updating the cluster centers of the initial clustering, and the expression is: where \(n\in[1,\ldots,N]\), \(N\) is the number of target user electricity consumption data samples, when \(u\) mn \( = 1\) indicates that the electricity consumption data of the \(n\)th target user belongs to the \(m\)th initial cluster, when \(u\) mn \( = 0\) indicates that the electricity consumption data of the \(n\)th target user does not belong to the \(m\)th initial cluster; Judging whether the objective function is less than a first threshold. If the objective function is greater than the first threshold, then continue to adjust the sum of the squares of the distances from the target user electricity consumption data to the K initial cluster centers until it is less than or equal to the first threshold; if the objective function is less than or equal to the first threshold, then stop the iteration and output the number of clusters and each cluster group matrix.

3. The method for user profiling of environmental protection enterprises according to claim 2, characterized in that, Constructing the target user classification model further includes constructing a CH index clustering evaluation model, and the CH index clustering evaluation model is used to evaluate the K-means clustering model; the CH index clustering evaluation model includes calculating the sum of squared errors between cluster groups, the sum of squared errors within cluster groups and the variance ratio; Calculating the sum of squared errors between cluster groups, and its expression is: Calculating the sum of squared errors within cluster groups, and its expression is: Calculating the variance ratio, and its expression is: Where N is the number of target user power consumption data samples, and n i is the power consumption data quantity of the nth target user of the ith type, and m i is the clustering center of the ith clustering group, and m is the average value of all clustering group centers, and c i is the matrix of the ith clustering group, and x is the target user power consumption data sample in c i ; If the variance ratio is less than or equal to a second threshold, it is determined that the result of the K-means clustering model is qualified, and the construction of the target user classification model is completed; If the variance ratio is greater than the second threshold, it is determined that the result of the K-means clustering model is unqualified, and the K-means clustering model is reconstructed until the variance ratio is less than or equal to the second threshold.

4. The method for user profiling of environmental protection enterprises according to claim 1, characterized in that, The target user label system is composed of a combination of labels at one or more levels, and among them, the labels at multiple levels have a subordinate relationship.

5. The method for the environmental protection enterprise user portrait according to claim 4, characterized in that, the target user tag system at least includes one first-level tag; the first-level tag is at least one of the qualification level, low-carbon emission reduction, and credit behavior tags; the first-level tag at least includes one second-level tag, the second-level tag belongs to the first-level tag, and one second-level tag belongs to only one first-level tag; the qualification level at least includes the second-level tags: basic information, report certificates, production manufacturing, and personnel and plant; the low-carbon emission reduction at least includes the second-level tags: low-carbon economy, low-carbon consumption, and low-carbon environment; the credit behavior at least includes the second-level tags: credit information, statement information, and target user enterprise environmental protection information.

6. The method for the environmental protection enterprise user portrait according to claim 5, characterized in that, the second-level tag at least includes one third-level tag, the third-level tag belongs to the second-level tag, and one third-level tag belongs to only one second-level tag; the basic information at least includes the third-level tags: unit type, registered capital, net output value, registered location, and total funds; the report certificates at least include the third-level tag: type and quantity of certificates; the production manufacturing at least includes the third-level tags: number of test equipment, purchase unit price of test equipment, number of production equipment, and purchase unit price of production equipment; the personnel and plant at least includes the third-level tags: total number of employees, number of management personnel, number of technical personnel, number of production personnel, location of the factory area, and total area of the factory area; the low-carbon economy at least includes the third-level tags: output value, output value growth rate, enterprise profit, per capita output value, carbon productivity, proportion of low-carbon technology investment, and proportion of low-carbon environmental protection investment; the low-carbon energy consumption at least includes the third-level tags: carbon emissions, energy conversion rate, non-coal energy ratio, energy carbon emission coefficient, and energy intensity; the low-carbon environment at least includes the third-level tags: pollutant emission intensity, pollutant control coefficient, green coverage rate of the factory area, effective utilization rate of waste, carbon emissions per unit output value, and pollution emissions per unit output value; the credit information at least includes the third-level tags: current liabilities and repaid debts; the statement information at least includes the third-level tags: information subject statement, reporting agency description, and credit reference center annotation; the enterprise environmental protection information at least includes the third-level tags: enterprise illegal information, penalty and execution information, environmental protection approval information, environmental protection certification information, and clean production audit information.

7. An environmental protection enterprise user portrait device, characterized in that, comprising: a classification model construction module for constructing a target user classification model, the target user classification model being constructed based on the electricity consumption big data of the target user; a tag system construction module for constructing a target user tag system, the target user tag system being constructed at least based on expert evaluation, data collection, and investigation and research; an attribute probability construction module for constructing an attribute probability model based on the target user classification model and the target user tag system to obtain a target user portrait; The attribute probability construction module is further configured to construct the attribute probability model based on each clustering group matrix in the target user classification model and the three-level labels in the target user label system; and set the set of tertiary label attributes corresponding to the k-th clustering group to a k , then a k =(a k,1 , a k,2 , …, a k,i , …, a k,T ), i ∈ [1, …, T], where T is the number of tertiary labels, k ∈ [1, …, K], and K is the number of clustering categories; According to the situation of the i-th third-level tag of all users in the k-th clustering group, the i-th third-level tag a of the k-th clustering group can be k,i divided into q categories, that is, the feature set b k,i =(b k,i,1 , b k,i,2 , …, b k,i,j , …, b k,i,q )(j = 1, …, q); Using the Naive Bayes formula, calculate the posterior probability, and the expression is: wherein, P(b k,i,1 |a k,i ), P(b k,i,2 |a k,i ), ……, and the category j corresponding to the maximum probability among P(b k,i,q |a k,i ) is used as the division result of the i-th three-level label a k,i of the electricity consumption data of the k-th category of target users; the division results of T three-level labels are integrated to form a complete portrait of the k-th clustering group.

8. A terminal device, including a memory and a processor, where a computer program that can run on the processor is stored in the memory, characterized in that, when the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, characterized in that, when the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • User interest portraying method and related equipment

    CN111552865A

  • Enterprise perspective portrait method based on energy big data and terminal

    CN111861262A