Method and system for power marketing service event risk hierarchical identification based on knowledge graph semantic modeling
By constructing a knowledge graph and hierarchical semantic modeling, the problems of data silos and semantic understanding in power marketing services have been solved, enabling precise and intelligent risk identification of power marketing service events, improving the efficiency and accuracy of risk identification, and adapting to new business scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MARKETING SERVICE CENT OF STATE GRID GANSU ELECTRIC POWER CO
- Filing Date
- 2025-10-20
- Publication Date
- 2026-07-03
AI Technical Summary
The electricity marketing service suffers from problems such as data silos, low semantic understanding accuracy, poor scenario adaptability, and high reliance on manual labor, resulting in low efficiency and insufficient accuracy in risk identification, making it difficult to cope with diversified customer needs and new business scenarios.
A knowledge graph-based risk stratification identification method for power marketing service events is constructed. By integrating multi-source data and employing hierarchical semantic modeling and dynamic optimization techniques, it achieves comprehensive data coverage, accurate feature extraction, and efficient risk identification.
It has enabled precise and intelligent risk identification of power marketing service events, improved data correlation coverage, risk feature extraction accuracy and model adaptability, reduced manual intervention costs, and improved customer satisfaction and service efficiency.
Smart Images

Figure CN121390877B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electricity marketing services, specifically involving a method and system for risk hierarchical identification of electricity marketing service events based on knowledge graph semantic modeling. It is particularly suitable for risk identification of electricity marketing service events during new business scenarios such as electric vehicle charging and live electricity consumption, as well as during peak electricity consumption periods. Background Technology
[0002] With the deep penetration of artificial intelligence technology into various industries, the power industry is accelerating its digital transformation to cope with the increasingly diversified customer demands and fierce market competition. In the field of power marketing services, the traditional model of relying on manual processing of massive service data is no longer sufficient to meet the requirements of service quality and efficiency. Specifically, this manifests in two ways: First, power marketing services involve multi-source heterogeneous data such as basic customer information, 95598 complaint records, power supply fault repair, electric vehicle charging, and electricity allocation during the heating and cooling seasons. The existing manual "case-by-case" processing method cannot achieve systematic integration and in-depth analysis of the data, resulting in slow response speed and low accuracy in handling customer requests. Second, the current lack of a comprehensive marketing event classification system and efficient risk identification technology makes it difficult to accurately capture potential risks from massive customer interaction data, posing a risk that general requests may escalate into 95598 complaints, spill over to the 12398 regulatory agency, or trigger negative public opinion.
[0003] Although big data and natural language processing technologies have been initially applied in the field of power customer service, and the industry has promoted the implementation of a large-scale semantic model for customer service, current power marketing service event risk identification technology still faces many problems: First, data silos are serious, with basic customer information, 95598 request records, and data from new business scenarios (such as electric vehicle charging and live electricity consumption) stored in a scattered manner, making it impossible to form a full-link association between customer, event, and risk, resulting in a relatively one-sided risk identification; Second, the semantic understanding accuracy is low, with general models struggling to interpret power-related terms such as peak-valley electricity prices and transformer line losses, leading to significant deviations in the extraction of risk features from customer request texts; Third, scenario adaptability is poor, with traditional fixed rules unable to cope with the dynamic risk characteristics of new business scenarios such as electric vehicle charging and live electricity consumption, requiring frequent manual adjustments; Fourth, there is a high degree of reliance on human intervention, with high-risk event screening and risk tracing relying heavily on human experience, resulting in low efficiency and a high risk of missed detections.
[0004] In summary, there is an urgent need to build a risk identification method and system that integrates the characteristics of the power industry with advanced algorithmic thinking in order to improve the intelligence level and risk prevention and control capabilities of power marketing services. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for risk hierarchical identification of power marketing service events based on knowledge graph semantic modeling. By constructing a complete system of data integration, feature extraction, risk identification and dynamic optimization, it can achieve accurate and intelligent risk identification of power marketing service events.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0007] A method for risk-stratified identification of electricity marketing service events based on knowledge graph semantic modeling includes the following steps:
[0008] S1: Construct a knowledge graph and multi-source data preprocessing system for power marketing services. By integrating multi-source data and domain knowledge from all aspects of power marketing, a knowledge graph including customer nodes, event nodes, risk nodes, and business nodes is established to form a data association network covering all business and all scenarios. This solves the problem of data silos, achieves comprehensive data coverage of power marketing service event risks, and provides a data foundation for subsequent risk feature extraction and identification.
[0009] S2: Design a hierarchical semantic modeling mechanism based on knowledge graph semantic embedding. By mapping knowledge graph nodes to semantic vector space and combining them with semantic encoding of customer demand text, a hierarchical semantic feature extraction module is constructed, consisting of a basic feature layer, an associated feature layer, and a risk feature layer. This forms a multi-dimensional feature matrix to achieve comprehensive extraction of risk features. At the same time, an attention mechanism is used to strengthen the expression of key risk features, solve the problem of semantic understanding bias, and provide feature support for risk layer identification.
[0010] S3: Construct a risk stratification identification model that integrates redundancy elimination and precise positioning. First, the first-level risk screening module filters the candidate set of high-risk events and removes redundant events. Then, the second-level risk detailed judgment module identifies the specific risk type. Finally, the third-level risk tracing module traces the risk trigger source and removes irrelevant interference information. At the same time, a dynamic risk threshold adjustment mechanism is set to improve the model's adaptability to different business scenarios. The model outputs risk stratification identification results, solves the problems of low risk identification efficiency and ambiguous tracing, and provides accurate basis for risk disposal.
[0011] S4: Establish a risk identification result verification and feedback optimization system. By constructing a multi-scenario verification dataset, the accuracy and recall of the risk stratification identification model are tested. A manual feedback mechanism is established to verify the model identification results. A model iteration and update module is designed to regularly optimize the knowledge graph and the model, solve the problem of poor model scenario adaptability, and ensure that the method adapts to changes in power marketing service business in the long term.
[0012] Furthermore, step S1 specifically includes the following sub-steps:
[0013] S1.1: Collect multi-source data from all aspects of electricity marketing, including basic customer information, 95598 complaint records, call transcripts, electricity payment data, power supply fault repair records, and new business scenario data. The 95598 complaint records include customer complaint text, complaint submission time, and affiliated transformer area information. The new business scenario data includes electric vehicle charging behavior data and live-streaming electricity consumption data. The electric vehicle charging behavior data includes charging time, power curve, and charging interruption records. The live-streaming electricity consumption data includes records of live-streaming time, load fluctuations, and voltage stability. At the same time, integrate professional knowledge in the field of electricity marketing, including business process specifications, historical cases of risk events, and expert experience rules, to provide raw data and domain knowledge support for knowledge graph construction.
[0014] S1.2: Preprocess the collected multi-source data, using automated cleaning tools combined with manual verification to remove redundant data, correct erroneous data, and fill in missing data. Perform word segmentation, part-of-speech tagging, and entity recognition on unstructured text data to extract customer demand entities, event type entities, and risk element entities, ensuring the data quality input into the knowledge graph.
[0015] S1.3: Based on preprocessed data and domain knowledge, construct a knowledge graph for power marketing services. The graph nodes include customer nodes, event nodes, risk nodes, and business nodes. The edge relationships include: the association between customers, demands, and events; the mapping relationship between events, risk elements, and risk levels; and the triggering relationship between business links and events. This forms a data association network covering all business and all scenarios, realizing the association mapping between multiple entities and providing a foundation for risk tracing.
[0016] Furthermore, step S2 specifically includes the following sub-steps:
[0017] S2.1: Based on the knowledge graph constructed in step S1, the TransH algorithm is used to map each node in the graph to the semantic vector space to generate node semantic embedding vectors; at the same time, the BERT-BiLSTM model is used to perform text semantic encoding on the customer request text to obtain text semantic feature vectors, thereby realizing the feature fusion of structured graph and unstructured text.
[0018] S2.2: Construct a hierarchical semantic feature extraction module. The first layer is the basic feature layer, which extracts basic semantic features such as event type, customer basic attributes, and business process. The second layer is the association feature layer, which extracts the association features between customer historical demands and current events, as well as the clustering features of similar events in the same region, based on knowledge graph association relationships. The third layer is the risk feature layer, which extracts risk semantic features such as risk element matching degree, customer emotional tendency, and event urgency, forming a multi-dimensional feature matrix to provide multi-dimensional feature input for risk classification.
[0019] S2.3: An attention mechanism is introduced to assign weights to the semantic features extracted in layers. Higher attention weights are given to high-risk correlation features of historical complaint records and power failure-prone areas, which strengthens the expression intensity of key risk features and improves the accuracy of risk identification.
[0020] Furthermore, step S3 specifically includes the following sub-steps:
[0021] S3.1: Construct a first-level risk screening module. Using the basic feature layer and associated feature layer constructed in step S2 as input, the LightGBM classification algorithm is used to perform preliminary risk classification of power marketing service events, screen out a candidate set of high-risk events, and exclude obviously risk-free routine business events. This completes the stripping of redundant events and reduces the computational load of the subsequent detailed judgment module.
[0022] S3.2: Construct a secondary risk sub-judgment module. Taking the high-risk event candidate set and risk semantic features after primary screening as input, the Stacking ensemble learning model is used to determine the risk sub-categorization. The base models of the Stacking ensemble learning model include XGBoost model, random forest model and SVM model, and the meta-model is logistic regression model. It identifies specific risk types such as electricity bill dispute risk, power supply quality risk and power supply guarantee risk in new business scenarios, so as to achieve accurate classification of risk types.
[0023] S3.3: Construct a three-level risk tracing module. Based on the association paths of the knowledge graph constructed in step S1, use path analysis algorithms to trace the triggering source of risk events, such as power quality risks caused by equipment failure and electricity bill disputes caused by policy interpretation errors. At the same time, exclude association paths that are not related to the current risk, complete the stripping of irrelevant interference information, and achieve accurate positioning of risk triggering sources.
[0024] S3.4: Set up a dynamic risk threshold adjustment mechanism. Using historical risk event handling results and manual feedback data as input, the mechanism uses reinforcement learning algorithms to optimize the judgment threshold of the risk stratification identification model in real time, thereby improving the model's adaptability to different business scenarios such as peak electricity consumption periods and new business promotion periods.
[0025] Furthermore, step S4 specifically includes the following sub-steps:
[0026] S4.1: Construct a multi-scenario verification dataset covering regular business scenarios, peak electricity consumption scenarios, and new business scenarios. The new business scenarios include electric vehicle charging scenarios and live streaming electricity consumption scenarios. Use this dataset to test the accuracy and recall of the risk stratification identification model and verify the model's performance in different scenarios.
[0027] S4.2: Establish a manual feedback mechanism, invite power marketing experts to verify the model identification results, and include erroneous identification cases such as misjudged risk levels and missed risk types into the feedback dataset to provide a basis for correction for iterative optimization of model parameters;
[0028] S4.3: Design model iteration and update module, with a monthly cycle, updates the nodes and relationships of the knowledge graph based on newly added power marketing service data and manual feedback data, and retrains and optimizes the hierarchical semantic modeling mechanism and risk hierarchical identification model to ensure that the model performance continuously adapts to business changes.
[0029] Furthermore, in step S1.1, data augmentation techniques such as synonym replacement and sentence restructuring are used to expand the data volume for new business scenario data, thereby solving the problem of poor model generalization ability caused by insufficient data samples in new business scenarios.
[0030] Furthermore, in step S1.3, entity linking technology is used in the knowledge graph construction process to match the entities identified in the customer's request text with the existing standard entities in the knowledge graph, thereby resolving entity ambiguity issues such as disputes over electricity bill calculation and delayed electricity bill payment, and improving the semantic consistency of the graph.
[0031] Furthermore, in step S2.1, the BERT-BiLSTM model is pre-trained in the field of electricity marketing. A domain pre-training corpus is constructed using professional text data on electricity marketing, such as business manuals and historical request records collected in step S1.1. Domain pre-training improves the model's semantic understanding of electricity industry terms such as peak-valley electricity prices and transformer area line losses.
[0032] In step S2.2, the extraction of customer sentiment tendencies uses the TextCNN model to identify sentiment keywords such as "dissatisfaction," "complaint," and "urgent" in the customer's complaint text, and generate sentiment tendency values by combining the contextual semantics; the range of the sentiment tendency value is [-1, 1], where -1 represents extremely negative and 1 represents extremely positive, and the sentiment tendency value is used as an important input feature of the risk feature layer.
[0033] Furthermore, in step S3.1, the feature inputs of the first-level risk screening module include the number of similar events in the past 30 days (i.e., event frequency), the number of complaints in the past year (i.e., customer complaint history), and the probability of risk occurrence in each business link based on historical data statistics (i.e., business link risk coefficient). The efficient screening of the candidate set of high-risk events is achieved through the combination of these three types of features.
[0034] Furthermore, in step S3.3, the risk tracing process uses a knowledge graph path scoring algorithm to score the associated paths of event nodes, risk element nodes, and trigger source nodes. The scoring formula is: path score = node semantic similarity × 0.4 + relationship confidence × 0.3 + historical path occurrence frequency × 0.3. The top 3 paths with the highest scores are selected as the main risk tracing results, and irrelevant paths with scores below the cross-validation determination threshold are excluded to improve the accuracy of risk tracing.
[0035] Furthermore, in step S4.3, an incremental training strategy is adopted during the model iteration update process. Only the knowledge graph nodes, semantic features and model parameters corresponding to the newly added data are updated, without retraining the entire model, thereby reducing the consumption of computing resources and improving the update efficiency.
[0036] Furthermore, the method also includes a risk warning output step: supplementing the knowledge graph constructed in step S1 with gridded service unit nodes and the relationship between customers and gridded service units; based on the risk stratification identification results output in step S3, generating risk warning information including risk event type, risk level, scope of affected customers, suggested handling measures and handling time limit; and pushing the risk warning information to the corresponding gridded service unit to support rapid risk response.
[0037] Furthermore, during the process of pushing risk warning information, the target service unit is located by combining the relationship between customers and gridded service units in the knowledge graph, so as to achieve accurate and targeted push of warning information, avoid ineffective push across regions and units, and improve service response efficiency.
[0038] A risk stratification identification system for power marketing service events based on knowledge graph semantic modeling includes a data acquisition and preprocessing module, a knowledge graph construction module, a stratified semantic modeling module, a risk stratification identification module, and a verification and feedback optimization module.
[0039] The data acquisition and preprocessing module is used to collect multi-source data and domain knowledge from all aspects of power marketing business, perform redundant data removal, error data correction, and missing data filling operations, and perform word segmentation, part-of-speech tagging, and entity recognition on unstructured text data, outputting preprocessed data and domain knowledge. The output of the data acquisition and preprocessing module is connected to the input of the knowledge graph construction module. The knowledge graph construction module is used to construct a knowledge graph including customer nodes, event nodes, risk nodes, and business nodes based on the preprocessed data and domain knowledge, forming a data association network covering all business and all scenarios, and outputting the knowledge graph. The output of the knowledge graph construction module is connected to the input of the hierarchical semantic modeling module. The hierarchical semantic modeling module is used to map knowledge graph nodes to semantic vector space and combine customer request text semantic encoding to construct a hierarchical semantic feature layer, a related feature layer, and a risk feature layer. The feature extraction module forms a multi-dimensional feature matrix and strengthens the expression of key risk features through an attention mechanism, outputting multi-dimensional risk features. The output of the hierarchical semantic modeling module is connected to the input of the risk hierarchical identification module, which performs risk hierarchical identification: a first-level risk screening module filters high-risk event candidate sets and removes redundant events; a second-level risk fine-tuning module identifies specific risk types; a third-level risk tracing module traces risk trigger sources and removes irrelevant interference information; and a threshold adjustment sub-module optimizes the judgment threshold, outputting the risk hierarchical identification result. The output of the risk hierarchical identification module is connected to the input of the verification and feedback optimization module, which constructs a multi-scenario verification dataset to test the risk hierarchical identification result, establishes a manual feedback mechanism to verify the model output, and regularly updates the knowledge graph and parameters of each module to ensure the system adapts to business changes.
[0040] Furthermore, the system also includes a domain pre-training submodule and a path scoring submodule. The domain pre-training submodule is integrated into the hierarchical semantic modeling module and is used to pre-train the BERT-BiLSTM model using professional text data on power marketing, thereby improving the model's semantic understanding of power industry terms. The path scoring submodule is integrated into the three-level risk tracing module of the risk hierarchical identification module and is used to score the knowledge graph association paths using a path scoring formula. The path scoring formula is: Path Score = Node Semantic Similarity × 0.4 + Relationship Confidence × 0.3 + Historical Path Occurrence Frequency × 0.3, and the top 3 scored paths are selected as the main risk tracing results.
[0041] The beneficial effects of this invention are:
[0042] This invention breaks through the traditional approach of optimizing a single link by constructing a full-link technology system encompassing data, graphs, semantics, recognition, and optimization. It deeply integrates multi-source data with knowledge graphs and deep learning technologies, achieving full data coverage, semantic accuracy, and process automation for risk identification. Customized technical solutions are tailored to the characteristics of the power industry, such as structuring data for 95598 requests and enhancing data for new business scenarios, solving the adaptability problem of general technologies in the power sector. Through the synergistic effect of methods and systems, the technical solutions are materialized into a modular system that can be implemented, ensuring that the technical effects are reproducible and scalable. This invention supports the digital transformation of enterprise electricity marketing, promotes the upgrade from human experience-driven to data intelligence-driven, reduces enterprise operating costs, and improves resource utilization efficiency. At the same time, it improves the quality of user service, shortening the customer request response time from 48 hours to 28.8 hours (based on a 3-month pilot data collection in 14 cities of a provincial power company), increasing the power supply quality problem resolution rate from 60% to 90% (the pilot covered 56 million electricity customers and 1.2 million request records), and increasing customer satisfaction from 82 points to 91 points. In particular, it ensures the stability of power supply for users of new businesses such as electric vehicles and live streaming, and can help promote the coordinated development of people's livelihood and emerging industries.
[0043] Multi-source data integration: Collect data from all aspects of electricity marketing, clarify the 95598 complaint records including customer complaint text, submission time, and affiliated transformer area information, and simultaneously integrate data from new business scenarios such as electric vehicle charging and live electricity consumption to form a full-dimensional data pool, completely breaking down data silos. Compared with the traditional decentralized data processing model, the data association coverage rate has increased from 40% to 95%. The inclusion of new business scenario data has expanded the risk identification scenarios from traditional businesses to emerging fields, solving the problem of lack of data support for new business risks. For example, the identification coverage rate of electric vehicle charging interruption events has increased from 30% to 92%.
[0044] Knowledge Graph Construction: A knowledge graph including customers, events, risks, and business nodes is constructed. Ambiguous entities in customer request texts are matched through entity linking technology. For example, electricity bill issues correspond to disputes over electricity bill calculation or payment delays, forming a full-business related network. This effectively improves the semantic consistency of data, and the accuracy rate of entity ambiguity resolution reaches 89%, avoiding risk misjudgment caused by entity confusion. The association relationship of the graph improves the efficiency of mining historical customer requests and current events, as well as similar events in the same region, by 60%. For example, a charging pile failure event in a certain area can be quickly linked to the core risk source of insufficient transformer capacity, improving the source tracing efficiency by 50%.
[0045] Hierarchical semantic modeling combines the semantic embedding of knowledge graph nodes with the text encoding of customer requests. Features are extracted from the basic layer (event type, customer attributes), the association layer (historical request association, regional event clustering), and the risk layer (emotional tendency, urgency). The attention mechanism is used to strengthen the weight of high-risk features, which effectively improves the extraction accuracy of risk features. The semantic understanding accuracy of power professional terms has increased from 60% to 85%, and the contribution of high-risk features has increased from 30% to 55%. Compared with the general semantic model, the misjudgment rate of risk type identification has been reduced by 42%. For example, the accuracy of identifying the risk type of unstable voltage for live power supply has reached 88%.
[0046] The three-tiered risk identification system consists of three levels: Level 1 screening, which filters out over 80% of redundant events based on event frequency, complaint history, and business risk coefficients; Level 2 detailed judgment, which identifies multiple risk types through an ensemble learning model; and Level 3 source tracing, which locates the trigger source using a path scoring algorithm (weighted by semantic similarity, relationship confidence, and historical frequency). Simultaneously, risk thresholds are dynamically adjusted, effectively reducing the cost of manual intervention. High-risk event screening efficiency is improved by 70%, and the average daily number of events handled by frontline staff is reduced from 200 to 56. Risk source tracing accuracy reaches 87%, and the time required for traditional manual source tracing is reduced from 2 hours / item to 1.1 hours / item. Dynamic threshold adjustment improves the model's adaptability to peak electricity consumption periods and new business promotion periods by 40%, avoiding missed detections due to changes in scenarios.
[0047] Dynamic optimization: Risk identification results are tested using multi-scenario validation datasets. Knowledge graph node relationships and model parameters are updated in conjunction with human feedback. Incremental training is used to shorten update time, ensuring the long-term effectiveness of the model. Monthly iterations keep the risk identification accuracy stable at over 88%. Incremental training reduces model update time from 8 hours to 1.5 hours, reducing computational resource consumption. At the same time, the human feedback mechanism allows the model to continuously adapt to business changes. For example, after adding new energy storage user electricity scenarios, only 15% of the parameters need to be adjusted to achieve risk identification coverage. Attached Figure Description
[0048] Figure 1 This is a flowchart of the method of the present invention;
[0049] Figure 2 This is a system structure diagram of the present invention. Detailed Implementation
[0050] To facilitate understanding of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the embodiments described are merely illustrative of the invention and should not be considered as specific limitations thereof.
[0051] In the following description of the embodiments, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0052] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of a described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that, as used in this specification and the appended claims, the term "and / or" refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0053] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0054] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. References to "one embodiment" or "some embodiments" in this application mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0055] Example 1: A Risk Hierarchy Identification Method for Electricity Marketing Service Events Based on Knowledge Graph Semantic Modeling
[0056] This embodiment is based on the business scenario of risk identification in electricity marketing services of a provincial power company. It addresses the needs of handling requests from 56 million electricity customers (including residential, commercial, industrial, and new business users) across 14 cities under the company's jurisdiction. The specific implementation process is as follows: Figure 1 As shown.
[0057] Step S1: Constructing a knowledge graph and multi-source data preprocessing system for electricity marketing services
[0058] S1.1 Multi-source data acquisition and domain knowledge integration: collecting data from all aspects of power marketing business.
[0059] Customer basic information: Collects electricity account number, address, electricity type (residential / commercial / electric vehicle users, etc.) and payment records of 56 million customers through the SG186 electricity marketing system;
[0060] Requests and business data: The system collects nearly 1.2 million request records (including voice-to-text transcripts) and 800,000 power supply fault repair records through the 95598 hotline system; it also collects electric vehicle charging behavior data (300,000 charging duration and power curves) through the charging pile management platform; and it collects heating and cooling season electricity consumption data (average daily electricity consumption per household) and production electricity consumption data (including load fluctuation data from 2,000 live streaming bases) through the electricity consumption information collection system.
[0061] Domain knowledge integration: Collect business process specifications such as the State Grid Electricity Marketing Service Standards and Electricity Customer Complaint Handling Guidelines, historical case studies of risk events in the past 5 years (such as the power outage incident of a charging pile in a certain area in 2023), and experience rules provided by 10 electricity marketing experts (such as the need to pay special attention to live broadcast electricity load fluctuations exceeding 30%).
[0062] New business scenario data enhancement: To address the problem of insufficient data samples for electric vehicle charging and live streaming electricity consumption (initially 500 samples each), synonym replacement (e.g., replacing charging failure with charging pile power outage) and sentence recombination (e.g., recombinizing power outage during live streaming as power interruption during live streaming) techniques were used to expand the data to 1500 samples each, thus solving the problem of model generalization ability.
[0063] S1.2 Multi-source data preprocessing: Automated preprocessing was implemented using Python + Pandas, combined with manual verification (1% of the data was reviewed by 3 experts).
[0064] Structured data cleaning: Remove duplicate records (approximately 5%) from electricity bill payment data, correct equipment model errors in power supply fault repair records (e.g., a transformer with a capacity of 10 kVA mistakenly written as 100 kVA), and fill in missing values in residential electricity consumption data with the mean (missing value rate 2.3%).
[0065] Unstructured text processing: Jieba segmentation was used to segment the 95598 complaint text, and BERT-CRF model was used for entity recognition to extract customer complaint entities (such as disputes over electricity bill calculation), event type entities (such as power supply quality problems), and risk element entities (such as equipment failures). The entity recognition accuracy rate reached 91.5%.
[0066] S1.3 Knowledge Graph Construction: The graph is built based on the Neo4j graph database, and entity linking technology is used to resolve ambiguities.
[0067] The graph nodes consist of customer nodes (including account number and electricity type attributes), event nodes (including event type and occurrence time attributes), risk nodes (including risk level and scope of impact attributes), and business nodes (including business process and processing time limit attributes), totaling 1.8 million nodes.
[0068] Edge relationships: the relationship between customer-demand-event (e.g., customer A-proposes-charging pile failure event), the mapping relationship between event-risk factor-risk level (e.g., charging pile failure-insufficient equipment capacity-high risk), and the triggering relationship between business process-event (e.g., charging pile installation-trigger-power supply load event), totaling 4.2 million edges;
[0069] Entity linking: Matching the electricity bill issue entity in the appeal text with the electricity bill calculation dispute and electricity bill payment delay standard entities in the graph, the ambiguity resolution accuracy rate reaches 89%, improving the semantic consistency of the graph.
[0070] Step S2: Design a hierarchical semantic modeling mechanism based on knowledge graph semantic embedding
[0071] S2.1 Semantic Vector Generation and Text Encoding
[0072] Semantic embedding of knowledge graph nodes: Based on the knowledge graph constructed in step S1, the TransH algorithm (parameters: marginal γ=1.0, vector dimension 128, number of iterations 1000) is used to map each node to the semantic vector space and generate node semantic embedding vectors. For example, the cosine similarity between the charging pile fault node and the electric vehicle user node is 0.87. The TransH algorithm can handle multiple types of relationships between entities and is more suitable for the complex relationship scenarios of customers, demands, risks and business in this solution.
[0073] Text semantic encoding: The BERT-BiLSTM model is used to encode customer request texts. First, a domain pre-training corpus is built using 100,000 professional texts on electricity marketing (business manuals, historical request records) collected in step S1.1. The model is then pre-trained (learning rate 5e-5, batch size=32, training epochs 15). Then, the request text is input to generate a 256-dimensional text semantic feature vector, which improves the semantic understanding accuracy of terms such as peak and valley electricity prices and transformer line loss by 32%.
[0074] S2.2 Hierarchical semantic feature extraction: A three-layer feature extraction module is constructed and implemented using PyTorch.
[0075] Basic Feature Layer: Extracts basic semantic features such as event type (e.g., charging pile failure), customer basic attributes (e.g., commercial users), and business process (e.g., charging pile installation), totaling 12 features;
[0076] Association Feature Layer: Based on the graph association relationship, extract the association features of customer historical requests and current events (such as the number of similar requests by customers in the past 3 months) and the clustering features of similar events in the same area (such as the number of charging pile failures in a certain area in the past 7 days), for a total of 8 features;
[0077] Risk Feature Layer: Extracts risk element matching degree (e.g., the matching degree between insufficient equipment capacity and the event is 0.92), customer sentiment tendency (using the TextCNN model to identify keywords such as dissatisfaction and urgency, generating sentiment tendency values: the value range is [-1, 1], in this embodiment the sentiment tendency value of the impact of live broadcast power outage on revenue is -0.85), and event urgency (e.g., the interruption of live broadcast power is marked as urgent), for a total of 6 features.
[0078] The above three layers of features are combined to form a 26-dimensional multi-dimensional feature matrix.
[0079] S2.3 Attention Mechanism Enhances Key Features: A Scaled Dot-Product attention mechanism is introduced to assign feature weights.
[0080] High-risk associated features (such as customer complaint records in the past year, and events in areas with frequent power outages) are assigned a weight of 0.35, while conventional features (such as customer gender) are assigned a weight of 0.05.
[0081] After attention adjustment, the contribution of high-risk characteristics increased from 30% to 55%, strengthening the expression of key risk characteristics.
[0082] Step S3: Construct a risk stratification identification model that integrates redundancy elimination and precise positioning.
[0083] S3.1 First-level risk screening. A first-level screening module is constructed based on the basic feature layer and associated feature layer output from step S2, using the LightGBM classification algorithm (parameters: learning rate 0.1, number of trees 100, maximum depth 8):
[0084] Feature inputs: event frequency (number of similar events in the past 30 days, such as 5 charging pile failures in a certain area), customer complaint history (number of complaints in the past year, such as 2 complaints from customer A), and business process risk coefficient (based on historical data statistics, such as a risk coefficient of 0.75 for the charging pile installation process).
[0085] Screening results: The 100,000 test events were initially classified, and a candidate set of high-risk events (18,000 in total) was selected. At the same time, 82,000 risk-free routine business events (such as routine electricity bill inquiries) were excluded. The redundant event removal rate reached 82%, which reduced the subsequent computing load.
[0086] S3.2 Secondary Risk Detailed Assessment. For the high-risk candidate set after the primary screening, a Stacking ensemble learning model is used, combining semantic features from the risk feature layer:
[0087] The Stacking ensemble learning model addresses the issue of insufficient accuracy in identifying power supply risks for new businesses by fusing multiple models. Base models include: XGBoost (maximum depth 6), Random Forest (200 trees), and SVM (RBF kernel). The meta-model is a logistic regression model (regularization coefficient 0.1).
[0088] Risk type identification: Eight specific risk types are output, including electricity bill dispute risk, power supply quality risk, and power supply guarantee risk for new business scenarios. The accuracy rate of risk type identification in the test set reached 88.5%, of which the accuracy rate of identifying power supply guarantee risk for new business scenarios reached 86%.
[0089] S3.3 Level 3 Risk Source Tracing. Based on knowledge graph-based path analysis, a path scoring algorithm is employed.
[0090] Path analysis: Tracing the associated paths of charging pile failure events, such as charging pile failure → insufficient capacity of transformer in the distribution area → 12 new charging piles added in the past 3 months;
[0091] Path score: Calculated according to the formula "Path score = Node semantic similarity × 0.4 + Relationship confidence × 0.3 + Historical path occurrence frequency × 0.3". In this embodiment, the node semantic similarity is 0.93, the relationship confidence is 0.88, and the historical path occurrence frequency is 0.75. Therefore, the score = 0.93 × 0.4 + 0.88 × 0.3 + 0.75 × 0.3 = 0.86.
[0092] Results filtering: The top 3 paths with the highest scores were selected as the main traceability results, and irrelevant paths with scores below the threshold (threshold 0.7, determined by 5-fold cross-validation) were excluded (such as charging pile failure → customer payment delay, score 0.52). The accuracy of risk trigger source location reached 87%.
[0093] S3.4 Dynamic adjustment of risk threshold. Based on the historical risk event processing results (50,000 records) and manual feedback data (10,000 records) from 2023, the DQN reinforcement learning algorithm is adopted:
[0094] State space: Business scenario type (regular / peak electricity consumption / new business), event feature vector;
[0095] Action space: threshold adjustment range (±0.05);
[0096] Reward function: Rewards are based on the accuracy of risk identification (accuracy > 90% reward +10, < 80% penalty -5);
[0097] Optimization results: The risk assessment threshold for peak electricity consumption periods (such as the summer cooling season) was adjusted from 0.65 to 0.58, and the threshold for the promotion period of new businesses (such as the concentrated installation period of charging piles) was adjusted from 0.7 to 0.62. The adaptability of the model to different scenarios was improved by 40%.
[0098] Step S4: Establish a risk identification result verification and feedback optimization system
[0099] S4.1 Multi-scenario Verification.
[0100] Construct a multi-scenario validation dataset (30,000 records in total):
[0101] Scenario coverage: Regular business scenarios (10,000 items, such as daily electricity bill inquiries), peak electricity consumption scenarios (10,000 items, such as summer power supply failures), and new business scenarios (10,000 items, such as electric vehicle charging and live streaming electricity consumption).
[0102] Test metrics: The risk stratification identification model was tested for accuracy and recall. In regular scenarios, the accuracy was 92% and the recall was 90%, while in new business scenarios, the accuracy was 88% and the recall was 86%, which meets business requirements.
[0103] S4.2 Human feedback.
[0104] A feedback panel of 15 electricity marketing experts (with ≥8 years of work experience) was invited.
[0105] Verification method: 5,000 model recognition results are randomly selected, and experts verify the risk level, type, and source tracing results;
[0106] Feedback Data: 180 misidentification cases (such as misjudging the fluctuation of electricity consumption during live broadcasts as low risk and failing to identify insufficient transformer capacity as a trigger source) were included in the feedback dataset, and the reasons for the errors were labeled (such as insufficient weight of sentiment tendency features).
[0107] S4.3 model iterative update.
[0108] The module is designed for monthly iterative updates, employing an incremental training strategy:
[0109] Knowledge graph update: Based on 10,000 new power marketing data entries, 5,000 new customer nodes and 3,000 new event nodes were added, and 2,000 edge relationships of new business-risk were updated without reconstructing the entire graph;
[0110] Model parameter updates: Only the semantic feature weights corresponding to the feedback data (such as adjusting the customer sentiment tendency feature weight from 0.2 to 0.3) and the risk stratification identification model parameters (such as adjusting the Stacking meta-model regularization coefficient from 0.1 to 0.08) are updated, reducing the training time from 8 hours of full training to 1.5 hours and reducing the consumption of computing resources.
[0111] Example 2: A Risk Hierarchy Identification System for Electricity Marketing Service Events Based on Knowledge Graph Semantic Modeling
[0112] This embodiment corresponds to the method in Embodiment 1, and discloses the hardware architecture, module implementation, and data interaction of the system.
[0113] (I) System Hardware Architecture
[0114] Servers: Two Huawei TaiShan 2280 servers (CPU: Kunpeng 920 32 cores, memory: 128 GB, hard disk: 2TB SSD) are used, one for data processing and model training, and the other for map storage and service deployment.
[0115] Database: Neo4j graph database (version 5.10, cluster mode, 3 nodes) stores knowledge graphs, MySQL database (version 8.0) stores structured data (such as customer information, event records);
[0116] Interactive terminals: 50 web terminals (browser: Chrome 110+) are deployed in grassroots power supply stations to view risk identification results and submit manual feedback.
[0117] (ii) System module implementation (e.g.) Figure 2 (As shown)
[0118] 1. Data Acquisition and Preprocessing Module
[0119] Interface design: The RESTful API is developed using Java and interfaces with the SG186 marketing system, the 95598 hotline system, and the charging pile management platform. Real-time data transmission is achieved through Apache Kafka (throughput: 1000 messages / second).
[0120] Preprocessing function: Integrates Kettle ETL tool to perform data cleaning, calls jieba word segmentation interface and BERT-CRF entity recognition interface (deployed on TensorFlow Serving), and outputs structured data in JSON format (example: {"Customer ID":"C00123", "Event Type":"Charging Pile Failure", "Preprocessing Status":"Completed"}).
[0121] 2. Knowledge Graph Construction Module
[0122] Entity and Relationship Extraction: An extraction tool based on Python + Spacy is developed to extract entities and relationships from structured data. Neo4j's Cypher statement is used to create nodes and edges (e.g., CREATE(c:customer{ID:'C00123'})-[:submit]->(e:event{type:'charging pile failure'})).
[0123] Entity linking function: Integrates a BERT-based entity disambiguation model (deployed on PyTorch Serving), takes the entity to be matched as input and the standard entity library as input, and returns the matching result. The disambiguation resolution response time is less than 1 second.
[0124] 3. Hierarchical Semantic Modeling Module
[0125] Domain pre-training submodule: Based on the Hugging Face Transformers library, load the bert-base-chinese model, pre-train it with 100,000 power texts from Example 1, generate a domain pre-trained model file (approximately 1.2GB), deploy it on TorchServe, and support batch inference (batch size=64, inference time <0.5 seconds / text).
[0126] Feature layer extraction submodule: Develop a Python script to call the TransH algorithm interface to generate node vectors, call the pre-trained model interface to generate text vectors, merge them into a multi-dimensional feature matrix, and store it in a Redis cache (expiration time of 1 hour) for subsequent modules to call.
[0127] 4. Risk Stratification and Identification Module
[0128] Level 1 Risk Screening Module: Implements the LightGBM classifier based on Scikit-learn, loads a pre-trained model (model file approximately 500 MB), inputs a feature matrix, outputs a high-risk candidate set, with an accuracy of 92%;
[0129] Level 2 Risk Assessment Module: Integrates base models including XGBoost, Random Forest and SVM models, and logistic regression meta-models, deployed in Docker containers, and supports hot model updates (update time < 5 minutes).
[0130] The Level 3 Risk Source Tracing Module (including the Path Scoring Submodule) develops path scoring algorithm scripts, calls Neo4j's path query interface (such as MATCH p=(e:event)-[r:association]->(f:risk factor)RETURN p), calculates path scores and filters results, with a source tracing response time of <2 seconds.
[0131] Threshold Adjustment Submodule: Integrates DQN reinforcement learning algorithm, inputs historical risk event processing results (50,000 records) and manual feedback data (10,000 records), defines the state space as business scenario type + event feature vector, the action space as threshold ±0.05 adjustment range, and the reward function as +10 for risk identification accuracy > 90% and -5 for < 80%, and optimizes the judgment threshold in real time, such as adjusting the threshold from 0.65 to 0.58 during peak electricity consumption periods.
[0132] 5. Verification and Feedback Optimization Module
[0133] Multi-scenario verification function: Develop Python test scripts, read the verification dataset in MySQL, call the risk identification module interface, and output accuracy and recall reports (format: Excel).
[0134] Human feedback function: The web terminal develops a feedback interface based on Vue.js, which supports experts to mark "correct / incorrect" and the reason for the error, and the feedback data is written to MySQL in real time;
[0135] Iterative update function: Develop a Shell scheduled script (executed at 0:00 on the 1st of each month) to call the Neo4j update interface and model training interface to automatically complete the map and model update, and store the logs in the / var / log / system_update / directory.
[0136] (III) System Data Interaction Process
[0137] 1) Data Acquisition and Preprocessing Module → Knowledge Graph Construction Module: Structured Data (JSON Format), Transmission Frequency: Real-time;
[0138] 2) Knowledge graph construction module → hierarchical semantic modeling module: graph node ID and attributes (e.g., event ID: E001, type: charging pile fault), transmission frequency: once per hour;
[0139] 3) Hierarchical semantic modeling module → Risk stratification identification module: Multi-dimensional feature matrix (CSV format), transmission frequency: real-time;
[0140] 4) Risk stratification identification module → Verification and feedback optimization module: Risk identification results (e.g., event ID: E001, risk type: new business power supply guarantee risk, source tracing result: insufficient transformer capacity), transmission frequency: real-time;
[0141] 5) Verification and Feedback Optimization Module → Knowledge Graph Construction Module / Hierarchical Semantic Modeling Module: Feedback data (e.g., Event ID: E001, Error Type: Risk Level Misjudgment, Adjustment Parameter: Threshold -0.05), Transmission Frequency: Once a day.
[0142] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0143] In various embodiments, the hardware implementation of the technology can directly utilize existing smart devices, including but not limited to industrial control computers, PCs, smartphones, handheld devices, and floor-standing devices. Its input device preferably uses an on-screen keyboard, its data storage and computing modules utilize existing memory, calculators, and controllers, its internal communication modules utilize existing communication ports and protocols, and its remote communication utilizes existing GPRS networks, the World Wide Web, etc.
[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0145] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0146] In the various embodiments of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units. If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
Claims
1. A method for risk-stratified identification of electricity marketing service events based on knowledge graph semantic modeling, characterized in that, Includes the following steps: S1: Construct a knowledge graph and multi-source data preprocessing system for power marketing services. By integrating multi-source data and domain knowledge from all aspects of power marketing, a knowledge graph including customer nodes, event nodes, risk nodes, and business nodes is established to form a data association network covering all business and all scenarios. S2: Design a hierarchical semantic modeling mechanism based on knowledge graph semantic embedding. Use the TransH algorithm to map each node in the knowledge graph to the semantic vector space to generate node semantic embedding vectors. At the same time, use the BERT-BiLSTM model pre-trained in the field of power marketing to encode the text of customer requests and obtain the text semantic feature vector. A hierarchical semantic feature extraction module is constructed. The first layer is the basic feature layer, which extracts basic semantic features of event type, customer basic attributes and business process. The second layer is the association feature layer, which extracts the association features between customers' historical demands and current events, as well as the clustering features of similar events in the same region, based on the knowledge graph association relationship; the third layer is the risk feature layer, which extracts the risk element matching degree, customer emotional tendency and event urgency degree risk semantic features, and merges them to form a multi-dimensional feature matrix; An attention mechanism is introduced to assign higher attention weights to high-risk correlation features of historical complaint records and events in areas with frequent power outages, thereby strengthening the expression of key risk features. S3: Construct a risk stratification identification model that integrates redundancy elimination and precise positioning, and execute a three-level progressive risk identification process: S3.1: Construct a first-level risk screening module. Using the basic feature layer and related feature layer of step S2 as input, the LightGBM classification algorithm is used to perform preliminary risk classification of power marketing service events, screen out the candidate set of high-risk events and exclude risk-free routine business events. The feature inputs of the first-level risk screening module include the frequency of event occurrence, customer complaint history and risk coefficient of business process. S3.2: Construct a secondary risk sub-judgment module. Taking the high-risk event candidate set and risk semantic features after primary screening as input, the Stacking ensemble learning model is used to determine the risk sub-categories. The base models of the Stacking ensemble learning model include XGBoost model, random forest model and SVM model, and the meta-model is logistic regression model. It identifies specific risk types including electricity bill dispute risk, power supply quality risk and power supply guarantee risk in new business scenarios. S3.3: Construct a three-level risk tracing module. Based on the association paths of the knowledge graph, a path scoring algorithm is used to trace the triggering source of risk events. The calculation formula of the path scoring algorithm is: Path Score = Node Semantic Similarity × 0.4 + Relationship Confidence × 0.3 + Historical Path Occurrence Frequency × 0.
3. The top 3 paths with the highest scores are selected as the main risk tracing results, and association paths that are irrelevant to the current risk are excluded. At the same time, a dynamic risk threshold adjustment mechanism is set up. Using the historical risk event processing results and manual feedback data as input, the DQN reinforcement learning algorithm is used to optimize the judgment threshold of the risk stratification identification model in real time, and output the risk stratification identification results. S4: Establish a risk identification result verification and feedback optimization system. Test the accuracy and recall of the risk stratification identification model by constructing a multi-scenario verification dataset, establish a manual feedback mechanism to verify the model identification results, and regularly optimize the knowledge graph and model.
2. The method for risk hierarchical identification of power marketing service events based on knowledge graph semantic modeling according to claim 1, characterized in that, Step S1 specifically includes the following sub-steps: S1.1: Collect multi-source data from all aspects of electricity marketing, including basic customer information, 95598 complaint records, call transcripts, electricity payment data, power supply fault repair records, and new business scenario data. The 95598 complaint records include customer complaint text, complaint submission time, and affiliated transformer area information. The new business scenario data includes electric vehicle charging behavior data and live electricity consumption data. At the same time, integrate professional knowledge in the field of electricity marketing, including business process specifications, historical cases of risk events, and expert experience rules. S1.2: Preprocess the collected multi-source data, using automated cleaning tools combined with manual verification to remove redundant data, correct erroneous data, and fill in missing data. Perform word segmentation, part-of-speech tagging, and entity recognition on unstructured text data to extract customer request entities, event type entities, and risk element entities. S1.3: Based on the preprocessed data and domain knowledge, construct a knowledge graph for power marketing services. The graph nodes include customer nodes, event nodes, risk nodes, and business nodes. The edge relationships include: the association between customers, demands, and events; the mapping relationship between events, risk factors, and risk levels; and the triggering relationship between business links and events, forming a data association network covering all business and all scenarios.
3. The method for risk-layered identification of power marketing service events based on knowledge graph semantic modeling according to claim 1, characterized in that, Step S4 specifically includes the following sub-steps: S4.1: Construct a multi-scenario verification dataset, covering regular business scenarios, peak electricity consumption scenarios, and new business scenarios, including electric vehicle charging scenarios and live streaming electricity consumption scenarios; S4.2: Establish a manual feedback mechanism, and have power marketing experts verify the model identification results, and include erroneous identification cases, including misjudged risk levels and missed risk types, into the feedback dataset; S4.3: Design model iteration and update module, with a monthly cycle, updates the nodes and relationships of the knowledge graph based on newly added power marketing service data and manual feedback data, and retrains and optimizes the hierarchical semantic modeling mechanism and risk hierarchical identification model.
4. The method for risk-layered identification of power marketing service events based on knowledge graph semantic modeling according to claim 2, characterized in that, In step S1.1, data augmentation technology is used to expand the data volume for new business scenario data.
5. The method for risk hierarchical identification of power marketing service events based on knowledge graph semantic modeling according to claim 2, characterized in that, In step S1.3, entity linking technology is used in the knowledge graph construction process to match the entities identified in the customer request text with the existing standard entities in the knowledge graph.
6. A risk-stratified identification system for electricity marketing service events based on knowledge graph semantic modeling, characterized in that, It includes a data acquisition and preprocessing module, a knowledge graph construction module, a hierarchical semantic modeling module, a risk hierarchical identification module, and a verification and feedback optimization module; The data acquisition and preprocessing module is used to collect multi-source data and domain knowledge from all aspects of power marketing, perform redundant data removal, erroneous data correction, and missing data filling operations, and perform word segmentation, part-of-speech tagging and entity recognition on unstructured text data, and output preprocessed data and domain knowledge. The knowledge graph construction module is used to construct a knowledge graph including customer nodes, event nodes, risk nodes and business nodes based on preprocessed data and domain knowledge, forming a data association network covering all business and all scenarios; The hierarchical semantic modeling module has a built-in domain pre-training sub-module, which uses the TransH algorithm to map knowledge graph nodes to semantic vector space to generate node semantic embedding vectors. At the same time, it uses a BERT-BiLSTM model pre-trained in the power marketing field to semantically encode customer complaint text to obtain text semantic feature vectors. It constructs a hierarchical semantic feature extraction module with basic feature layer, associated feature layer and risk feature layer to form a multi-dimensional feature matrix. At the same time, it assigns higher attention weights to high-risk associated features of historical complaint records and power failure-prone area events through an attention mechanism to output multi-dimensional risk features. The risk stratification identification module includes a first-level risk screening submodule, a second-level risk detailed judgment submodule, a third-level risk tracing submodule, and a threshold adjustment submodule. The first-level risk screening submodule uses the LightGBM algorithm to screen the candidate set of high-risk events and remove redundant events. The input features include the frequency of event occurrence, customer complaint history, and risk coefficient of business process. The secondary risk assessment submodule uses a Stacking ensemble learning model to identify specific risk types; the tertiary risk tracing submodule has a built-in path scoring submodule, which uses a path scoring formula to trace the risk trigger source. The path scoring formula is: Path Score = Node Semantic Similarity × 0.4 + Relationship Confidence × 0.3 + Historical Path Occurrence Frequency × 0.
3. The top 3 paths with the highest scores are selected as the main risk tracing results. The threshold adjustment submodule uses the DQN reinforcement learning algorithm to optimize the judgment threshold and outputs the risk stratification identification result. The verification and feedback optimization module is used to construct a multi-scenario verification dataset to test the risk stratification identification results, establish a manual feedback mechanism to verify the model output, and regularly update the knowledge graph and the parameters of each module.