Fusion deep mode mining of electricity theft inspection rule intelligent agent construction method
Patent Information
- Application Number
- CN202611247787.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]综合来看,电力营销稽查的违约用电与反窃电场景迫切需要一种融合深度模式挖掘与大语言模型语义理解的规则智能生成方法,实现从业务场景到结构化稽查规则的端到端自动化生成,克服人工规则制定效率低、传统挖掘方法难以处理高维数据以及挖掘结果缺乏业务语义、难以直接应用于稽查实践等不足
(1)本发明提出一种融合深度模式挖掘的违窃电稽查规则智能体构建方法。该方法构建了自动化规则生成框架,通过场景数据映射、规则生成及规则去噪总结,将业务场景理解、数据映射与获取、深度模式挖掘及规则生成等关键环节有机耦合,依托智能体根据任务状态驱动关键流程,实现了从用户业务需求到结构化稽查规则的端到端自动生成。
Smart Images

Figure CN122798445A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of anti-electricity theft and illegal electricity use technology, specifically involving a method for constructing an intelligent agent for investigating illegal electricity theft rules by incorporating deep pattern mining. Background Technology
[0002] Anti-electricity theft and illegal electricity use (collectively referred to as illegal electricity theft) are key scenarios for electricity marketing audits. The most direct manifestation of illegal electricity theft is its historical electricity consumption data, which is usually accompanied by abnormal changes in multiple dimensions of data such as user profiles, metering equipment, electricity consumption behavior, electricity consumption and charges, and business transactions. Moreover, the two scenarios have certain differences. For example, the anti-electricity theft scenario focuses more on continuous numerical characteristics such as electricity consumption and charges, while the illegal electricity use scenario focuses more on discrete fields related to user profiles.
[0003] In generating rules for electricity theft scenarios, traditional methods primarily rely on domain experts manually formulating rules based on their business experience. Business personnel summarize historical audit cases, combine metering principles and business specifications, and design audit judgment logic and threshold parameters for each rule. While this approach ensures that rules align with business semantics, it has a long development cycle, limited coverage, and difficulty in uncovering potential complex correlation patterns in massive amounts of data. Especially in high-dimensional business data scenarios, manual rules can only cover common, superficial anomaly patterns; deep nonlinear correlations and cross-field combination anomaly patterns hidden in the data are difficult to discover through manual experience.
[0004] Association rule mining is a typical method for extracting rules from structured data, mainly including algorithms such as Apriori and FP-Growth. It discovers inter-item relationships in transactional data through frequent itemset enumeration and association metrics. However, electricity marketing audit data is characterized by high dimensionality, continuous values, and multi-table associations. Traditional frequent itemset enumeration faces the combinatorial explosion problem and struggles to handle nonlinear association patterns in continuous data.
[0005] In the areas of intelligent agents and large language models, large language models have achieved significant breakthroughs in natural language understanding and generation tasks in recent years, demonstrating powerful semantic understanding and knowledge reasoning capabilities. Intelligent agent technology combines large language models with external tools, enabling the model to autonomously plan tasks, invoke tools, and orchestrate multi-step workflows to complete complex automated tasks. This architecture uses the large language model as the core inference engine and external tools as the execution unit, allowing the model to autonomously select tools, pass parameters, and schedule tasks, providing a new paradigm for multi-stage automated processing.
[0006] In summary, the scenarios of electricity marketing audits involving illegal electricity use and anti-electricity theft urgently require an intelligent rule generation method that integrates deep pattern mining and semantic understanding of large language models. This method would enable end-to-end automated generation of structured audit rules from business scenarios, overcoming the shortcomings of low efficiency in manual rule formulation, the difficulty of traditional mining methods in handling high-dimensional data, and the lack of business semantics in mining results, making them difficult to directly apply to audit practices. Summary of the Invention
[0007] To overcome the problems in the prior art, this invention proposes a method for constructing an intelligent agent for investigating illegal electricity theft based on deep pattern mining.
[0008] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: This invention provides a method for constructing an intelligent agent for investigating illegal electricity theft based on deep pattern mining, comprising the following steps: A scenario knowledge configuration and scenario data mapping tool is constructed between anti-electricity theft and illegal electricity use business scenarios and data resources. The scenario data mapping tool adopts a scenario parsing and data mapping method driven by scenario knowledge configuration, and completes scenario parsing, data resource location and data acquisition SQL generation according to user input. A rule generation tool is constructed, which adopts a rule generation method that integrates data preprocessing and deep autoencoder. Based on the data acquisition SQL, the tool acquires data and sequentially performs preprocessing, deep pattern learning and association rule extraction processes to generate a set of candidate rules. A rule denoising and summarization tool is constructed, which adopts a rule denoising and semantic structuring method based on a large language model to transform candidate rules into structured rules with practical business meaning; Based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, an intelligent agent is constructed for generating rules for investigating illegal electricity theft.
[0009] Furthermore, the scenario parsing and data mapping method based on scenario knowledge configuration-driven methods includes: Based on scenario knowledge configuration, a scenario data acquisition Prompt template is designed, which injects scenario knowledge configuration, database schema and user requirements into the model context in a unified manner, and guides the model to complete the generation of data acquisition SQL through an SQL validity verification mechanism.
[0010] Furthermore, the scene data acquisition Prompt template Defined as: ; in, Represents the schema content of a database table; Represents specific information for a particular scene; Indicates the current time; This indicates important reminders related to scene data acquisition; Indicates the time range included in the user's question; This indicates other filter criteria in the user's question; This indicates the original user question.
[0011] Furthermore, the rule generation method that integrates deep autoencoders with data preprocessing includes: After receiving the data retrieval SQL, a connection is established with the target database based on the data retrieval SQL, the data retrieval SQL is executed to obtain the required business data, the validity of the business data is verified, and the verified business data is preprocessed. The preprocessed data is converted into one-hot encoded form; a deep autoencoder is constructed and trained; the trained deep autoencoder is used to perform inference analysis on the preprocessed data to extract candidate rules; the candidate rules are evaluated and filtered to obtain the final set of candidate rules.
[0012] Furthermore, the data preprocessing includes: The original power data is subjected to hierarchical processing for missing values. The field types of the original power data include numeric fields and categorical fields. After processing for missing values, outlier identification and filtering are performed on the numeric fields, and the numeric fields are discretized.
[0013] Furthermore, a supervised optimal binning method is adopted to discretize numerical fields based on business target fields, thus preserving the differences in business patterns in the data.
[0014] Furthermore, rule-based denoising and semantic structuring methods based on large language models include: Based on the candidate rule set and combined with business scenario knowledge, a rule denoising summary Prompt template is designed. The candidate rules are denoised and their effectiveness is screened through a large language model. The screened candidate rules are then converted into a unified structured format, which is a structured rule with actual business meaning.
[0015] Furthermore, rule-based denoising summarizes the Prompt template. Defined as: ; in, Indicates the name of the current inspection scenario; This indicates the data table and field structure information; Important reminders regarding rule-based noise reduction summary; This represents the set of candidate rules.
[0016] Furthermore, based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, an intelligent agent for generating rules for investigating illegal electricity theft is constructed, including: The constructed scene data mapping tool, rule generation tool, and rule denoising and summarization tool will be uniformly registered in the intelligent agent tool library; Based on the agent tool library, an agent Prompt template is designed to guide the agent to call the scene data mapping tool, rule generation tool and rule denoising and summarizing tool in sequence, so as to realize the automatic connection between data acquisition SQL generation, candidate rule generation and rule denoising and semantic structuring. Once the rule denoising and summarizing tool has finished executing, it outputs the final structured rule set.
[0017] Furthermore, the agent Prompt template Defined as: ; in, G Indicate the task objective; This indicates an important reminder that the intelligent agent has completed its task; The toolspace represents the tool space, containing each tool and its calling description; D Represents database resource information; Q This indicates the original user question.
[0018] Compared with the prior art, the present invention has the following technical effects: (1) This invention proposes a method for constructing an intelligent agent for investigating illegal electricity theft rules by integrating deep pattern mining. The method constructs an automated rule generation framework, which organically couples key links such as business scenario understanding, data mapping and acquisition, deep pattern mining and rule generation through scenario data mapping, rule generation and rule denoising and summarization. Relying on the intelligent agent to drive key processes according to task status, it realizes end-to-end automatic generation from user business needs to structured investigation rules.
[0019] (2) This invention proposes a scenario parsing and data mapping method based on scenario knowledge configuration. It constructs scenario knowledge configuration between anti-electricity theft and non-compliant electricity use business scenarios and data resources, combines scenario data acquisition Prompt templates, introduces database schema, scenario information, and user requirement constraints, automatically generates data acquisition SQL, and improves the accuracy of SQL generation through SQL syntax verification and self-correction mechanisms, providing a reliable data source for rule generation.
[0020] (3) This invention proposes a rule generation method that integrates data preprocessing and deep autoencoders. A data preprocessing workflow for electricity marketing data is constructed, improving data quality through missing value handling and outlier filtering. A supervised optimal binning discretization method combining target business category information is adopted, optimizing numerical interval division by utilizing the correlation between field values and business categories. Based on this, a deep autoencoder is used to learn potential correlation patterns in high-dimensional business data, and effective rules are selected through rule extraction and rule quality evaluation mechanisms, enhancing the ability to discover potential patterns in complex business scenarios.
[0021] (4) This invention proposes a rule denoising and semantic structuring method based on a large language model. By constructing a rule denoising summary tool and a Prompt template, business scenario information, field semantic information and candidate rule set are input into the large language model. Through business semantic analysis, the validity of rules is judged, redundant rules are eliminated, similar rules are merged and rules are expressed in a structured way, and candidate rules are converted into structured rules with business meaning, which are understandable and executable. Attached Figure Description
[0022] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of the method proposed in this invention; Figure 2 Example diagram of the configuration of the scenario for detecting illegal electricity use and anti-electricity theft constructed for the experiment of this invention; Figure 3 The specific content of the Prompt template for obtaining scene data is shown; Figure 4 The specific content of the rule-based noise reduction summary Prompt template is shown; Figure 5 The specific content of the agent Prompt template is shown. Detailed Implementation
[0024] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the technical solutions proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. Specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] To address the aforementioned issues, this invention proposes a method for constructing an intelligent agent for electricity theft investigation rules that integrates deep pattern mining. This method targets anti-electricity theft and illegal electricity use investigation scenarios. Based on large-scale model and intelligent agent technology, it constructs an intelligent agent for generating electricity theft investigation rules, with deep pattern mining at its core and the large model driving the invocation of various tools. Through the coordinated scheduling and execution of scenario parsing and data mapping, data preprocessing, deep pattern mining, and rule denoising and semantic structuring, the intelligent agent automatically discovers business patterns from massive amounts of data and generates structured investigation rules, providing a highly efficient and high-quality automatic rule generation technology for electricity theft investigation.
[0026] In this embodiment, refer to Figures 1-5 This paper provides a method for constructing an intelligent agent for investigating illegal electricity theft based on deep pattern mining, including the following steps: Step 100: Scene Analysis and Data Mapping: Construct a scene knowledge configuration and scene data mapping tool between anti-electricity theft and illegal electricity use business scenarios and data resources. The scene data mapping tool adopts a scene analysis and data mapping method driven by scene knowledge configuration, and completes scene analysis, data resource location and data acquisition SQL generation according to user input. Step 200: Preprocessing scheme construction: Pre-build a data preprocessing process to transform low-quality raw power data into high-quality rule mining data, providing a reliable data foundation for subsequent deep pattern mining; Step 300: Rule generation based on deep pattern mining: Construct a rule generation tool that adopts a rule generation method that integrates data preprocessing with a deep autoencoder. Based on the data acquisition SQL, the tool acquires data and sequentially performs preprocessing, deep pattern learning, and association rule extraction processes to generate a candidate rule set. Step 400: Rule Denoising and Semantic Structuring: Construct a rule denoising and summarizing tool, which adopts a rule denoising and semantic structuring method based on a large language model to transform candidate rules into structured rules with practical business meaning; Step 500: Rule-based intelligent agent construction: Based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, construct an intelligent agent for generating rules for investigating illegal electricity theft.
[0027] The following is a detailed explanation of each of the above steps: Step 100: Scene Analysis and Data Mapping: Construct a scene knowledge configuration and scene data mapping tool between anti-electricity theft and illegal electricity use business scenarios and data resources. The scene data mapping tool adopts a scene analysis and data mapping method driven by scene knowledge configuration, and completes scene analysis, data resource location and data acquisition SQL generation according to user input.
[0028] Because the data resources, business objects, and key indicators differ significantly across different inspection scenarios, such as anti-electricity theft and illegal electricity use, this invention designs a scenario parsing and data mapping method driven by scenario knowledge configuration. By constructing scenario knowledge configuration, scenario data acquisition Prompt templates, and scenario data mapping tools, it achieves scenario parsing based on user natural language requirements, automatic data resource location, and automatic generation of data acquisition SQL, providing standardized data input for subsequent rule mining. The specific steps are as follows: Step 110: Construct scenario knowledge configuration between anti-electricity theft and illegal electricity use business scenarios and data resources.
[0029] For the anti-electricity theft and illegal electricity use scenarios in electricity marketing auditing, scenario knowledge configurations are constructed. Each scenario knowledge configuration includes information such as the target business scenario, related data tables, core fields, inter-table relationships, business semantic descriptions, and filtering condition descriptions.
[0030] For any scenario, its scenario knowledge configuration structure can be represented as: ; in, This indicates specific information about a particular scene; This represents the set of associated data tables corresponding to this scenario; This represents the set of core fields corresponding to each table; This describes the relationships between multiple tables. This describes the business semantics of the scenario. This indicates a description of the conditions that need to be filtered, such as filtering by data time.
[0031] Step 120: Adopt a scenario parsing and data mapping method driven by scenario knowledge configuration to automatically complete scenario parsing, data resource location and data acquisition SQL generation based on user input.
[0032] To achieve automatic mapping from business scenarios to data resources, a scenario data mapping tool was designed. This tool, acting as the data acquisition unit within the intelligent agent, automatically completes data resource location, data acquisition SQL generation, and SQL validity verification based on scenario knowledge configuration.
[0033] Step 121: Define the scene data mapping tool description.
[0034] The purpose of defining a tool description is to enable the intelligent agent to understand the tool's capabilities, so as to automatically select and invoke the tool. The tool description is defined as: "Scenario data mapping tool, which receives power business scenarios input by the user and returns the corresponding business data query SQL statements."
[0035] Step 122: Define the scene data mapping tool input.
[0036] To ensure that the scene analysis results can be accurately mapped to the corresponding underlying data resources, this tool defines the following input parameters: scene input, time input, other filtered inputs, and original requirement input, to fully understand user issues. The information contained herein: ; in, Indicates the time range included in the user's question; This indicates other filter criteria in the user's question; This indicates the original user question; This indicates the name of the scenario to which the current problem belongs. To ensure the accuracy of the mapping, the scenario parameter must be limited to the scenario defined in step 110, i.e.: ; This represents the complete set of business scenarios defined in the scenario knowledge configuration library constructed in step 110.
[0037] Based on the above parameter definitions, and with the help of the large model and intelligent agent framework, the large model can be guided to understand the user's natural language needs and effectively transform them into structured input parameters.
[0038] Step 123: Design the scenario data acquisition Prompt template.
[0039] To address the issue of large language models easily generating misinterpretations due to a lack of knowledge of the data structure of internal business systems, a scenario-based data acquisition prompt template was designed. This template injects scenario knowledge configuration, database schema, and user requirements into the model context. Furthermore, by setting strict SQL specifications and error correction rules, it guides the model to accurately generate the data acquisition SQL. See details below. Figure 3 .
[0040] Scene data acquisition Prompt template defined as : ; in, This represents the schema content of the database table, used to provide the model with information about the table structure, field definitions, and field attributes. This represents specific information about a scenario defined in step 110; Indicates the current time, used to process relative time expressions in user input; Indicates the time range included in the user's question; This indicates other filter criteria in the user's question; This indicates important reminders related to scene data acquisition; This indicates the original user question.
[0041] Step 124: Design an SQL validity verification mechanism.
[0042] To avoid generating SQL statements with missing fields, incorrect table names, or invalid joins in large language models, this invention designs an SQL validity verification mechanism, including two phases: pure syntax verification and multi-round self-correction. (1) Pure syntax checking: The candidate SQL generated by the large language model is wrapped into a subquery, with an outer WHERE 1=0 constant false condition. The principle is that the database engine will fully parse the syntax structure of the subquery, but because the WHERE condition is always false, it will not trigger actual data scanning and return. The execution cost is extremely low. It is suitable for fast syntax verification of SQL generated by the large language model. It can detect common generation errors such as missing fields, incorrect table names, and undefined aliases without consuming data query resources.
[0043] (2) Multi-round self-correction mechanism: When the SQL fails syntax validation, the system automatically retrieves the error information returned by the database, extracts and normalizes the error content, and appends the error reason as feedback information to the model context. The system then calls the large language model again to regenerate the SQL. Based on the previous generation results and the database feedback information, the model specifically corrects the errors in the SQL and re-enters the validation process.
[0044] The maximum number of iterations is preset (3 times in this invention). When the SQL passes the validation, the iteration ends immediately and the result is output. If the maximum number of iterations is reached but the validation still fails, the current SQL generation is determined to have failed, and the corresponding failure status and error information are returned.
[0045] Step 125: Define the output of the scene data mapping tool.
[0046] Once the SQL statement passes validation, the tool does not directly execute the data query and return the full data. Instead, it outputs the validated SQL statement directly to avoid the performance overhead caused by transmitting large amounts of information and the limitations of large model context. When the tool fails to execute a task, the output includes the task status and the reason for failure, which helps the agent to handle exceptions and make adjustments.
[0047] Step 200: Preprocessing scheme construction: Pre-construct a raw power data preprocessing method flow to transform low-quality raw power data into high-quality rule mining data.
[0048] Raw power data commonly suffers from missing values and outliers. Directly using it for rule mining can easily lead to rule bias, increased noise, and decreased model stability. To improve the quality of rule mining data, this invention differentiates missing and outlier values based on data type and quality characteristics, and discretizes continuous values, providing a high-quality data foundation for subsequent deep pattern mining. The specific steps are as follows: Step 210: Use a missing value processing method based on both column type and missing rate to perform hierarchical processing of missing values on the original power data.
[0049] To address the issue of numerical and categorical columns having different missing characteristics, a missing value handling method based on dual criteria of column type and missing rate is designed. According to the data type and missing ratio of the field, the corresponding missing value handling strategy is adaptively selected to ensure data integrity while preserving the statistical distribution characteristics and business semantic information of the data as much as possible.
[0050] For dataset D any field First, calculate its missing rate. : ; in, Representation field The number of missing values, This indicates the total number of records in the dataset.
[0051] Different processing strategies are adopted based on the field type and the magnitude of the missing data: (1) For numeric fields, perform stratification based on the field missing rate.
[0052] Based on data integrity requirements, a first threshold α1 and a second threshold α2 are set. When the proportion of missing data in a field is high, the field contains little valid information, making it difficult to support rule pattern mining, and therefore, it is deleted. When a field has a certain proportion of missing data but still contains some valid information, the data existence characteristic of the field is retained. When the proportion of missing data in a field is low, statistical imputation is used to restore the field's integrity. Specifically, this includes: when If the field is deemed to lack sufficient valid information, it will be deleted directly. when When this happens, the field is converted to a binary value of "NULL / NOT NULL" to retain only the field's existence information; when When missing values are filled, the median is used. The median is calculated as follows: ; in, ( () represents the median calculation function; This represents the calculated median value.
[0053] The first threshold α1 and the second threshold α2 are determined according to the requirements of the rule mining task for data integrity and feature retention capability, and are used to balance the field information retention rate and data reliability. In this invention, α1=0.8 and α2=0.5 are set.
[0054] (2) For categorical fields, if the missing rate exceeds the preset threshold (set to 0.8), the field is deleted directly; otherwise, "NULL" is used as the missing value identifier to maintain the integrity of the categorical information.
[0055] Step 220: Use a Z-score-based outlier detection method to identify and filter outliers in the numeric fields after missing value processing.
[0056] The purpose of this step is to remove extreme outliers that deviate from the overall data distribution, thereby improving the data quality for subsequent rule mining. The specific process is as follows: Step 221: Perform Z-score standardization on the numeric fields after stratification of missing values to obtain standardized statistical values.
[0057] For each numeric field after the missing value stratification in step 210, calculate the mean and standard deviation of the numeric field, and convert each record in the numeric field into a standardized statistical value. Let the first k fields The Middle i The value of each record is The standardization calculation method is as follows: ; in, and Representing fields respectively C k The mean and standard deviation.
[0058] When the standard deviation of a field is zero, it means that the field has a constant value and is not included in the anomaly detection.
[0059] Step 222: Filter out abnormal data based on standardized statistical values and preset anomaly detection thresholds.
[0060] Let the anomaly detection threshold be... In this invention, =3, meaning outlier detection is performed using the 3-standard-deviation principle. For each data record, if any of its numeric fields satisfies: ; If the record contains an outlier, it will be removed from the dataset; otherwise, the record will be retained as normal data, forming the outlier-filtered dataset.
[0061] Step 230: Discretize the numerical fields using a supervised optimal binning method.
[0062] After missing value processing and outlier filtering, the numerical fields in the data remain continuous variables. Directly using continuous values for rule mining not only easily leads to an excessive number of candidate rules but also makes it difficult to form rule intervals with clear business guidance. Therefore, this invention employs a supervised optimal binning method to discretize continuous values, thereby better leveraging the advantages of pattern rule mining algorithms that are better suited to handling discrete attributes. The specific steps are as follows: Step 231: Based on each numeric field after outlier identification and filtering, convert the corresponding business field value into a label category, combine the numeric field value with the business category label to form a supervised learning dataset; use the business category label as a supervisory variable to guide binning boundary optimization.
[0063] Unlike traditional unsupervised discretization methods, this invention utilizes the service tag information of power data as a supervisory variable to guide the optimization of bin boundaries. The supervisory variable adopts user category or electricity consumption category, and the resulting numerical ranges more closely reflect actual service distribution, thus improving the range's ability to distinguish service patterns.
[0064] For each numeric field processed in step 220 k The corresponding business field values are converted into label categories (e.g., 0, 1), and then the field values are combined with the business category labels to form a supervised learning dataset. ; in, Indicates the first k The first field i Values can be selected; This indicates the supervision label for the corresponding sample.
[0065] Step 232: Generate candidate binning intervals based on the distribution of valid sample data. Combine the distribution of the target variable in each candidate binning interval with the optimization objective function to iteratively optimize the binning boundary and obtain the optimal binning boundary.
[0066] The optimization objective function can be expressed as: ; ; in, Represents the optimal set of bin boundaries; Represent the objective function; Indicates the first k Each container compartment; m Indicates the number of boxes; D Indicates the first k Evaluation function for the discriminative power of the target variable within each binning interval.
[0067] By solving the objective function, the final set of bin boundaries is obtained as follows: ; in, For the first i The optimal split point; This represents the final set of bin boundaries. m Indicates the number of boxes.
[0068] Step 233: Construct a complete binning boundary sequence based on the obtained optimal binning boundary, approximate the binning boundary to generate the final binning interval, and use the final binning interval to map the numerical field to complete the discretization of the numerical field.
[0069] Based on the obtained optimal bin boundaries, construct the complete bin boundary sequence: ; Considering that the split points obtained by supervised binning contain a large number of high-precision values, directly using them as rule conditions will reduce the readability of the rules and may lead to a decrease in the generalization ability of the rules due to overly fine boundaries.
[0070] Therefore, the bin boundaries are approximated by mapping the dividing points to business-meaning integer boundaries with a preset precision, preferably using a hundreds-digit rounding process. ; in, Represents the approximate bin boundaries; Round( The ) indicates rounding to the specified precision; p indicates the precision to retain, preferably to the hundreds place. For example, when p=100, 814 is approximately 800, and 894 is approximately 900.
[0071] Subsequently, the final binning sequence is constructed based on the approximated boundaries: ; The final container compartments are generated according to the principle of left-open and right-closed: ; in, Indicates the final binning sequence The first in k One boundary; Indicates the first k Each container compartment.
[0072] The original numerical fields are mapped using the final binning intervals, converting continuous values into corresponding interval labels, thus completing the discretization process of the numerical fields and obtaining discretized data.
[0073] Step 300: Rule generation based on deep pattern mining: Construct a rule generation tool that adopts a rule generation method that integrates data preprocessing with a deep autoencoder. Based on the data acquisition SQL, the tool acquires data and sequentially performs preprocessing, deep pattern learning, and association rule extraction processes to generate a candidate rule set.
[0074] Traditional association rule mining methods are prone to combinatorial explosion problems in high-dimensional data scenarios such as anti-electricity theft and illegal electricity use. Therefore, this invention employs a deep autoencoder as the rule mining model. By learning the latent feature representations of business data, it discovers hidden business patterns and constructs a rule generation tool. This tool encapsulates data acquisition, preprocessing, deep pattern mining, and association rule extraction into a unified execution process, which is automatically invoked by the intelligent agent to achieve automatic generation of candidate rules. The specific steps are as follows: Step 310: Build a rule generation tool.
[0075] Step 311: Define the description of the rule generation tool.
[0076] The rule generation tool is described as follows: "A candidate rule generation tool that acquires the SQL statements output by the scene data mapping tool, executes the data acquisition SQL, performs preprocessing, and executes the algorithm to generate candidate rules." This tool description enables the agent to understand the applicable scenarios and, after completing scene parsing and data mapping, automatically selects and invokes the tool based on the task planning results.
[0077] Step 312: Define the input for the rule generation tool.
[0078] The input parameters for the rule generation tool are defined as SQL statements, rather than directly transmitting business data, thus avoiding large-scale data transfer. The SQL is derived from the scenario data mapping tool and automatically passed in by an intelligent agent that identifies the context.
[0079] Step 313: Define the rule generation tool process.
[0080] After receiving the input SQL, the rule generation tool first establishes a connection with the database and executes the data retrieval SQL; then, it processes the data according to the preprocessing scheme constructed in step 200; based on this, it uses a deep autoencoder to learn the potential association patterns in the data and generate candidate rules. The specific execution steps are shown in steps 320 to 340.
[0081] Step 314: Define tool output.
[0082] After rule mining is completed, the tool will persistently store the generated candidate rule set in the form of a file and return the rule file storage path and execution status information. When the tool encounters an error, the output will also include the corresponding error description.
[0083] Step 320: Data Acquisition and Validation: Establish a connection with the target database based on the data acquisition SQL, execute the data acquisition SQL to obtain the required business data, and perform validity validation on the business data.
[0084] Based on the data generated in step 100, an SQL query is executed to establish a connection with the target database, retrieve the business data required for rule mining, and then the data is validated, specifically including: (1) Minimum data volume verification. When the number of samples is lower than the preset threshold (2000 in this invention), it is considered that the data is insufficient to support reliable pattern learning, the current rule generation task is terminated, and the corresponding prompt information is returned.
[0085] (2) Data scale control. For data with a sample size exceeding the preset upper limit (500,000 in this invention), random sampling is used to compress the data, thereby controlling the computational scale and improving the efficiency of rule mining while ensuring the representativeness of the samples.
[0086] Step 330: Preprocessing scheme execution: Perform data preprocessing on the verified business data.
[0087] Following the preprocessing scheme constructed in step 200, missing value processing, outlier filtering, and numerical discretization are performed sequentially on the acquired data to provide reliable data input for subsequent processing.
[0088] Step 340: Deep pattern mining and rule extraction: Utilize a deep autoencoder to learn the association patterns in the latent representation space of the preprocessed data, namely association features and behavioral patterns, to form a set of candidate rules.
[0089] Traditional association rule mining methods mainly rely on frequent itemset enumeration to generate candidate rules. When the feature dimension is high, the number of candidate combinations grows exponentially, easily leading to combinatorial explosion and making it difficult to learn nonlinear relationships in complex business data. This invention employs a rule mining and extraction method based on a deep autoencoder. It learns the latent association patterns in power data through a deep neural network and converts the implicit knowledge learned by the network into interpretable association rules, thereby generating candidate rules. The specific steps are as follows: Step 341: Data Encoding: The discretized data is converted into one-hot encoded form.
[0090] Convert the discretized data obtained in step 330 into one-hot encoded form. Assume the current data contains F categorical fields, and the... j The fields contain Each record can take one value and is represented as follows: ; ; in, x This represents a data record after one-hot encoding conversion; Indicates the data dimension.
[0091] Step 342: Construct a deep autoencoder and train it.
[0092] An undercomplete deep autoencoder network consisting of an encoder and a decoder is constructed to map high-dimensional one-hot encoded data to a low-dimensional latent feature space, and to learn the internal relationships of the data by reconstructing the data. Its forward propagation process is represented as follows: ; ; Where z is the low-dimensional latent feature representation learned by the encoder; The output vector reconstructed by the decoder; , These represent the weight matrices of the encoder and decoder, respectively. , Each represents its respective bias value; ( ) and ψ( ) is a non-linear activation function.
[0093] The model training uses the reconstruction error between the input data and the reconstruction result as the optimization objective, and employs the binary cross-entropy loss function as the objective function. ,Right now: ; in, Indicates data dimension, Represents the first of the original inputs. i Each dimension value; Indicates the first output of the decoder i Each dimension value.
[0094] By continuously optimizing network parameters to minimize the objective function value, i.e., the reconstruction error, we can learn stable behavioral patterns and potential correlations in the data.
[0095] Step 343: Rule extraction: Use the trained deep autoencoder to perform inference analysis on the preprocessed data and extract candidate rules from the latent feature representation.
[0096] After model training is completed, inference analysis is performed on the trained deep autoencoder to extract explicit pattern-related rules from the latent feature representation.
[0097] First, construct the test input vector to be analyzed. One or more feature values to be analyzed are used as candidate rule prefixes, with their corresponding positions set to the active state. The remaining features remain in their initial states. This vector is then fed into the trained autoencoder for inference, outputting the predicted probability of each feature value. The calculation process can be represented as follows: ; in, For the test input vector; These are the probabilities of each feature value output by the network. , These are the encoder and the decoder, respectively.
[0098] Based on the output analysis, if the implied probability of the candidate antecedent is higher than the preset pattern frequency threshold α (α=0.1 in this invention), then the feature combination is considered to have a high probability of pattern occurrence and can be used as the antecedent of the rule. Subsequently, the remaining features are traversed. If their predicted probability is higher than the rule strength threshold β (β=0.6 in this invention), and they do not belong to the same business field as the antecedent, then the feature is considered to have a stable correlation with the antecedent of the rule, and is determined as the consequent of the rule, forming a candidate rule. A→C ; Where A represents the preceding rule and C represents the following rule.
[0099] By iterating through different feature combinations and repeating the above reasoning process, candidate rules covering different patterns are generated, ultimately resulting in a candidate rule set. R : ; in, Indicates the first i Candidate rules; n Indicates the number of rules.
[0100] Step 350: Rule quality assessment: Evaluate and filter the candidate rules to obtain the final set of candidate rules.
[0101] To ensure the high reliability and business value of the generated rules, candidate rules are evaluated and filtered. This invention uses lift and Zhang correlation as rule selection indicators. Lift measures the actual correlation between the preceding and following terms of a rule; a lift greater than 1 indicates a positive correlation between the preceding and following terms. The calculation formula is as follows: ; ; in, Indicates the confidence level of the rule; Indicates the elevation of the rule. This indicates the number of times the preceding and following terms of a rule appear simultaneously in the dataset. and This represents the total number of occurrences of the preceding and following terms of the rule; N represents the total number of data points in the dataset.
[0102] Zhang correlation degree comprehensively considers rule coverage and random co-occurrence factors, is less affected by marginal probability, and avoids the randomness brought about by relying solely on lift degree. Its calculation formula is: ; Based on the set evaluation threshold, only candidate rules that satisfy the minimum lift (1.01 in this invention) and Zhang correlation (0.2 in this invention) are retained to form the final candidate rule set. : ; Where m represents the number of rules remaining after filtering.
[0103] Step 400: Rule Denoising and Semantic Structuring: Construct a rule denoising and summarizing tool, which adopts a rule denoising and semantic structuring method based on a large language model to transform candidate rules into structured rules with practical business meaning.
[0104] To address the issues of noisy, redundant, and low-value rules in the candidate rule set generated by deep pattern mining, this invention constructs a rule denoising and summarizing tool, designs a rule denoising and summarizing Prompt template, and utilizes a large language model to perform business semantic analysis on the candidate rule set. This achieves rule validity screening, redundant rule elimination, similar rule merging, and structured semantic expression of rules, transforming the candidate rules obtained from deep pattern mining into structured rules with clear business meaning, providing business personnel with high-quality and easily understandable rule results. The specific steps are as follows: Step 410: Define the rule-based noise reduction and summary tool.
[0105] The rule denoising and summarizing tool receives the candidate rule set generated during the deep pattern mining stage and, combined with business scenario knowledge, performs rule filtering, semantic understanding, and structure transformation through a large language model. The specific definition is as follows: Step 411: Define the rule-based noise reduction summary tool description.
[0106] The rule denoising and summarizing tool is defined as: "Obtaining the results of the rule generation tool, filtering, summarizing and refining candidate rules, and generating the final effective rules." Based on the above tool description, after the rule generation tool completes, the intelligent agent can automatically invoke this tool according to the task flow.
[0107] Step 412: Define rules for noise reduction and summarize the tool input.
[0108] The input parameters for the rule denoising and summarizing tool are defined as the business scenario name and the candidate rule file identifier. The business scenario defines the business context for rule analysis, guiding the large language model to understand the rules using corresponding domain knowledge. The candidate rule file identifier is used to read the candidate rule set file, also to avoid data transfer issues when there are too many candidate rules.
[0109] Step 413: Design a rule-based noise reduction summary Prompt template.
[0110] Since large language models lack business data structures and related knowledge, they cannot directly determine the business rationality of candidate rules. Therefore, a rule denoising summary Prompt template was designed. The business scenario, table and field structure information, and rule set are uniformly injected into the model context, guiding the model to complete rule analysis as an expert in anti-electricity theft and illegal electricity use. Specific details are as follows: Figure 4 As shown. The rule-based denoising summary Prompt template can be defined as: ; in, This indicates the name of the current inspection scenario, used to limit the business context of the rule-based noise reduction. It represents the data table and field structure information, and is used to provide the model with data table field definitions and field attributes; Important reminders regarding rule-based noise reduction summary; This represents the set of candidate rules generated in step 300.
[0111] Based on the above rules for denoising and summarizing the Prompt template, the large language model performs the following processing on the candidate rules: (1) Rule-based noise reduction and validity screening: Based on business scenarios and field semantics, a rationality analysis is conducted on candidate rules, filtering out low-value rules that lack business logic support, are caused by data quality issues, have logical conflicts, or are duplicated or included. For rules with inclusion relationships, rules with finer business granularity and stronger interpretability are prioritized for retention.
[0112] (2) Rule-based structure generation: The filtered candidate rules are converted into a unified structured format, including rule name, rule text, and rule description. The rule text expresses the rule logic in the form of "condition item → result item," while the rule description explains the business implications and application value.
[0113] Step 414: Define rules for denoising and summarize the tool output.
[0114] After successful execution, the tool will persistently store the generated set of structured rules as a file and return the rule file storage path, the number of valid rules, and execution status information. If the tool encounters an error, the output will also include a corresponding error description.
[0115] Step 500: Rule-based intelligent agent construction: Based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, construct an intelligent agent for generating rules for investigating illegal electricity theft.
[0116] This invention constructs a rule-generating intelligent agent for electricity theft investigation scenarios. Using a large language model as its inference engine, it uniformly registers the scenario data mapping tool, rule generation tool, and rule denoising and summarizing tool constructed in steps 100 to 400 as tools that the intelligent agent can recognize and invoke. Based on the user's natural language description requirements, the large model drives the intelligent agent to automatically complete tool selection, parameter passing, and task scheduling, achieving automatic execution of the entire rule generation process. The specific steps are as follows: Step 510: Tool registration.
[0117] The scene data mapping tool built in step 100, the rule generation tool built in step 300, and the rule denoising and summarizing tool built in step 400 are uniformly registered in the agent tool library, including tool name, tool description, and input and output parameters, so that the agent can automatically identify the applicable scenarios of each tool based on the tool description.
[0118] Step 520: Design the agent Prompt template.
[0119] To enable the agent to autonomously complete task planning, tool selection, and process execution based on user-described needs in natural language, a task execution prompt template is designed. This template integrates the agent's task objectives, tool capabilities, and database resources into the context of the large language model, guiding the agent to complete the rule-generated task step-by-step according to the process. See details below. Figure 5 The agent's Prompt template can be represented as: ; in, G Indicate the task objective; This indicates an important reminder that the intelligent agent has completed its task; The tool space represents the tools constructed in steps 100, 300, and 400, along with their invocation instructions. D This indicates database resource information, used to connect to the database and retrieve data. Q This indicates the original user question.
[0120] Step 530: Task execution and result output.
[0121] After receiving user input, the agent performs task analysis according to the agent Prompt template designed in step 520, identifying business scenarios, time conditions, and other content. Subsequently, the agent sequentially calls the scenario data mapping tool, rule generation tool, and rule denoising and summarizing tool, using the output of the previous tool as the input of the next tool, thus achieving automatic connection between processes such as data acquisition SQL generation, candidate rule generation, rule denoising, and semantic structuring.
[0122] After the rule-based denoising and summarization tool completes its execution, it outputs the final structured set of rules, which can be represented as: ; ; in, Indicates the first i Structured rules, , , These represent the rule name, rule text, and rule description, respectively.
[0123] Upon completion of all tasks, the agent outputs structured rules, the storage path of the rule files, and execution status information. If an exception occurs during execution, it returns the corresponding error message, providing the user with a complete list of task execution results.
[0124] Method and effect demonstration: Figure 2 It is a data mapping knowledge configured for anti-electricity theft and illegal electricity use scenarios. The illegal electricity use scenario involves two tables with more than 30 fields, and the anti-electricity theft scenario involves two tables with more than 10 fields. It also has additional information such as scenario description, table relationship conditions, and time filtering conditions to help the large model understand the scenario.
[0125] Based on this, the following example of rule generation in a scenario of electricity misuse will be used to demonstrate the effectiveness of the present invention.
[0126] After a user inputs "Retrieve all data from the past three years to generate rules for electricity violations," the intelligent agent identifies the task type and time range, automatically invoking a scenario data mapping tool. The tool's execution results show that it generated a data query SQL covering over 30 fields, including user profiles, electricity consumption, and electricity charges. This process automatically converts natural language requirements into business data retrieval SQL, reducing the cost of manual scenario analysis and SQL writing.
[0127] The tool receives the data query SQL generated in step 100 and automatically retrieves business data. The tool's execution results show that, based on the missing value handling and outlier filtering in the data preprocessing workflow, a total of 298,924 valid records and 16 valid fields were retained. Based on this data, 29 candidate rules were generated through deep rule mining and rule validity filtering, and saved as a rule file.
[0128] The tool reads the candidate rule file, inputs the business scenario, field semantics, and candidate rules into a large language model, and refines the 29 candidate rules into 3 high-quality structured audit rules through business rationality analysis, redundant rule elimination, and similar rule merging. This process effectively reduces noisy rules generated during the mining process and improves the business understandability and application value of the rule results.
[0129] As can be seen, after discretization, the original continuous values are transformed into interval features with discriminative capabilities. For example, one interval for total active power is 35-200, which aligns with business patterns. All rules are expressed in a unified structured format, including rule name, rule text, and rule description. For example, "When the comprehensive multiplier is low, mixed wiring corresponds to Class IV metering devices." The rule means: "When a user's comprehensive multiplier does not exceed 5 and a mixed wiring method is used, its metering device is usually classified as Class IV, which can be used to investigate risks such as abnormal metering records and abnormal equipment configurations." The generated rules are all derived from the statistical patterns of actual business data, and possess clear business semantics and executable conditions. Business experts have confirmed the usability of the rules.
[0130] Experiments show that this invention enables the intelligent agent to collaboratively call various tools. Users only need to input business requirements in natural language to automatically complete the entire process, including business scenario identification, data resource location, SQL generation, data processing, pattern mining, rule filtering, and semantic summarization, and finally output high-quality structured audit rules.
[0131] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for constructing an intelligent agent for investigating illegal electricity theft based on deep pattern mining, characterized in that, Includes the following steps: A scenario knowledge configuration and scenario data mapping tool is constructed between anti-electricity theft and illegal electricity use business scenarios and data resources. The scenario data mapping tool adopts a scenario parsing and data mapping method driven by scenario knowledge configuration, and completes scenario parsing, data resource location and data acquisition SQL generation according to user input. A rule generation tool is constructed, which adopts a rule generation method that integrates data preprocessing and deep autoencoder. Based on the data acquisition SQL, the tool acquires data and sequentially performs preprocessing, deep pattern learning and association rule extraction processes to generate a set of candidate rules. A rule denoising and summarization tool is constructed, which adopts a rule denoising and semantic structuring method based on a large language model to transform candidate rules into structured rules with practical business meaning; Based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, an intelligent agent is constructed for generating rules for investigating illegal electricity theft.
2. The method for constructing an intelligent agent for investigating illegal electricity theft based on fusion deep pattern mining as described in claim 1, characterized in that, Scene parsing and data mapping methods based on scene knowledge configuration drive include: Based on scenario knowledge configuration, a scenario data acquisition Prompt template is designed, which injects scenario knowledge configuration, database schema and user requirements into the model context in a unified manner, and guides the model to complete the generation of data acquisition SQL through an SQL validity verification mechanism.
3. The method for constructing an intelligent agent for investigating illegal electricity theft based on fusion deep pattern mining as described in claim 2, characterized in that, Scene data acquisition Prompt template Defined as: ; in, Represents the schema content of a database table; Represents specific information for a particular scene; Indicates the current time; This indicates important reminders related to scene data acquisition; Indicates the time range included in the user's question; This indicates other filter criteria in the user's question; This indicates the original user question.
4. The method for constructing an intelligent agent for investigating electricity theft rules based on fusion deep pattern mining as described in claim 1, characterized in that, Rule generation methods that integrate deep autoencoders with data preprocessing include: After receiving the data retrieval SQL, a connection is established with the target database based on the data retrieval SQL, the data retrieval SQL is executed to obtain the required business data, the validity of the business data is verified, and the verified business data is preprocessed. The preprocessed data is converted into one-hot encoded form; a deep autoencoder is constructed and trained; the trained deep autoencoder is used to perform inference analysis on the preprocessed data to extract candidate rules; the candidate rules are evaluated and filtered to obtain the final set of candidate rules.
5. The method for constructing an intelligent agent for investigating electricity theft rules based on fusion deep pattern mining as described in claim 4, characterized in that, The data preprocessing includes: The original power data is subjected to hierarchical processing for missing values. The field types of the original power data include numeric fields and categorical fields. After processing for missing values, outlier identification and filtering are performed on the numeric fields, and the numeric fields are discretized.
6. The method for constructing an intelligent agent for investigating electricity theft rules based on fusion deep pattern mining as described in claim 5, characterized in that, A supervised optimal binning method is adopted to discretize numerical fields based on business target fields.
7. The method for constructing an intelligent agent for investigating electricity theft rules based on fusion deep pattern mining as described in claim 1, characterized in that, Rule-based denoising and semantic structuring methods based on large language models include: Based on the candidate rule set and combined with business scenario knowledge, a rule denoising summary Prompt template is designed. The candidate rules are denoised and their effectiveness is screened through a large language model. The screened candidate rules are then converted into a unified structured format, which is a structured rule with actual business meaning.
8. The method for constructing an intelligent agent for investigating illegal electricity theft based on fusion deep pattern mining as described in claim 7, characterized in that, Rule-based noise reduction summary Prompt template Defined as: ; in, Indicates the name of the current inspection scenario; This indicates the data table and field structure information; Important reminders regarding rule-based noise reduction summary; This represents the set of candidate rules.
9. The method for constructing an intelligent agent for investigating illegal electricity theft based on fusion deep pattern mining as described in claim 1, characterized in that, Based on scene data mapping tools, rule generation tools, and rule denoising and summarization tools, an intelligent agent for generating rules for investigating illegal electricity theft is constructed, including: The constructed scene data mapping tool, rule generation tool, and rule denoising and summarization tool will be uniformly registered in the intelligent agent tool library; Based on the agent tool library, an agent Prompt template is designed to guide the agent to call the scene data mapping tool, rule generation tool and rule denoising and summarizing tool in sequence, so as to realize the automatic connection between data acquisition SQL generation, candidate rule generation and rule denoising and semantic structuring. Once the rule denoising and summarizing tool has finished executing, it outputs the final structured rule set.
10. The method for constructing an intelligent agent for investigating illegal electricity theft based on fusion deep pattern mining as described in claim 9, characterized in that, Agent Prompt Template Defined as: ; in, G Indicate the task objective; This indicates an important reminder that the intelligent agent has completed its task; The toolspace represents the tool space, containing each tool and its calling description; D Represents database resource information; This indicates the original user question.