A distributed photovoltaic accounting exception analysis processing method and system
Patent Information
- Application Number
- CN202611299989.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本发明的目的是提供一种分布式光伏核算异常分析处理方法及系统,通过以发电户号为索引聚合形成结构化数据包,对发电分摊公式进行解析与校验,基于大语言模型与规则引擎进行异常判定及加权融合,并根据分级判定结果自动确定处置路径,以解决现有技术存在的核算异常依赖人工核查,缺乏确定性校验,且检出后需人工跨系统处置的问题
[0062]1、本发明通过以发电户号为索引聚合分布式光伏发电业务数据形成结构化数据包,对发电分摊公式进行解析与校验,并采用大语言模型与规则引擎分别进行异常判定与确定性硬校验修正,再将公式校验结果、经校验的异常项列表及规则判定结果进行加权融合与分级判定后根据字段归属确定处置路径,实现了核算异常从数据聚合、公式校验、多源融合判定到分级处置的全流程自动化处理,克服了现有技术中依赖人工核查、单一检测手段准确性不足、大语言模型推理结果缺少确定性校验机制以及异常检出后仍需人工跨系统完成处置操作的技术缺陷,提高了分布式光伏核算异常分析的效率和准确性。
Smart Images

Figure CN122796451A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for analyzing and processing anomalies in distributed photovoltaic accounting, belonging to the field of photovoltaic power generation operation and maintenance technology. Background Technology
[0002] In recent years, with the advancement of market-oriented reforms in renewable energy feed-in tariffs, distributed photovoltaic (PV) electricity billing has added several policy parameters, such as new and existing capacity indicators, mechanism-based tariffs, and mechanism-based electricity volumes, significantly increasing the complexity of anomaly detection. Due to the dispersed and complex data relationships between user files, metering point configurations, allocation formulas, and issuance fees, anomaly location requires repeated queries across multiple business systems, with single-household investigations taking several days. A large number of anomaly work orders are concentrated in provincial-level professional positions for processing, resulting in low efficiency, high error rates, and significant personnel pressure.
[0003] In terms of anomaly detection and localization technology, existing solutions include: rule screening methods that perform judgments on marketing data item by item by setting a fixed set of business rules, marking data that hits the rules as anomalies and outputting details. However, this method relies on manually predefined rules, resulting in incomplete rule coverage, requiring development and deployment processes for updates, and being unable to handle complex anomaly scenarios that require cross-field semantic understanding; and power marketing verification methods based on vector retrieval engines, which encode audit rule text and user profile text into vectors and store them in a knowledge base, retrieving relevant rules and performing judgments through vector similarity matching. However, in distributed photovoltaic business scenarios, the mapping relationship between database table names, field names, standard code tables, and natural language is difficult to accurately establish through general vector encoding, leading to a decrease in retrieval accuracy when facing scenarios with multiple knowledge points overlapping, resulting in increased understanding bias of the large language model. Public literature introduces the exploration idea of applying the large language model to electricity billing anomaly screening, but does not propose a fusion mechanism between the large language model inference results and the programmatic rule verification, nor does it involve a closed-loop handling process after anomaly detection. The following technical problems still need to be solved in the existing technology: First, the reasoning results of large language models lack a deterministic verification mechanism; Second, after anomaly detection, manual cross-system processing is still required, and there is a lack of automated connection from detection to resolution. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for analyzing and processing anomalies in distributed photovoltaic accounting. By aggregating structured data packets using the generator number as an index, the system parses and verifies the generator allocation formula, performs anomaly judgment and weighted fusion based on a large language model and rule engine, and automatically determines the handling path based on the hierarchical judgment results. This solves the problems of existing technologies where accounting anomalies rely on manual verification, lack deterministic verification, and require manual cross-system handling after detection.
[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution:
[0006] In a first aspect, the present invention provides a method for anomaly analysis and processing in distributed photovoltaic accounting, comprising:
[0007] Distributed photovoltaic power generation business data is aggregated using the power generation account number as an index to form a structured data package;
[0008] The power generation allocation formula in the structured data packet is parsed and verified to generate a formula verification result including anomaly markers;
[0009] The structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base are input into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, the corresponding field values are read from the structured data packet to perform anomaly judgment and generate a list of anomalies in the large language model.
[0010] Based on the rule engine, deterministic hard validation rules are independently executed on the structured data packets to generate rule judgment results. The rule judgment results are then used to validate and correct the list of anomalies in the large language model to obtain a validated list of anomalies.
[0011] The formula verification results, the verified list of anomalies, and the rule judgment results are weighted and fused to calculate the fusion confidence of each anomaly. The fusion confidence of each anomaly is then graded according to a preset threshold and a hard error priority rule. The output includes an anomaly analysis list that includes anomalies marked with confidence level.
[0012] The handling path is determined based on the field attribution of the anomaly item marked with the confidence level in order to process the anomaly in the distributed photovoltaic accounting.
[0013] Furthermore, the distributed photovoltaic (PV) power generation business data includes at least customer profile data, power generation project profile data, metering point configuration data, issuance fee data, and historical snapshot data; wherein, the distributed PV power generation business data is aggregated using the power generation account number as an index to form a structured data package, including:
[0014] Data on distributed photovoltaic power generation is obtained based on the power generation account number and then aggregated to obtain an aggregated data package.
[0015] Organize the data in JSON format, and simultaneously append English field text and Chinese field tags to each field in the aggregated data packet to obtain a formatted data packet;
[0016] The historical snapshot data in the formatted data packet is compared field by field with the current distributed photovoltaic power generation business data to generate comparison difference markers. The comparison difference markers are then written into the formatted data packet to obtain a structured data packet.
[0017] Furthermore, the generation allocation formula in the structured data packet is parsed and verified to generate a formula verification result including anomaly markers, including:
[0018] Based on regular expression segmentation, the code string corresponding to the power generation allocation formula in the structured data packet is parsed using the recursive descent parsing method to identify the metering point code, field type, operator and parentheses in the Token;
[0019] According to the preset metering point mapping table, the machine identifier corresponding to the metering point code is replaced with a readable variable, while retaining the operators and parentheses, and a mathematical expression is generated;
[0020] The mathematical expression is sequentially subjected to bracket pairing verification, denominator non-zero verification, metering point attribution verification, and field type verification. Specifically, the bracket pairing verification checks the symmetry and nesting validity of brackets based on the bracket type in the Token; the denominator non-zero verification checks the validity of the denominator in the division operation based on the electricity values of each metering point participating in the allocation in the structured data packet; the metering point attribution verification checks whether all metering points referenced in the mathematical expression belong to the current generator account number based on the metering point code in the Token; and the field type verification checks whether there are incompatible cross-references or substitution operations between the total active power and the forward active power based on the field type in the Token.
[0021] Mathematical expressions that fail any of the following checks—bracket pairing, non-zero denominator, measurement point attribution, and field type—are marked as anomalies. The corresponding check type and the location information determined based on the Token parsing location are recorded. The anomaly type of the anomaly is determined according to the check type, and a formula verification result containing the anomaly type and location information is generated.
[0022] Furthermore, based on the natural language rules in the verification rule knowledge base and the field mapping relationships in the mathematical model mapping knowledge base, corresponding field values are read from the structured data packet for anomaly detection, generating a list of language model anomalies, including:
[0023] Based on the digital model mapping knowledge base, business terms in natural language are mapped to specific database tables and fields, and state values described in natural language are mapped to corresponding standard code values to obtain field mapping relationships;
[0024] Based on the field mapping relationship, the field values corresponding to each natural language rule in the verification rule knowledge base are read from the structured data packet, and the field values are judged to be abnormal to obtain a preliminary list of abnormal items;
[0025] Based on the structured data packet, the field mapping relationship of the numerical model mapping knowledge base, and the natural language rules of the verification rule knowledge base, prompt words are generated;
[0026] Driven by the prompt words, the large language model parses the structured data packet based on the preliminary anomaly list and the generator number as the index. It then rewrites the natural language rules into data query statements according to the digital model mapping knowledge base. The corresponding field values in the structured data packet are confirmed through the field mapping relationship, and the queried field values are analyzed using rule-based reasoning based on the natural language rules in the verification rule knowledge base, generating a large language model anomaly list.
[0027] Furthermore, based on the rule engine, deterministic hard validation rules are independently executed on the structured data packets to generate rule judgment results. These results are then used to validate and correct the large language model's anomaly list, resulting in a validated anomaly list, including:
[0028] According to the preset deterministic hard validation rules, the rule engine independently performs deterministic logic validation on each field in the structured data packet to generate rule judgment results; the rule judgment results include the pass flag, fail flag and corresponding exception flag for each deterministic hard validation rule;
[0029] The rule determination result is compared item by item with the list of anomalies in the large language model, and a correction step is performed to obtain a verified list of anomalies.
[0030] The correction steps include:
[0031] If a certain deterministic hard validation rule in the rule judgment result is judged to be unsuccessful, and the corresponding item in the large language model anomaly list is not marked as an anomaly, then the corresponding anomaly item is added to the large language model anomaly list.
[0032] If a deterministic hard validation rule in the rule determination result is determined to pass, and the corresponding item in the large language model anomaly list is marked as an anomaly, then the anomaly mark of the corresponding anomaly item is removed from the large language model anomaly list.
[0033] Furthermore, the formula validation results, the validated list of outliers, and the rule-based judgment results are weighted and fused to calculate the fusion confidence of each outlier. Based on a preset threshold and a hard error priority rule, the fusion confidence of each outlier is graded and judged, outputting an anomaly analysis list including outliers marked with confidence levels, including:
[0034] Based on the formula verification results, the verified list of anomalies, and the rule judgment results, the verification rules are divided into deterministic rules of status archive, rules of power value calculation, and rules of cross-field semantic understanding. The large language model weight, rule engine weight, and formula verification weight are configured for each verification rule to obtain the weight configuration corresponding to each verification rule.
[0035] Based on the weight configuration corresponding to each verification rule, the formula verification result, the list of verified anomalies, and the judgment value corresponding to each rule judgment result are obtained, and the fusion confidence of each verification rule is calculated according to the preset weighted summation formula; when the judgment value is 1, it is judged as an anomaly, and when the judgment value is 0, it is judged as normal.
[0036] The fusion confidence level of each anomaly is graded and determined according to a preset threshold and a hard error priority rule, wherein the hard error priority rule determines that the rule result is an anomaly.
[0037] If the fusion confidence reaches a preset threshold, or the rule judgment result is abnormal, or the formula verification result, the verified list of abnormal items, and the judgment result of the rule judgment are all abnormal, and the fusion confidence plus a preset boost value reaches the preset threshold, then the abnormal item is determined to be a high-confidence abnormal item; otherwise, the abnormal item is determined to be a low-confidence item pending review.
[0038] Based on the classification judgment results, each anomaly in the anomaly analysis list is marked with a corresponding confidence level, and an anomaly analysis list including anomalies marked with confidence levels is output.
[0039] Furthermore, based on the field attribution of the anomaly item marked with the confidence level, a handling path is determined to process the distributed photovoltaic accounting anomalies, including:
[0040] Based on each anomaly marked with a confidence level, obtain the field attribution information corresponding to each anomaly; the field attribution information includes the archive field, the metering point configuration field, and the power generation allocation formula configuration field;
[0041] Based on the field attribution information, the abnormal items are classified and handled as follows:
[0042] If the field corresponding to the exception belongs to the archive field, it is classified as an archive exception and the archive maintenance process is triggered to handle the distributed photovoltaic accounting exception.
[0043] If the field corresponding to the anomaly belongs to the metering point configuration field, it is classified as a metering anomaly and a metering operation and maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly.
[0044] If the field corresponding to the anomaly belongs to the power generation allocation formula configuration field, it is classified as a formula-related anomaly, and a system maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly.
[0045] Furthermore, the network structure of the large language model includes:
[0046] An input encoding layer is used to receive an input text sequence and convert the input text sequence into a vector representation, wherein the input text sequence includes at least the structured data packet, the field mapping relationship in the digital model mapping knowledge base, and the natural language rules in the verification rule knowledge base;
[0047] The deep inference layer includes a multi-head self-attention module and a feedforward neural network module. The multi-head self-attention module is used to capture the dependencies between different positions in the vector representation, and the feedforward neural network module is used to perform a non-linear transformation on the output of the multi-head self-attention module to generate an enhanced semantic representation. The deep inference layer is used to perform inference operations according to the enhanced semantic representation, sequentially parsing the structured data packet with the generator number as an index, rewriting the natural language rules into data query statements, querying the corresponding field values in the structured data packet, and performing rule-based inference diagnosis on the queried field values.
[0048] The result generation layer is used to generate word sequence combinations of a large language model anomaly list by decoding word by word according to the enhanced semantic representation using an autoregressive generation method.
[0049] The output layer is used to combine the word sequence into a large language model anomaly list in text form and output it.
[0050] Furthermore, the training method for the large language model includes:
[0051] Based on historical distributed photovoltaic power generation business data and the corresponding historical verification anomaly list, a training sample set is constructed. The training sample set includes multiple training samples. The input data of each training sample includes structured data package samples, field mapping relationship samples in the digital model mapping knowledge base, and natural language rule samples in the verification rule knowledge base. The labels of each training sample are the corresponding anomaly list labels.
[0052] Based on the training sample set, the large language model is pre-trained using the cross-entropy loss function. The cross-entropy loss function is used to measure the difference between the word sequence output by the large language model and the labels of the anomaly list, so as to update the network parameters of the large language model and obtain the pre-trained large language model.
[0053] Based on the pre-trained large language model, a scenario verification sample set for the large language model in the distributed photovoltaic accounting anomaly analysis scenario is obtained, and the pre-trained large language model is fine-tuned. The fine-tuning training adopts a low-rank adaptive method, which freezes the pre-training parameters of the large language model and injects a trainable low-rank decomposition matrix to update the parameters of the low-rank decomposition matrix, thereby obtaining the fine-tuned large language model.
[0054] Secondly, the present invention provides a distributed photovoltaic accounting anomaly analysis and processing system, comprising:
[0055] The data aggregation module is used to aggregate distributed photovoltaic power generation business data using the power generation account number as an index, forming a structured data package;
[0056] The formula verification module is used to parse and verify the power generation allocation formula in the structured data packet, and generate a formula verification result including anomaly markers.
[0057] The language model determination module is used to input the structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, it reads the corresponding field values from the structured data packet to determine anomalies and generates a list of anomalies in the large language model.
[0058] The rule verification module is used to independently execute deterministic hard verification rules on the structured data packet according to the rule engine, generate rule judgment results, and use the rule judgment results to verify and correct the list of anomalies in the large language model to obtain a verified list of anomalies.
[0059] The fusion and grading module is used to weight and fuse the formula verification results, the verified list of anomalies, and the rule judgment results to calculate the fusion confidence of each anomaly. Based on the preset threshold and the hard error priority rule, the fusion confidence of each anomaly is graded and judged, and the output includes an anomaly analysis list including anomalies marked with confidence level.
[0060] The disposal execution module is used to determine the disposal path based on the field belonging to the anomaly item marked with the confidence level in order to handle the distributed photovoltaic accounting anomaly.
[0061] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0062] 1. This invention aggregates distributed photovoltaic (PV) power generation business data into structured data packages by using the power generation account number as an index. It parses and verifies the power generation allocation formula, and uses a large language model and a rule engine to perform anomaly detection and deterministic hard verification correction, respectively. Then, it performs weighted fusion and hierarchical judgment on the formula verification results, the list of verified anomalies, and the rule judgment results, and determines the handling path based on the field attribution. This realizes the fully automated processing of accounting anomalies from data aggregation, formula verification, multi-source fusion judgment to hierarchical handling. It overcomes the technical defects of existing technologies, such as reliance on manual verification, insufficient accuracy of single detection methods, lack of deterministic verification mechanism for large language model inference results, and the need for manual cross-system handling operations after anomaly detection. This improves the efficiency and accuracy of distributed PV accounting anomaly analysis.
[0063] 2. This invention aggregates customer profile data, power generation project profile data, metering point configuration data, issuance fee data, and historical snapshot data using the power generation account number as an index to form a structured data package. It then compares the historical snapshot data in the structured data package with the current distributed photovoltaic power generation business data field by field to generate comparison difference markers. Simultaneously, it parses the code string corresponding to the power generation allocation formula based on regular expression segmentation and recursive descent parsing methods. Through four-dimensional verification—bracket pairing verification, denominator non-zero verification, metering point attribution verification, and field type verification—it generates a formula verification result containing anomaly type and location information. This achieves structured integration of distributed photovoltaic accounting data and automated parsing and verification of the power generation allocation formula, overcoming the technical shortcomings of existing technologies where anomaly investigation requires repeated queries across multiple business systems and single-account investigation is time-consuming, thus improving anomaly location efficiency.
[0064] 3. This invention inputs structured data packets, a mathematical model mapping knowledge base, and a verification rule knowledge base into a large language model. The large language model then performs anomaly detection based on the natural language rules in the verification rule knowledge base and the field mapping relationships in the mathematical model mapping knowledge base, generating an anomaly list for the large language model. Simultaneously, based on the rule engine, deterministic hard verification rules are independently executed on the structured data packets. The rule detection results are used to supplement and correct the anomaly list of the large language model by adding new entries and removing anomaly markers, resulting in a verified anomaly list. This overcomes the technical deficiency of the large language model's inference results lacking a deterministic verification mechanism, reducing false positives and false negatives.
[0065] 4. This invention calculates the fusion confidence level of each anomaly by weighting and fusing the formula verification results, the verified anomaly list, and the rule judgment results. It then classifies the fusion confidence level of each anomaly according to a preset threshold and a hard error priority rule. Based on the classification judgment results, it marks the confidence level of each anomaly in the anomaly analysis list. Finally, based on the field affiliation of the anomaly marked with a confidence level, the anomaly is categorized into archive anomalies, measurement anomalies, or formula anomalies, triggering archive maintenance processes, measurement operation and maintenance work orders, or system operation and maintenance work orders respectively. This achieves fully automated connection from anomaly detection and confidence level classification to the triggering of the handling path, overcoming the technical defect in existing technologies where manual cross-system handling operations are still required after anomaly detection, thus reducing labor costs. Attached Figure Description
[0066] Figure 1 This is a flowchart illustrating a distributed photovoltaic accounting anomaly analysis and processing method provided in an embodiment of the present invention. Detailed Implementation
[0067] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0068] Example 1
[0069] like Figure 1 As shown in the figure, this embodiment introduces a method for anomaly analysis and processing in distributed photovoltaic accounting, including:
[0070] Step 1: Aggregate distributed photovoltaic power generation business data using the power generation account number as an index to form a structured data package.
[0071] This embodiment aggregates distributed photovoltaic (PV) power generation business data using the power generation account number as an index. The distributed PV power generation business data includes at least customer profile data, power generation project profile data, metering point configuration data, issuance fee data, and historical snapshot data. After aggregation processing to obtain an aggregated data package, it is organized in JSON format. English field text and Chinese field tags are added to each field. Then, the historical snapshot data and the current distributed PV power generation business data are compared field by field to generate comparison difference markers, which are written into the formatted data package to obtain a structured data package. This realizes the unified integration of heterogeneous data scattered in multiple business systems, overcomes the technical defects of existing technologies that require repeated queries across multiple business systems for anomaly location and are time-consuming for single-account investigation, and improves the efficiency of data acquisition and anomaly location.
[0072] Step 2: Parse and verify the power generation allocation formula in the structured data packet to generate a formula verification result including anomaly markers.
[0073] This embodiment uses regular expression segmentation and recursive descent parsing to parse the code string corresponding to the power generation allocation formula to identify the metering point code, field type, operators, and parentheses in the token. Based on the metering point mapping table, the machine identifier corresponding to the metering point code is replaced with a readable variable, while retaining the operators and parentheses to generate a mathematical expression. Then, the mathematical expression is sequentially subjected to parentheses pairing verification, denominator non-zero verification, metering point attribution verification, and field type verification. Mathematical expressions that fail any of the verifications are marked as anomalies, and the corresponding verification type and location information are recorded. The anomaly type of the anomaly is determined based on the verification type, and a formula verification result containing anomaly type and location information is generated. This realizes automated parsing and four-dimensional verification of the power generation allocation formula code, overcoming the technical defects of existing technologies where power generation allocation formulas are stored in code or script form and lack automated parsing and verification methods, making it difficult to detect formula configuration errors in a timely manner. This improves the detection rate and location accuracy of formula configuration errors.
[0074] Step 3: Input the structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, read the corresponding field values from the structured data packet to determine anomalies and generate a list of anomalies for the large language model.
[0075] This embodiment inputs structured data packets, a digital model mapping knowledge base, and a verification rule knowledge base into the large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationships in the digital model mapping knowledge base, it reads the corresponding field values from the structured data packets to determine anomalies and generates a list of anomalies for the large language model. The digital model mapping knowledge base stores the correspondence between database table names, field English names, field Chinese names, standard code values, and code meanings in a structured table format. It maps business terms in natural language to specific database tables and fields, and maps state values described in natural language to corresponding standard code values. This solves the technical problem that general vector encoding is difficult to accurately establish the mapping relationship between database table names, field names, and natural language, and improves the ability of the large language model to detect cross-field semantic understanding anomalies in distributed photovoltaic accounting scenarios.
[0076] Step 4: Based on the rule engine, independently execute deterministic hard validation rules on the structured data packets to generate rule judgment results. Use the rule judgment results to verify and correct the list of anomalies in the large language model to obtain a verified list of anomalies.
[0077] This embodiment uses a rule engine to independently execute deterministic hard validation rules on structured data packets to generate rule judgment results. These results are then used to validate and correct the large language model's anomaly list. When a deterministic hard validation rule fails in the rule judgment result but the corresponding item in the large language model's anomaly list is not marked as an anomaly, the corresponding anomaly item is added to the large language model's anomaly list. Conversely, when a deterministic hard validation rule passes in the rule judgment result but the corresponding item in the large language model's anomaly list is marked as an anomaly, the anomaly mark for the corresponding item is removed from the large language model's anomaly list. By using the rule engine's deterministic logic validation to correct the large language model's inference results, this approach overcomes the technical shortcomings of the large language model's inference results, which lack a deterministic validation mechanism and are prone to false positives and false negatives, thereby improving the accuracy and reliability of the anomaly list.
[0078] Step 5: The formula verification results, the verified list of outliers, and the rule judgment results are weighted and fused to calculate the fusion confidence of each outlier. The fusion confidence of each outlier is then graded according to the preset threshold and the hard error priority rule, and the output includes an anomaly analysis list that includes outliers marked with confidence level.
[0079] This embodiment configures adaptive weights for formula verification results, verified anomaly lists, and rule judgment results according to the verification rule type. It configures large language model weights, rule engine weights, and formula verification weights for deterministic rules of status archives, power value calculation rules, and cross-field semantic understanding rules, respectively. The fusion confidence level is calculated using a weighted summation formula, and then graded according to preset thresholds and hard error priority rules. High-confidence anomalies are distinguished from low-confidence anomalies awaiting review, achieving quantitative fusion and confidence grading of multi-source heterogeneous anomaly judgment results. This avoids misjudgments caused by insufficient accuracy of single detection methods and improves the credibility and interpretability of anomaly judgment results.
[0080] Step 6: Determine the handling path based on the field attribution of the anomaly item marked with the confidence level in order to process the distributed photovoltaic accounting anomaly.
[0081] This embodiment obtains the corresponding field attribution information based on each anomaly item marked with a confidence level in the anomaly analysis list. The field attribution information includes archive fields, metering point configuration fields, and power generation allocation formula configuration fields. If the field corresponding to the anomaly item belongs to the archive field, it is classified as an archive anomaly and the archive maintenance process is triggered. If it belongs to the metering point configuration field, it is classified as a metering anomaly and a metering operation and maintenance work order is triggered. If it belongs to the power generation allocation formula configuration field, it is classified as a formula anomaly and a system operation and maintenance work order is triggered. This realizes the automatic classification of anomalies by field attribution and the automatic triggering of the handling path. It overcomes the technical defects of the prior art, which still requires manual cross-system handling operations after anomaly detection and lacks the technical defects of automatic connection from detection to resolution, thus shortening the anomaly handling cycle.
[0082] Example 2
[0083] Based on the same inventive concept as Embodiment 1, this embodiment describes the implementation steps of a distributed photovoltaic accounting anomaly analysis and processing method, including:
[0084] Step 1: Aggregate distributed photovoltaic power generation business data using the power generation account number as an index to form a structured data package.
[0085] Step 1.1: Obtain distributed photovoltaic power generation business data based on the power generation account number, and perform aggregation processing to obtain aggregated data packets.
[0086] In this embodiment, using a specific generator account number as the query key, the external service interface of the power marketing business system is invoked to query the customer file table to obtain the account name, user status, user type, and voltage level; the generator project association table is queried to obtain the project name, registered capacity, grid connection date, increase / decrease flag, and mechanism electricity price; the metering point configuration table is queried to obtain the metering point number, metering point type, CT (Current Transformer) ratio, and PT (Potential Transformer) ratio; the metering point reading table is queried to obtain the total active and reactive power of each metering point; the issuance fee table is queried to obtain the on-grid electricity and issuance amount; and the generator allocation formula table is queried to obtain the underlying code text of the allocation formula; the above query results are aggregated into the same data packet.
[0087] Step 1.2: Organize in JSON format, and simultaneously attach English field text and Chinese field tags to each field in the aggregated data packet to obtain a formatted data packet.
[0088] In this embodiment, the fields in the aggregated data packet are organized in key-value pairs. Each field contains both an English field name and a Chinese field label, and the code value field is given a Chinese definition. For example, the user status field "09" is identified as "account closed", the user type field "02" is identified as "distributed photovoltaic", and the metering point type field "01" is identified as "power generation point".
[0089] Step 1.3: Compare the historical snapshot data in the formatted data packet with the current distributed photovoltaic power generation business data field by field, generate comparison difference markers, and write the comparison difference markers into the formatted data packet to obtain a structured data packet.
[0090] In this embodiment, the historical snapshot table is queried, and the snapshot archive of the previous issuance period is read. The snapshots are then compared field by field with the customer profile data and meter reading data of the current period. The comparison fields include at least the user status, user type, voltage level, and total active power of each meter point. The changed fields are recorded as difference items, which include the field name, the value before the change, the value after the change, and the change flag. All difference items are written into the header area of the formatted data packet in the form of a difference flag list to generate a structured data packet.
[0091] Step 2: Parse and verify the power generation allocation formula in the structured data packet to generate a formula verification result including anomaly markers.
[0092] Step 2.1: Based on regular expression segmentation, use the recursive descent parsing method to parse the code string corresponding to the power generation allocation formula in the structured data packet to identify the metering point code, field type, operator and parentheses in the Token.
[0093] In this embodiment, the underlying code string of the power generation allocation formula includes a conditional judgment function, a summation function, metering point identifiers, arithmetic operators, parentheses, and constants. Regular word segmentation combined with recursive descent parsing is used to scan the underlying code string, identifying conditional judgment keywords, summation function names, metering point codes, field type identifiers, arithmetic operators, left and right parentheses, numbers, and separators in the token sequence. The parser constructs the token sequence into an abstract syntax tree, where the root node of the abstract syntax tree is the conditional judgment, the conditional branch is the comparison of the summation function expression with zero, the truth branch is the multiplication and division operation between the metering point electricity and the summation result, and the false branch is zero.
[0094] Step 2.2: Based on the preset metering point mapping table, replace the machine identifier corresponding to the metering point code with a readable variable, and retain the operators and parentheses to generate a mathematical expression.
[0095] In this embodiment, based on a preset metering point mapping table, the code of each metering point in the Token sequence is replaced with the corresponding readable variable, the grid-connected electricity identifier is replaced with the grid-connected electricity readable variable, and all operators and parentheses are retained to generate a mathematical expression composed of readable variables. Combining the total active power value and grid-connected electricity value of each metering point in the structured data packet, the mathematical expression is substituted to calculate the allocated electricity value of each power generation point.
[0096] Step 2.3: Perform bracket matching verification, denominator non-zero verification, measurement point attribution verification, and field type verification on the mathematical expression in sequence.
[0097] In this embodiment, the bracket pairing verification is used to verify the symmetry and nesting legality of brackets based on the bracket type in the Token; the denominator non-zero verification is used to verify whether the denominator in the division operation is legal based on the electricity values of each metering point participating in the allocation in the structured data packet; the metering point affiliation verification is used to verify whether all metering points referenced in the mathematical expression belong to the current generator number based on the metering point code in the Token; and the field type verification is used to verify whether there are incompatible cross-references or substitution operations between the total active power and the positive active power based on the field type in the Token.
[0098] In this embodiment, the bracket pairing verification traverses the token sequence and maintains a stack structure. Left brackets are pushed onto the stack sequentially, and when a right bracket is encountered, a matching left bracket is popped out. Finally, the stack is empty, and the verification passes. The non-zero denominator verification calculates the sum of the electricity to be allocated based on the total active power of each metering point in the structured data packet. The verification result is that the sum is not equal to zero, and the verification passes. The metering point attribution verification confirms that all metering point codes referenced in the mathematical expression belong to the current generator number, and the verification passes. The field type verification confirms that all fields referenced in the mathematical expression are total active power fields and are not mixed with positive active power fields, and the verification passes.
[0099] Step 2.4: Mark any mathematical expression that fails any of the following checks: bracket pairing check, denominator non-zero check, meter point attribution check, and field type check, as an anomaly. Record the corresponding check type and the location information determined based on the Token parsing location. Determine the anomaly type of the anomaly based on the check type and generate a formula check result containing the anomaly type and location information.
[0100] In this embodiment, all four checks passed, no formula-level anomalies were detected, and the judgment values corresponding to each rule in the formula check results were all normal.
[0101] Step 3: Input the structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, read the corresponding field values from the structured data packet to perform anomaly judgment and generate a list of anomalies in the large language model.
[0102] In this embodiment, the network structure of the large language model includes an input encoding layer, a deep inference layer, a result generation layer, and an output layer. The input encoding layer receives an input text sequence and converts it into a vector representation. The input text sequence includes at least the structured data packet, field mapping relationships in the mathematical model mapping knowledge base, and natural language rules in the verification rule knowledge base. The deep inference layer includes a multi-head self-attention module and a feedforward neural network module. The multi-head self-attention module captures dependencies between different positions in the vector representation, and the feedforward neural network module performs a non-linear transformation on the output of the multi-head self-attention module to generate an enhanced semantic representation. The deep inference layer performs inference operations based on the enhanced semantic representation, sequentially parsing the structured data packet using the generator number as an index, rewriting the natural language rules into data query statements, querying the corresponding field values in the structured data packet, and performing rule-based inference diagnosis on the queried field values. The result generation layer uses an autoregressive generation method to decode the enhanced semantic representation word by word and generate a combination of lexical sequences for a large language model anomaly list. The output layer combines the lexical sequences into a text-based large language model anomaly list and outputs it.
[0103] The large language model employs a pre-trained language model based on a Transformer decoder architecture. The specific parameter configuration of the pre-trained language model network structure is as follows: the input encoding layer uses positional encoding vectors with a word embedding dimension of 4096, and the maximum input sequence length is 8192 words; the deep inference layer consists of 32 stacked Transformer decoders, each containing a multi-head self-attention module and a feedforward neural network module. The multi-head self-attention module has 32 heads, each with a dimension of 128, and the hidden layer dimension of the feedforward neural network module is 16384. The activation function is a Gaussian error linear unit; the result generation layer uses an autoregressive generation method, with a beam search width of 4 and a temperature coefficient of 0.7; the output layer converts the hidden states into a probability distribution of the vocabulary size through a linear mapping, with a vocabulary size of 50272.
[0104] In this embodiment, the training method of the large language model includes:
[0105] Based on historical distributed photovoltaic power generation business data and the corresponding historical verification anomaly list, a training sample set is constructed. The training sample set includes multiple training samples. The input data of each training sample includes structured data package samples, field mapping relationship samples in the digital model mapping knowledge base, and natural language rule samples in the verification rule knowledge base. The labels of each training sample are the corresponding anomaly list labels.
[0106] Based on the training sample set, the large language model is pre-trained using the cross-entropy loss function. The cross-entropy loss function is used to measure the difference between the word sequence output by the large language model and the labels of the anomaly list, so as to update the network parameters of the large language model and obtain the pre-trained large language model.
[0107] Based on the pre-trained large language model, a scenario verification sample set for the large language model in the distributed photovoltaic accounting anomaly analysis scenario is obtained, and the pre-trained large language model is fine-tuned. The fine-tuning training adopts a low-rank adaptive method, which freezes the pre-training parameters of the large language model and injects a trainable low-rank decomposition matrix to update the parameters of the low-rank decomposition matrix, thereby obtaining the fine-tuned large language model.
[0108] In the pre-training phase of this embodiment, a large-scale general corpus was used, comprising power industry texts, technical documents, and general Chinese corpora, totaling approximately 1.2TB. The AdamW optimizer was used during pre-training with a learning rate of 3e-4 and a cosine annealing scheduling strategy. The warm-up steps were 2000, the batch size was 256 samples, and the training epochs were 3. The label smoothing factor in the cross-entropy loss function was set to 0.1. In the fine-tuning phase, a low-rank adaptive method was used, setting the rank of the low-rank decomposition matrix to 8 and the scaling factor to 16. Trainable parameters were injected only into the projection matrices of the query vector and value vector. During fine-tuning, the learning rate was set to 2e-5, the batch size was 16, the training epochs were 5, and an early stopping strategy on the validation set was used with a patience value of 2 epochs. The training sample set contained at least 100,000 sets of historical distributed photovoltaic power generation business data samples, covering different power generation account numbers, different business periods, and combinations of various anomaly types.
[0109] Step 3.1: Based on the digital model mapping knowledge base, map the business terms in natural language to specific database tables and fields, and map the state values described in natural language to the corresponding standard code values to obtain the field mapping relationship.
[0110] In this embodiment, the digital-analog mapping knowledge base stores the correspondence between database table names, field English names, field Chinese names, standard code values, and code meanings in a structured table format. The field mapping relationship includes at least the user status field CUST_STATUS and its codes 01 ("normal"), 02 ("in transit"), and 09 ("account closed"); the user type field CUST_TYPE and its code 02 ("distributed photovoltaic"); the metering point type field METER_TYPE and its codes 01 ("generation point") and 03 ("grid connection point"); the increase / existence flag field TYPE_FLAG and its values "increment" and "existence"; and the field mapping relationship between the total active power field AP and the grid-connected power field E_GRID.
[0111] Step 3.2: Based on the field mapping relationship, read the field values corresponding to each natural language rule in the verification rule knowledge base from the structured data packet, and perform anomaly judgment on the field values to obtain a preliminary list of anomalies.
[0112] In this embodiment, the verification rule knowledge base stores the verification rule text in pure Chinese natural language, including at least nine rules: distributed photovoltaic power generation users cannot be in a closed state; distributed photovoltaic users must have at least one power generation point metering point; the on-grid electricity should be equal to the sum of the electricity allocated to each power generation point; the current AP value of each power generation point metering point cannot be empty; the mechanism price field of incremental projects cannot be empty; the product of the CT and PT ratios of the power generation point metering point equals the comprehensive ratio and is consistent with the file registration; normal issuance is allowed when the user status is not in transit; when the AP value of the power generation point suddenly drops by more than 50% compared to the previous period, it is marked as pending review; and the mechanism price of existing projects should be consistent with the policy price of the current year. The large language model reads the corresponding field values in the structured data packet according to each rule, performs the judgment, and outputs the abnormal hit flag and field-level details of each rule.
[0113] Step 3.3: Generate prompt words based on the structured data packet, the field mapping relationship of the digital-to-analog mapping knowledge base, and the natural language rules of the verification rule knowledge base.
[0114] In this embodiment, the complete JSON content of the structured data packet, all mapping entries of the mathematical model mapping knowledge base, and all rule texts of the verification rule knowledge base are concatenated into a prompt word for a single reasoning operation. The prompt word includes a role definition section, a verification rule list section, a database field English-Chinese mapping table section, a structured data packet section to be verified section, and an output requirement section. The output requirement section specifies that the large language model outputs anomaly hit judgment and field-level details for each rule one by one, and summarizes and outputs an anomaly list after all rules have been judged.
[0115] Step 3.4: Drive the large language model according to the prompt words, parse the structured data packet based on the preliminary anomaly list and the generator number as the index, rewrite the natural language rules into data query statements according to the digital model mapping knowledge base, query and confirm the corresponding field values in the structured data packet through the field mapping relationship, and perform rule reasoning diagnosis on the queried field values according to the natural language rules in the verification rule knowledge base to generate the large language model anomaly list.
[0116] In this embodiment, after receiving the prompt word, the large language model performs reasoning diagnosis in the following four steps: parsing and confirming the target account number, rewriting the verification intent into a standard query form, confirming that the structured data packet has been fully loaded, and performing natural language rule reasoning diagnosis one by one. In the reasoning process, natural language business terms are mapped to database table fields and standard code values according to the digital model mapping knowledge base. The corresponding field values are read from the structured data packet for anomaly judgment, and the hit flag of each rule is output. For the hit anomaly, the machine name, current value, expected value, reason, and suggestion of the anomaly field are also output. After all rule judgments are completed, the large language model anomaly list is summarized and output.
[0117] Step 4: Based on the rule engine, independently execute deterministic hard validation rules on the structured data packets to generate rule judgment results. Use the rule judgment results to verify and correct the list of anomalies in the large language model to obtain a verified list of anomalies.
[0118] Step 4.1: According to the preset deterministic hard validation rules, the rule engine independently performs deterministic logic validation on each field in the structured data packet to generate rule judgment results.
[0119] In this embodiment, the rule determination result includes the pass flag, fail flag, and corresponding anomaly flag for each deterministic hard check rule.
[0120] In this embodiment, the rule engine adopts a rule execution framework based on Drools, injecting the value of each field in the structured data packet as a Fact object into the working memory; the deterministic hard validation rules are defined in .drl files, and the rule engine matches the Fact object with the condition part of the deterministic hard validation rule one by one. When the condition is met, the corresponding judgment action is triggered, and the rule judgment result is output; the rule judgment result includes the pass or fail flag of each deterministic hard validation rule and the corresponding exception item flag.
[0121] In this embodiment, at least the following deterministic hard verification rules are included: triggering a cancellation status anomaly determination for distributed photovoltaic users whose user status field CUST_STATUS is "09" and user type field CUST_TYPE is "02"; triggering a no-generator-point-metering-point anomaly determination for users whose user type is distributed photovoltaic and who do not have a metering point of type "01"; triggering an empty reading anomaly determination for metering points whose AP reading is empty or zero; and triggering an in-transit unavailable anomaly determination for users whose user status is "02" and who are in transit.
[0122] Step 4.2: Compare the rule determination result with the list of anomalies in the large language model item by item and perform correction steps to obtain the verified list of anomalies.
[0123] In this embodiment, the correction step includes:
[0124] If a certain deterministic hard validation rule in the rule judgment result is judged to be unsuccessful, and the corresponding item in the large language model anomaly list is not marked as an anomaly, then the corresponding anomaly item is added to the large language model anomaly list.
[0125] If a deterministic hard validation rule in the rule determination result is determined to pass, and the corresponding item in the large language model anomaly list is marked as an anomaly, then the anomaly mark of the corresponding anomaly item is removed from the large language model anomaly list.
[0126] In this embodiment, the 0 or 1 judgment value of each rule in the rule judgment result is recorded as the R vector, and the 0 or 1 judgment value of the corresponding rule in the large language model anomaly list is recorded as the L vector. The two are compared item by item: if the R vector is 1 and the L vector is 0, the anomaly corresponding to the rule is added to the large language model anomaly list and marked as a high confidence anomaly; if the L vector is 0 and the L vector is 1, the anomaly is retained but the rule engine weight contribution is zero in the subsequent fusion, making it tend to be low confidence and need to be reviewed; if the R vector is equal to the L vector, the original judgment result remains unchanged.
[0127] Step 5: The formula verification results, the verified list of outliers, and the rule judgment results are weighted and fused to calculate the fusion confidence of each outlier. The fusion confidence of each outlier is then graded according to the preset threshold and the hard error priority rule. The output includes an outlier analysis list with outliers marked with confidence level.
[0128] Step 5.1: Based on the formula verification results, the list of verified anomalies, and the rule judgment results, the verification rules are divided into deterministic rules of status archive, rules of power value calculation, and rules of cross-field semantic understanding. The large language model weight, rule engine weight, and formula verification weight are configured for each verification rule to obtain the weight configuration corresponding to each verification rule.
[0129] In this embodiment, the verification rules are divided into three types based on their business characteristics: Status profile deterministic rules, which include at least user status verification and generator point metering point existence verification, focusing on the deterministic judgment of the rule engine, with a large language model weight of 0.2, a rule engine weight of 0.5, and a formula verification weight of 0.3; Electricity value calculation rules, which include at least verification of the sum of on-grid electricity and allocated electricity and verification of the non-empty value of the metering point AP, focusing on the accurate calculation capability of the formula verification, with a large language model weight of 0.3, a rule engine weight of 0.3, and a formula verification weight of 0.4; and Cross-field semantic understanding rules, which include at least verification of the non-empty value of incremental project mechanism electricity price and verification of the consistency between the existing project mechanism electricity price and policy electricity price, focusing on the semantic reasoning advantages of the large language model, with a large language model weight of 0.6, a rule engine weight of 0.3, and a formula verification weight of 0.1.
[0130] Step 5.2: Based on the weight configuration corresponding to each verification rule, obtain the formula verification result, the list of verified anomalies, and the judgment value corresponding to each rule judgment result, and calculate the fusion confidence of each verification rule according to the preset weighted summation formula; when the judgment value is 1, it is judged as an anomaly, and when the judgment value is 0, it is judged as normal.
[0131] Step 5.3: Based on a preset threshold and a hard error priority rule, classify and determine the fusion confidence level of each anomaly. The hard error priority rule is defined as follows: [The rule determines the anomaly as follows].
[0132] If the fusion confidence reaches a preset threshold, or the rule judgment result is abnormal, or the formula verification result, the verified list of abnormal items, and the judgment result of the rule judgment are all abnormal, and the fusion confidence plus a preset boost value reaches the preset threshold, then the abnormal item is determined to be a high confidence abnormality.
[0133] Otherwise, the anomaly is determined to be of low confidence and requires further review.
[0134] Step 5.4: Based on the classification judgment results, mark the corresponding confidence level for each anomaly in the anomaly analysis list, and output the anomaly analysis list including anomalies marked with confidence levels.
[0135] In this embodiment, anomalies classified as high-confidence anomalies are marked as "high-confidence anomalies," and anomalies classified as low-confidence anomalies requiring review are marked as "low-confidence anomalies requiring review." An anomaly analysis list is generated only for anomalies marked with confidence levels. Each anomaly in the anomaly analysis list includes at least the corresponding rule ID, the anomaly field corresponding to the anomaly, the current value, the expected value, the reason, and the confidence level.
[0136] Step 6: Determine the handling path based on the field attribution of the anomaly item marked with the confidence level in order to process the distributed photovoltaic accounting anomaly.
[0137] Step 6.1: Based on each anomaly item marked with a confidence level, obtain the field attribution information corresponding to each anomaly item.
[0138] In this embodiment, the field attribution information includes the archive field, the metering point configuration field, and the power generation allocation formula configuration field.
[0139] Step 6.2: Based on the field attribution information, classify and handle the abnormal items:
[0140] If the field corresponding to the exception belongs to the archive field, it is classified as an archive exception and the archive maintenance process is triggered to handle the distributed photovoltaic accounting exception.
[0141] If the field corresponding to the anomaly belongs to the metering point configuration field, it is classified as a metering anomaly, and a metering operation and maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly.
[0142] If the field corresponding to the anomaly belongs to the power generation allocation formula configuration field, it is classified as a formula-related anomaly, and a system maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly.
[0143] In this embodiment, the currently detected exception is rule R001, and its corresponding exception field is CUST_STATUS. According to the field attribution mapping table, CUST_STATUS belongs to the archive field, so this exception is classified as an archive-related exception. The archive maintenance entry is displayed on the verification result interface to redirect to the customer basic information archive maintenance page of the marketing system. At the same time, interactive guidance text is generated through the large language model to ask the user in natural language whether to initiate an archive maintenance work order. In response to the user's confirmation instruction, a work order request body containing the account number, verification task ID, fields to be modified, and suggested values is constructed, and the work order initiation interface of the marketing system is called to trigger the archive maintenance approval process.
[0144] In this embodiment, for metering-related anomalies, the metering operation and maintenance work order initiation portal is displayed on the verification results interface; for formula-related anomalies, the system operation and maintenance work order initiation portal is displayed on the verification results interface.
[0145] Example 3
[0146] Based on the same inventive concept as other embodiments, this embodiment introduces a distributed photovoltaic accounting anomaly analysis and processing system, including:
[0147] The data aggregation module is used to aggregate distributed photovoltaic power generation business data using the power generation account number as an index, forming a structured data package;
[0148] The formula verification module is used to parse and verify the power generation allocation formula in the structured data packet, and generate a formula verification result including anomaly markers.
[0149] The language model determination module is used to input the structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, it reads the corresponding field values from the structured data packet to determine anomalies and generates a list of anomalies in the large language model.
[0150] The rule verification module is used to independently execute deterministic hard verification rules on the structured data packet according to the rule engine, generate rule judgment results, and use the rule judgment results to verify and correct the list of anomalies in the large language model to obtain a verified list of anomalies.
[0151] The fusion and grading module is used to weight and fuse the formula verification results, the verified list of anomalies, and the rule judgment results to calculate the fusion confidence of each anomaly. Based on the preset threshold and the hard error priority rule, the fusion confidence of each anomaly is graded and judged, and the output includes an anomaly analysis list including anomalies marked with confidence level.
[0152] The disposal execution module is used to determine the disposal path based on the field belonging to the anomaly item marked with the confidence level in order to handle the distributed photovoltaic accounting anomaly.
[0153] The specific functions of each module described above are explained in the relevant content of Embodiment 1 or 2, and will not be repeated here.
[0154] Example 4
[0155] Based on the same inventive concept as other embodiments, this embodiment describes a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the methods of Embodiment 1 or 2 described above.
[0156] Example 5
[0157] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions that, when executed by a processor, implement the steps of the methods described in Embodiment 1 or 2 above.
[0158] In summary, this invention aggregates distributed photovoltaic power generation business data into structured data packets by using the power generation account number as an index. It then parses and verifies the power generation allocation formula, employs a large language model and a rule engine for anomaly detection and deterministic hard verification correction, respectively. Finally, it weights and fuses the formula verification results, the list of verified anomalies, and the rule judgment results, and determines the handling path based on field affiliation. This achieves fully automated processing of accounting anomalies from data aggregation, formula verification, multi-source fusion judgment to hierarchical handling. It overcomes the technical shortcomings of existing technologies, such as reliance on manual verification, insufficient accuracy of single detection methods, lack of deterministic verification mechanisms for large language model inference results, and the need for manual cross-system handling after anomaly detection. This improves the efficiency and accuracy of distributed photovoltaic accounting anomaly analysis.
[0159] This invention aggregates customer profile data, power generation project profile data, metering point configuration data, issuance fee data, and historical snapshot data using the power generation account number as an index to form a structured data package. It then performs a field-by-field comparison between the historical snapshot data and the current distributed photovoltaic power generation business data within the structured data package to generate comparison difference markers. Simultaneously, it parses the code string corresponding to the power generation allocation formula based on regular expression segmentation and recursive descent parsing methods. Through four-dimensional verification—bracket pairing verification, denominator non-zero verification, metering point attribution verification, and field type verification—it generates a formula verification result containing anomaly type and location information. This achieves structured integration of distributed photovoltaic accounting data and automated parsing and verification of the power generation allocation formula, overcoming the technical shortcomings of existing technologies where anomaly investigation requires repeated queries across multiple business systems and is time-consuming for single-account investigations, thus improving anomaly location efficiency.
[0160] This invention inputs structured data packets, a mathematical model mapping knowledge base, and a verification rule knowledge base into a large language model. The large language model then performs anomaly detection based on the natural language rules in the verification rule knowledge base and the field mapping relationships in the mathematical model mapping knowledge base, generating a list of anomalies for the large language model. Simultaneously, based on the rule engine, deterministic hard verification rules are independently executed on the structured data packets. The rule detection results are used to supplement and correct the list of anomalies in the large language model by adding new entries and removing anomaly markers, resulting in a verified list of anomalies. This overcomes the technical deficiency of the large language model's inference results lacking a deterministic verification mechanism, reducing false positives and false negatives.
[0161] This invention calculates the fusion confidence level of each anomaly by weighting and fusing the formula verification results, the verified anomaly list, and the rule judgment results. It then classifies the fusion confidence level of each anomaly according to a preset threshold and a hard error priority rule. Based on the classification judgment results, it marks the confidence level of each anomaly in the anomaly analysis list. Finally, based on the field affiliation of the anomaly marked with a confidence level, the anomaly is categorized into archive anomalies, measurement anomalies, or formula anomalies, triggering archive maintenance processes, measurement operation and maintenance work orders, or system operation and maintenance work orders respectively. This achieves fully automated connection from anomaly detection and confidence level classification to the triggering of the handling path, overcoming the technical defect in existing technologies where manual cross-system handling operations are still required after anomaly detection, thus reducing labor costs.
[0162] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0163] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0164] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0165] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0166] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for anomaly analysis and processing in distributed photovoltaic accounting, characterized in that, include: Distributed photovoltaic power generation business data is aggregated using the power generation account number as an index to form a structured data package; The power generation allocation formula in the structured data packet is parsed and verified to generate a formula verification result including anomaly markers; The structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base are input into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, the corresponding field values are read from the structured data packet to perform anomaly judgment and generate a list of anomalies in the large language model. Based on the rule engine, deterministic hard validation rules are independently executed on the structured data packets to generate rule judgment results. The rule judgment results are then used to validate and correct the list of anomalies in the large language model to obtain a validated list of anomalies. The formula verification results, the verified list of anomalies, and the rule judgment results are weighted and fused to calculate the fusion confidence of each anomaly. The fusion confidence of each anomaly is then graded according to a preset threshold and a hard error priority rule. The output includes an anomaly analysis list that includes anomalies marked with confidence level. The handling path is determined based on the field attribution of the anomaly item marked with the confidence level in order to process the anomaly in the distributed photovoltaic accounting.
2. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 1, characterized in that, The distributed photovoltaic (PV) power generation business data includes at least customer profile data, power generation project profile data, metering point configuration data, issuance fee data, and historical snapshot data; among which, the distributed PV power generation business data is aggregated using the power generation account number as an index to form a structured data package, including: Data on distributed photovoltaic power generation is obtained based on the power generation account number and then aggregated to obtain an aggregated data package. Organize the data in JSON format, and simultaneously append English field text and Chinese field tags to each field in the aggregated data packet to obtain a formatted data packet; The historical snapshot data in the formatted data packet is compared field by field with the current distributed photovoltaic power generation business data to generate comparison difference markers. The comparison difference markers are then written into the formatted data packet to obtain a structured data packet.
3. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 2, characterized in that, The generation allocation formula in the structured data packet is parsed and verified to generate a formula verification result including anomaly markers, including: Based on regular expression segmentation, the code string corresponding to the power generation allocation formula in the structured data packet is parsed using the recursive descent parsing method to identify the metering point code, field type, operator and parentheses in the Token; According to the preset metering point mapping table, the machine identifier corresponding to the metering point code is replaced with a readable variable, while retaining the operators and parentheses, and a mathematical expression is generated; The mathematical expression is sequentially subjected to bracket pairing verification, denominator non-zero verification, metering point attribution verification, and field type verification. Specifically, the bracket pairing verification checks the symmetry and nesting validity of brackets based on the bracket type in the Token; the denominator non-zero verification checks the validity of the denominator in the division operation based on the electricity values of each metering point participating in the allocation in the structured data packet; the metering point attribution verification checks whether all metering points referenced in the mathematical expression belong to the current generator account number based on the metering point code in the Token; and the field type verification checks whether there are incompatible cross-references or substitution operations between the total active power and the forward active power based on the field type in the Token. Mathematical expressions that fail any of the following checks—bracket pairing, non-zero denominator, measurement point attribution, and field type—are marked as anomalies. The corresponding check type and the location information determined based on the Token parsing location are recorded. The anomaly type of the anomaly is determined according to the check type, and a formula verification result containing the anomaly type and location information is generated.
4. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 3, characterized in that, Based on the natural language rules in the verification rule knowledge base and the field mapping relationships in the mathematical model mapping knowledge base, the corresponding field values are read from the structured data packet for anomaly detection, generating a list of language model anomalies, including: Based on the digital model mapping knowledge base, business terms in natural language are mapped to specific database tables and fields, and state values described in natural language are mapped to corresponding standard code values to obtain field mapping relationships; Based on the field mapping relationship, the field values corresponding to each natural language rule in the verification rule knowledge base are read from the structured data packet, and the field values are judged to be abnormal to obtain a preliminary list of abnormal items; Based on the structured data packet, the field mapping relationship of the numerical model mapping knowledge base, and the natural language rules of the verification rule knowledge base, prompt words are generated; Driven by the prompt words, the large language model parses the structured data packet based on the preliminary anomaly list and the generator number as the index. It then rewrites the natural language rules into data query statements according to the digital model mapping knowledge base. The corresponding field values in the structured data packet are confirmed through the field mapping relationship, and the queried field values are analyzed using rule-based reasoning based on the natural language rules in the verification rule knowledge base, generating a large language model anomaly list.
5. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 4, characterized in that, Based on the rule engine, deterministic hard validation rules are independently executed on the structured data packets to generate rule judgment results. These results are then used to validate and correct the list of outliers in the large language model, resulting in a validated list of outliers, including: According to the preset deterministic hard validation rules, the rule engine independently performs deterministic logic validation on each field in the structured data packet to generate rule judgment results; the rule judgment results include the pass flag, fail flag and corresponding exception flag for each deterministic hard validation rule; The rule determination result is compared item by item with the list of anomalies in the large language model, and a correction step is performed to obtain a verified list of anomalies. The correction steps include: If a certain deterministic hard validation rule in the rule judgment result is judged to be unsuccessful, and the corresponding item in the large language model anomaly list is not marked as an anomaly, then the corresponding anomaly item is added to the large language model anomaly list. If a deterministic hard validation rule in the rule determination result is determined to pass, and the corresponding item in the large language model anomaly list is marked as an anomaly, then the anomaly mark of the corresponding anomaly item is removed from the large language model anomaly list.
6. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 5, characterized in that, The formula validation results, the validated list of outliers, and the rule-based judgment results are weighted and fused to calculate the fusion confidence of each outlier. Based on a preset threshold and a hard error priority rule, the fusion confidence of each outlier is graded and judged, outputting an anomaly analysis list including outliers marked with confidence levels, including: Based on the formula verification results, the verified list of anomalies, and the rule judgment results, the verification rules are divided into deterministic rules of status archive, rules of power value calculation, and rules of cross-field semantic understanding. The large language model weight, rule engine weight, and formula verification weight are configured for each verification rule to obtain the weight configuration corresponding to each verification rule. Based on the weight configuration corresponding to each verification rule, the formula verification result, the list of verified anomalies, and the judgment value corresponding to each rule judgment result are obtained, and the fusion confidence of each verification rule is calculated according to the preset weighted summation formula; when the judgment value is 1, it is judged as an anomaly, and when the judgment value is 0, it is judged as normal. The fusion confidence level of each anomaly is graded and determined according to a preset threshold and a hard error priority rule, wherein the hard error priority rule determines that the rule result is an anomaly. If the fusion confidence reaches a preset threshold, or the rule judgment result is abnormal, or the formula verification result, the verified list of abnormal items, and the judgment result of the rule judgment are all abnormal, and the fusion confidence plus a preset boost value reaches the preset threshold, then the abnormal item is determined to be a high-confidence abnormal item; otherwise, the abnormal item is determined to be a low-confidence item pending review. Based on the classification judgment results, each anomaly in the anomaly analysis list is marked with a corresponding confidence level, and an anomaly analysis list including anomalies marked with confidence levels is output.
7. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 6, characterized in that, Based on the field attribution of the anomalies marked with confidence level, a handling path is determined to process the distributed photovoltaic accounting anomalies, including: Based on each anomaly marked with a confidence level, obtain the field attribution information corresponding to each anomaly; the field attribution information includes the archive field, the metering point configuration field, and the power generation allocation formula configuration field; Based on the field attribution information, the abnormal items are classified and handled as follows: If the field corresponding to the exception belongs to the archive field, it is classified as an archive exception and the archive maintenance process is triggered to handle the distributed photovoltaic accounting exception. If the field corresponding to the anomaly belongs to the metering point configuration field, it is classified as a metering anomaly and a metering operation and maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly. If the field corresponding to the anomaly belongs to the power generation allocation formula configuration field, it is classified as a formula-related anomaly, and a system maintenance work order is triggered to handle the distributed photovoltaic accounting anomaly.
8. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 7, characterized in that, The network structure of the large language model includes: An input encoding layer is used to receive an input text sequence and convert the input text sequence into a vector representation, wherein the input text sequence includes at least the structured data packet, the field mapping relationship in the digital model mapping knowledge base, and the natural language rules in the verification rule knowledge base; The deep inference layer includes a multi-head self-attention module and a feedforward neural network module. The multi-head self-attention module is used to capture the dependencies between different positions in the vector representation, and the feedforward neural network module is used to perform a non-linear transformation on the output of the multi-head self-attention module to generate an enhanced semantic representation. The deep inference layer is used to perform inference operations according to the enhanced semantic representation, sequentially parsing the structured data packet with the generator number as an index, rewriting the natural language rules into data query statements, querying the corresponding field values in the structured data packet, and performing rule-based inference diagnosis on the queried field values. The result generation layer is used to generate word sequence combinations of a large language model anomaly list by decoding word by word according to the enhanced semantic representation using an autoregressive generation method. The output layer is used to combine the word sequence into a large language model anomaly list in text form and output it.
9. The distributed photovoltaic accounting anomaly analysis and processing method according to claim 8, characterized in that, The training method for the large language model includes: Based on historical distributed photovoltaic power generation business data and the corresponding historical verification anomaly list, a training sample set is constructed. The training sample set includes multiple training samples. The input data of each training sample includes structured data package samples, field mapping relationship samples in the digital model mapping knowledge base, and natural language rule samples in the verification rule knowledge base. The labels of each training sample are the corresponding anomaly list labels. Based on the training sample set, the large language model is pre-trained using the cross-entropy loss function. The cross-entropy loss function is used to measure the difference between the word sequence output by the large language model and the labels of the anomaly list, so as to update the network parameters of the large language model and obtain the pre-trained large language model. Based on the pre-trained large language model, a scenario verification sample set for the large language model in the distributed photovoltaic accounting anomaly analysis scenario is obtained, and the pre-trained large language model is fine-tuned. The fine-tuning training adopts a low-rank adaptive method, which freezes the pre-training parameters of the large language model and injects a trainable low-rank decomposition matrix to update the parameters of the low-rank decomposition matrix, thereby obtaining the fine-tuned large language model.
10. A distributed photovoltaic accounting anomaly analysis and processing system, characterized in that, include: The data aggregation module is used to aggregate distributed photovoltaic power generation business data using the power generation account number as an index, forming a structured data package; The formula verification module is used to parse and verify the power generation allocation formula in the structured data packet, and generate a formula verification result including anomaly markers. The language model determination module is used to input the structured data packet, the preset digital model mapping knowledge base, and the preset verification rule knowledge base into the pre-trained large language model. Based on the natural language rules in the verification rule knowledge base and the field mapping relationship in the digital model mapping knowledge base, it reads the corresponding field values from the structured data packet to determine anomalies and generates a list of anomalies in the large language model. The rule verification module is used to independently execute deterministic hard verification rules on the structured data packet according to the rule engine, generate rule judgment results, and use the rule judgment results to verify and correct the list of anomalies in the large language model to obtain a verified list of anomalies. The fusion and grading module is used to weight and fuse the formula verification results, the verified list of anomalies, and the rule judgment results to calculate the fusion confidence of each anomaly. Based on the preset threshold and the hard error priority rule, the fusion confidence of each anomaly is graded and judged, and the output includes an anomaly analysis list including anomalies marked with confidence level. The disposal execution module is used to determine the disposal path based on the field belonging to the anomaly item marked with the confidence level in order to handle the distributed photovoltaic accounting anomaly.