A method and system for assisting power marketing professional electricity fraud discrimination using a large language model

By using a large language model to perform cross-dimensional cross-validation of multi-source heterogeneous power business data, the problem that traditional power anti-fraud technologies cannot understand unstructured text and perform cross-system verification is solved, enabling accurate identification and characterization of electricity bill fraud.

CN122433031APending Publication Date: 2026-07-21SICHUAN ZHIHE NEW ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN ZHIHE NEW ENERGY TECH CO LTD
Filing Date
2026-03-10
Publication Date
2026-07-21

Smart Images

  • Figure CN122433031A_ABST
    Figure CN122433031A_ABST
Patent Text Reader

Abstract

The application discloses a method and system for power marketing professional electricity fraud discrimination assisted by a large language model, and relates to the technical field of power system automation.The application constructs an extremely rigorous intelligent anti-fraud logic closed loop through deep integration of multi-source heterogeneous data and cross-domain verification mechanism by a large language model.Firstly, by using semantic entailment analysis, the field investigation working condition is high-dimensionally aligned with the complex electricity price policy boundary, and the hidden policy application violation problem is effectively identified.Then, the large language model is endowed with the tool capability of active cross-domain calling, which can convert natural language expression into physical material query instruction, realizing the automatic verification of account and reality coincidence.Furthermore, meteorological data and bottom layer network log environment comparison are innovatively introduced, through strict spatiotemporal causal chain tracing, the logic of continuous criminal acts from text forgery to account tampering is crossed, and the high-risk level of collaborative perjury behavior is accurately characterized and exposed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system automation technology, and specifically relates to a method and system for identifying electricity fraud in power marketing using a large language model. Background Technology

[0002] In daily electricity marketing operations, processes such as user applications for electricity usage status, on-site survey records for business expansion, and approvals for electricity bill refunds or reversals generate massive amounts of unstructured text containing natural language descriptions. Furthermore, the underlying physical equipment of the power grid is widely distributed in the natural environment, its operating status is objectively affected by weather conditions, and low-level alarm logs are left in the distribution automation system. Ensuring the compliance of electricity billing is crucial. Traditional electricity anti-fraud technologies are often limited to purely numerical load characteristic anomaly detection algorithms within a single system. On the one hand, these models cannot understand the deep semantic relationships between complex unstructured business texts and electricity price policy constraints, and are often powerless against users who exploit ambiguities in policy wording to falsely report electricity usage. On the other hand, traditional risk control systems lack cross-system calls and physical entity verification capabilities, and cannot penetrate compliant accounting approval processes to verify the authenticity of the underlying equipment ledgers. Moreover, collaborative fraud can even lead to the creation of perfectly forged text records and ledger data. Existing technologies lack a cross-verification mechanism with independent, objective physical environment data, making it impossible to verify the spatiotemporal logical causal relationship of reasons such as environmental damage in refund documents. They are easily misled by seemingly reasonable series of false evidence. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method and system for identifying electricity fraud in the power marketing field using a large language model, thereby solving the aforementioned technical problems.

[0004] A method for identifying electricity fraud in the power marketing field using a large language model includes the following steps: Acquire multi-source heterogeneous power business data of the target user. The multi-source heterogeneous power business data includes at least structured business file data and unstructured business flow text. The unstructured business flow text includes electricity fee refund or reversal approval text containing natural language reasons for refunds, as well as on-site survey record text. When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is started. In response to the discrimination task, a large language model is used to perform semantic analysis and feature extraction on the unstructured business flow text, and cross-dimensional cross-validation is performed. The cross-validation includes: policy applicability verification by comparing the business file data with the electricity price policy vector database, physical ledger authenticity verification by comparing with the equipment asset system, and spatiotemporal correlation verification by comparing with the objective environment database. Based on the results of the above verifications, an electricity fraud detection report containing qualitative information about abnormal states is output.

[0005] Preferably, the acquisition of multi-source heterogeneous power business data further includes: acquiring the electricity load characteristics and payment behavior data of the target user; When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is initiated, which specifically includes the following steps: A pre-trained tree-based machine learning algorithm is used to process the electricity load characteristics and payment behavior data to obtain a basic risk score. When the basic risk score exceeds a preset threshold, it is determined that the target user's business data meets the preset risk triggering condition, triggering a large language model discrimination task for the unstructured business flow text.

[0006] Preferably, the policy applicability verification by comparing the business file data with the electricity price policy vector database specifically includes the following steps: Based on the current electricity consumption characteristics of the target user in the business file data, retrieve the corresponding electricity price policy constraint clauses from the electricity price policy vector database; Using the large language model, semantic implication analysis is performed to determine whether the actual operating conditions described in the on-site survey record text meet the applicable conditions of the electricity price policy constraints. The result that does not meet the applicable conditions is used as the anomaly determination result of the policy applicability verification.

[0007] Preferably, retrieving the corresponding electricity price policy constraint clauses from the electricity price policy vector database specifically includes the following steps: The historical and current electricity price implementation rules are obtained and segmented into paragraphs. A high-dimensional semantic vector is generated using a text embedding model and stored in the electricity price policy vector database. Extract industry classification tags and billing strategy identifiers from the business file data as query statements, perform vector similarity retrieval, and recall electricity price policy constraint clauses with the highest relevance that contain scope restrictions and exclusion clauses.

[0008] Preferably, the verification of the authenticity of the physical ledger compared with the equipment asset system specifically includes the following steps: The verification of the authenticity of the physical ledger compared with the equipment asset system specifically includes the following steps: The large language model is used to extract named entities from the electricity fee refund or reversal approval text to infer the implicit physical equipment change actions. Using the tool call function of the large language model, a corresponding standard database query statement is generated, and a cross-domain query command is sent to the equipment asset system to verify the work order record and material entry and exit record with the metering point corresponding to the target user as the anchor point. If the cross-domain query command returns an empty result, a ledger missing risk flag is generated as an anomaly determination result for the verification of the authenticity of the physical ledger.

[0009] Preferably, the spatiotemporal correlation verification by comparing with an objective environmental database specifically includes the following steps: Extract the cause of loss entity, timestamp entity, and spatial location entity from the natural language refund reason; When the entity causing the damage falls within the category of environmental meteorological or power grid external force damage characteristics, the large language model is used to automatically generate spatiotemporal causal tracing instructions to call the objective environmental database independent of the marketing system for spatiotemporal data cross-comparison.

[0010] Preferably, the step of performing spatiotemporal data cross-comparison by calling the objective environment database independent of the marketing system specifically includes the following steps: The spatial location entity or the power supply address of the target user is converted into geographical latitude and longitude coordinates, and the information of the transformer equipment associated with the target user is obtained; The entity representing the cause of the damage, the geographical latitude and longitude coordinates, the equipment information of the transformer area, and the timestamp entity are encapsulated into request parameters and sent to the meteorological data server and the power distribution automation server, which serve as the objective environment database. If the results returned by the meteorological data server and the power distribution automation server are received, and the results indicate that no severe weather event occurred within the time period corresponding to the timestamp entity, or that there is no corresponding lightning trip or grounding fault alarm log in the transformer equipment information associated with the target user, then the large language model is used to determine that there is a cross-system data contradiction, and a spatiotemporal correlation anomaly conclusion with the highest risk level is generated.

[0011] Preferably, after acquiring the multi-source heterogeneous power business data of the target user and before performing semantic analysis on the unstructured business flow text using a large language model, a data space mapping processing step is further included: Using the target user's electricity account number as the primary key, a unique mapping relationship is established between the user's account number and the physical electricity meter asset number, spatial geographic coordinates, and the topology node of the connected distribution transformer through a ledger mapping table between the marketing system, equipment asset system, and preset geographic information system.

[0012] Preferably, the method for generating the electricity fraud detection report that includes qualitative analysis of abnormal states includes the following steps: Based on the cross-validation results of the large language model, conflicting policy citations, missing segments that failed the ledger verification, and spatiotemporal contradictory evidence chains that failed the objective environment database verification are structurally assembled to generate a natural language detection report containing the reasoning chain for restoring abnormal data.

[0013] A system for identifying electricity fraud in electricity marketing using a large language model includes: at least one processor and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 9.

[0014] The beneficial effects of this invention are as follows: By deeply integrating multi-source heterogeneous data and cross-domain verification mechanisms through a large language model, an extremely rigorous intelligent anti-fraud logic loop is constructed. Firstly, semantic implication analysis is used to align on-site survey conditions with the boundaries of complex electricity pricing policies in a high dimension, effectively identifying hidden policy application violations. Then, the large language model is endowed with the tool capability of proactive cross-domain invocation, enabling the transformation of natural language expressions into physical material query instructions, achieving automated verification of consistency between accounts and physical assets. Furthermore, it innovatively introduces environmental comparison between meteorological data and underlying network logs, and through rigorous spatiotemporal causal chain tracing, bypasses the chain of criminal logic from text forgery to ledger tampering, accurately identifying and exposing high-risk collaborative perjury behavior. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 The present invention provides a flowchart of the steps for a method of identifying electricity fraud in the power marketing industry using a large language model. Detailed Implementation

[0017] The following disclosure provides many different embodiments or examples for implementing various embodiments of the invention. To simplify the disclosure, specific embodiments are described below. Of course, these are merely examples and are not intended to limit the scope of the invention.

[0018] The embodiments of the invention will now be described in detail with reference to the accompanying drawings.

[0019] like Figure 1 As shown, a method for identifying electricity fraud in the power marketing field using a large language model includes the following steps: Acquire multi-source heterogeneous power business data of the target user. The multi-source heterogeneous power business data includes at least structured business file data and unstructured business flow text. The unstructured business flow text includes electricity fee refund or reversal approval text containing natural language reasons for refunds, as well as on-site survey record text. When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is started. In response to the discrimination task, a large language model is used to perform semantic analysis and feature extraction on the unstructured business flow text, and cross-dimensional cross-validation is performed. The cross-validation includes: policy applicability verification by comparing the business file data with the electricity price policy vector database, physical ledger authenticity verification by comparing with the equipment asset system, and spatiotemporal correlation verification by comparing with the objective environment database. Based on the results of the above verifications, an electricity fraud detection report containing qualitative information about abnormal states is output.

[0020] The business data of the power system is naturally distributed in mutually isolated marketing accounts, equipment assets, and production operation systems. Highly concealed fraud is often accompanied by internal personnel falsifying compliance documents for approval processes. Leveraging the natural language understanding capabilities of large language models and the ability to invoke external tools, isolated unstructured business application reasons can be transformed into structured cross-domain query commands. By utilizing external, immutable, objective spatiotemporal data and underlying physical ledgers, the upper-level accounting logic can be reverse-engineered for verification or falsification, thereby completely blocking seemingly compliant and logically consistent fraudulent accounting operations at the underlying physical logic level.

[0021] Therefore, in practical implementation, the system first accesses the structured business files of the marketing system and the unstructured transfer text containing natural language content such as refund and supplement approvals and on-site inspections. When the system identifies that the target user's business data has triggered preset risk conditions, it automatically calls the large language model to take over the deep discrimination process. The large language model performs information extraction on the unstructured text and initiates three-dimensional cross-validation in parallel: first, it extracts the working condition features from the on-site inspection text, and combines them with the business files to call the retrieval enhancement generation module to search for exclusion clauses in the electricity price policy vector database for policy applicability comparison; second, it parses the refund and supplement approval... The system retrieves fault actions from the approval document and sends query commands to the equipment asset system to compare their authenticity with physical work assignment and return ledgers. Finally, it extracts the natural disaster or external force-damaged entities and their temporal and spatial coordinates from the refund document and verifies their spatiotemporal correlation with the meteorological platform and production operation alarm system. For example, when a refund approval document initiates a large electricity bill reversal based on "meter burnout caused by lightning strike," it not only verifies with the asset system whether there is a genuine fault return record for the meter, but also verifies with the meteorological platform and distribution network whether there were thunderstorms and tripping alarm logs for that geographical coordinate during the time period of the incident. Compared to existing technologies that rely solely on numerical load characteristic anomaly detection, this solution effectively overcomes the barriers of semantic understanding and cross-system information isolation, achieving accurate identification and characterization of exploiting policy loopholes for high-price under-connections and internal collaborative fraud through forged system transfer documents.

[0022] More specifically, the acquisition of multi-source heterogeneous power business data also includes: acquiring the electricity load characteristics and payment behavior data of target users; When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is initiated, which specifically includes the following steps: A pre-trained tree-based machine learning algorithm is used to process the electricity load characteristics and payment behavior data to obtain a basic risk score. When the basic risk score exceeds a preset threshold, it is determined that the target user's business data meets the preset risk triggering condition, triggering a large language model discrimination task for the unstructured business flow text.

[0023] The electricity marketing system deals with real-time structured numerical data from tens of millions of users. Directly using a large language model to perform undifferentiated semantic parsing of the entire business flow data would incur enormous computational costs and system latency. Machine learning models based on decision trees possess extremely high computational efficiency and accuracy in processing tabular and high-dimensional structured numerical features and capturing nonlinear numerical anomalies; while large language models specialize in processing logically complex unstructured natural language text and performing cross-domain, cross-modal reasoning. Therefore, deploying the low-resource-consumption tree model at the front line to perform high-frequency, wide-area numerical anomaly capture, and only triggering the high-resource-consumption large language model for low-frequency, deep text logic and physical ledger review for a very small number of highly suspicious targets that exceed the risk threshold, perfectly aligns with the principle of economical allocation of underlying hardware resources, achieving the optimal technical balance between the breadth of coverage and the accuracy of conviction in anti-fraud investigations.

[0024] In practical implementation, when collecting multi-source heterogeneous power business data, the time-series electricity load characteristics and historical payment behavior data of the target user are extracted simultaneously. The extracted features include the daily average load fluctuation rate, the number of days with zero electricity consumption, and the frequency of late payment penalties. The above structured numerical features are input into a pre-trained machine learning algorithm based on gradient boosting tree or random forest for quantification, and the basic risk score reflecting the probability of anomalies for a single user is output. For example, when a large industrial user's daytime load shows a non-productive cliff-like drop, and there are multiple abnormal refund payment records with manual intervention in the historical business, the basic risk score calculated by the tree model reaches 85 points, which exceeds the set warning threshold of 80 points. This indicates that the preset risk triggering condition is met, and the large language model is automatically activated to perform deep semantic recognition tasks on the unstructured text such as the on-site inspection and refund approval recently generated by the user. Compared to existing technologies that directly input all data into complex networks for indiscriminate computation, this implementation method effectively eliminates the massive amount of redundant business data generated by normal power consumption behavior, significantly reduces the overall computing power consumption of the underlying server, and improves the timeliness of risk location. In the overall solution, it achieves an efficient transition from broadband numerical monitoring to precise logical detection.

[0025] More specifically, the policy applicability verification by comparing the business file data with the electricity price policy vector database includes the following steps: Based on the current electricity consumption characteristics of the target user in the business file data, retrieve the corresponding electricity price policy constraint clauses from the electricity price policy vector database; Using the large language model, semantic implication analysis is performed to determine whether the actual operating conditions described in the on-site survey record text meet the applicable conditions of the electricity price policy constraints. The result that does not meet the applicable conditions is used as the anomaly determination result of the policy applicability verification.

[0026] The essence of high-price, low-connection fraud in the power system lies in the logical contradiction between the objective production processes at the user's site and the abstract electricity pricing policy text issued by the National Development and Reform Commission (NDRC). Since site survey records are highly unstructured natural language, traditional risk control methods cannot directly read and understand their business process attributes. This solution first uses vector retrieval to obtain legally valid objective policy benchmarks as the major premise, then extracts the site survey text as the minor premise, forcing the large language model to perform strict semantic implications and contradictory logical judgments within the given policy context. This design not only utilizes the powerful cross-domain common-sense reasoning capabilities of the large language model but also rigidly constrains the model's output boundaries with external legal rules, completely eliminating the illusionary risk of AI. This allows for the automatic determination, with the rigor of legal auditing, of whether non-standardized equipment descriptions at the site substantially exceed the legal electricity pricing boundaries.

[0027] In practice, the electricity consumption nature tags declared by target users in the business files are extracted and used as a retrieval dimension to accurately recall the corresponding electricity price implementation details text containing applicable scope and exclusion clauses from the electricity price policy vector database. Subsequently, a large language model is called to extract the execution entities and process flow from the unstructured on-site inspection record text, obtain the list of production equipment actually operating on-site and the processing level attributes, and perform semantic implication analysis on them with the recalled policy provisions to determine whether the actual on-site working conditions fall within the physical boundaries allowed by the policy or touch the exclusion clauses. For example, when a user's file records "primary processing of agricultural products" enjoying off-peak electricity prices, but the on-site inspection text records "extrusion injection molding machine and electroplating tank are configured on-site", the large language model identifies "electroplating and injection molding" as belonging to the category of deep processing and industrial manufacturing based on semantic implication analysis, which forms a direct semantic conflict with the policy application conditions of primary processing of agricultural products, and then generates an abnormal judgment result of policy non-compliance. Compared to existing technologies that rely on manual visual inspection or simple regular expression matching, this implementation method effectively overcomes the failure of traditional rule engines caused by non-standardized text descriptions or the deliberate use of ambiguous natural language. It achieves text-level penetration detection of covert "high-price, low-connection" electricity fraud, establishing a core barrier in the overall invention that bridges the gap between pure data verification and policy / legal-level compliance checks. Combining the technical effects of this semantic alignment and compliance review, future technological improvements can extend to deploying edge-side compliance guidance agents on mobile power inspection terminals. This allows for real-time semantic prediction and blocking of irregular orders the instant on-site personnel enter inspection text, achieving an architectural upgrade from passive post-event inspection to proactive in-process interception of violations.

[0028] More specifically, retrieving the corresponding electricity price policy constraints from the electricity price policy vector database includes the following steps: The historical and current electricity price implementation rules are obtained and segmented into paragraphs. A high-dimensional semantic vector is generated using a text embedding model and stored in the electricity price policy vector database. Extract industry classification tags and billing strategy identifiers from the business file data as query statements, perform vector similarity retrieval, and recall electricity price policy constraint clauses with the highest relevance that contain scope restrictions and exclusion clauses.

[0029] Traditional retrieval technologies based on BM25 or SQL suffer from severe lexical mismatch problems. Industry tags in user business profiles often cannot be directly equated with formal terminology in official documents from the National Development and Reform Commission (NDRC), such as pig farms not directly corresponding to the livestock breeding industry. Converting text into high-dimensional semantic vectors essentially encodes the meaning of the text into a mathematical space. Semantically similar words will naturally converge in spatial distance, thus enabling the understanding of the essential content of policies despite literal differences.

[0030] In practice, during the initialization phase, historical and current electricity price implementation rules issued by national and local energy management departments are collected. These rules are then segmented into fine-grained text paragraphs based on semantic integrity. A pre-trained text embedding model is then used to map these policy segments from natural language into high-dimensional dense semantic vectors, which are persistently deployed to a dedicated electricity price policy vector database. The text embedding model employs a vertical domain embedding model fine-tuned using power industry corpus to compress variable-length policy texts into fixed-dimensional dense vectors. When conducting compliance assessments, industry classification tags and billing strategy identifiers, such as those for large industrial electricity consumption and single-rate electricity pricing, are extracted from the target user's profile. These tags are then concatenated to construct a natural language query, which is mapped to a query vector. High-dimensional spatial distance calculations are performed in the vector database to accurately recall binding clauses containing the applicable conditions and punitive exclusions for that type of electricity price. For example, when the profile records a user tag as "agricultural irrigation and drainage electricity," it is converted into a query vector and subjected to near-nearest neighbor similarity retrieval. The vector database not only recalls the basic definition of "applicable standards for agricultural irrigation and drainage electricity prices" but also accurately recalls exclusion clauses containing clauses such as "it is strictly prohibited to connect agricultural and sideline product processing equipment to irrigation and drainage transformers." Before the policy text is entered into the database, a preprocessing stage is required. Specifically, large language models or natural language processing rules can be used to tag the segmented policy paragraphs with metadata tags such as "scope of application," "exclusion clauses," and "penalty standards." During retrieval, a hybrid retrieval mechanism combining vector similarity recall and metadata filtering is employed. This mechanism mandates that the recalled results must include text blocks tagged with "exclusion clauses," ensuring sufficient reverse rejection criteria for subsequent semantic implication analysis of the large model. When performing vector similarity retrieval, the cosine similarity or inner product between the query vector and the database's stored vectors is used as a metric, and the retrieval process is accelerated through near-nearest neighbor algorithms such as hierarchical navigable small-world algorithms. Compared to existing technologies using traditional relational database fuzzy matching or keyword-based inverted indexes, this approach effectively overcomes the problems of missed detections and mismatches caused by frequent changes in electricity pricing policy documents and differences between business terminology and policy legal terminology. Furthermore, by calculating semantic similarity in a multi-dimensional vector space, it achieves accurate cross-contextual positioning of deep policy logic and exclusion clauses, laying the technical foundation for high-quality recall from the external knowledge base of the retrieval enhancement generation mechanism.

[0031] More specifically, the verification of the authenticity of the physical ledger compared with the equipment asset system includes the following steps: The verification of the authenticity of the physical ledger compared with the equipment asset system specifically includes the following steps: The large language model is used to extract named entities from the electricity fee refund or reversal approval text to infer the implicit physical equipment change actions. Using the tool call function of the large language model, a corresponding standard database query statement is generated, and a cross-domain query command is sent to the equipment asset system to verify the work order record and material entry and exit record with the metering point corresponding to the target user as the anchor point. If the cross-domain query command returns an empty result, a ledger missing risk flag is generated as an anomaly determination result for the verification of the authenticity of the physical ledger.

[0032] The core design of cross-domain physical ledger verification lies in leveraging the code generation capabilities and tool invocation mechanisms of large language models. By dynamically injecting database table structure metadata into the model and configuring a standard list of external tools, it essentially endows artificial intelligence with the proxy auditing capability to understand the underlying architecture of heterogeneous systems and proactively initiate empirical retrieval. Internal fraudsters can easily fabricate compliant refund reasons in the marketing system, but it is extremely difficult for them to simultaneously cross network isolation to tamper with the underlying dispatch logs and warehousing logistics ledgers of the equipment asset system. This solution mandates that any abnormal loss declaration at the text level must find corresponding mapping evidence in the physical world's equipment movement trajectory. Simultaneously, a read-only sandbox execution mechanism eliminates the risk of AI overstepping its authority or maliciously deleting or modifying underlying data. Thus, it successfully transforms the highly subjective, experience-dependent, pure text security review into a deterministic, objective database existence verification. In practical implementation, after receiving unstructured electricity bill refund or reversal approval text, the large language model is invoked to perform named entity extraction, parsing out key elements such as fault type and handling method, and then logically inferring the underlying physical equipment changes that must accompany the accounting operation, such as meter installation / removal or transformer replacement. Subsequently, the tool invocation function of the large language model is triggered, using a pre-fine-tuned tool invocation interface, defining the query interface of the equipment asset system as a list of external tools that conforms to specific specifications. In the query instruction generation stage, data definition language fragments or data dictionary metadata of the equipment asset system are dynamically injected into the system prompts of the large language model, guiding the model to accurately translate the natural language inference intent into a standard cross-domain database query that strictly matches the field names of the underlying system. To prevent model illusion risks, the generated query statements are sent to a sandbox execution engine with read-only permissions for execution. This engine issues a retrieval command to the independently running equipment asset system, using the target user's metering point as the primary key, to verify the work order trajectory and material inbound / outbound flow within the corresponding time window. For example, when a marketing approval text record states "Due to a lightning strike causing the electricity meter to burn out and resulting in metering anomalies, the monthly electricity fee will be fully refunded," the large language model infers that there must be a physical action of "returning the old meter and issuing the new meter." It then combines this with the injected table structure metadata to generate a structured query language, which verifies the user's meter replacement work order in the equipment asset system within the sandbox environment. If the returned result is empty, a ledger missing risk indicator is immediately generated, and the refund process is suspected of internal forgery. Compared to existing technologies that only focus on the closed-loop logic of a single system's accounting records, this implementation method breaks down the data isolation barrier between marketing accounts and production materials, and achieves secure automatic translation of unstructured fraudulent intent in natural language into a structured chain of physical evidence. It can not only penetrate seemingly complete fake approval forms, but also expose highly concealed internal collaborative fraud behavior of "faking accounts but not producing physical goods," establishing a hardcore defense mechanism of physical evidence in the overall solution.

[0033] More specifically, the spatiotemporal correlation verification by comparing with an objective environmental database includes the following steps: Extract the cause of loss entity, timestamp entity, and spatial location entity from the natural language refund reason; When the entity causing the damage falls within the category of environmental meteorological or power grid external force damage characteristics, the large language model is used to automatically generate spatiotemporal causal tracing instructions to call the objective environmental database independent of the marketing system for spatiotemporal data cross-comparison.

[0034] The design principle of the aforementioned spatiotemporal correlation verification lies in introducing external independent corroborating evidence to break the "logical self-circulation and self-proof of innocence" loopholes in the internal business system. In complex power fraud scenarios, fraudsters often use force majeure events such as "lightning strikes" and "flooding" as reasonable excuses to cover up abnormal refunds or unauthorized meter alterations. Such environmental damage events naturally lack direct means of falsification within pure financial or marketing databases. The core design of this solution is to use a large language model as a cross-modal translation hub, reducing unstructured case statements to three-dimensional structured physical coordinates, and using these as parameters to call a macroscopic objective database with physical authenticity and immutability. Its underlying logic is that any artificially created false disaster excuse cannot forge corresponding microclimate or external force physical imprints in an independently operating spatiotemporal environmental objective coordinate system, thus achieving a dimensionality-reduction anti-fraud attack by directly penetrating subjectively forged texts with the objective truth of nature.

[0035] In practical implementation, when processing refunds or reversals, the large language model is invoked to perform named entity recognition on the natural language refund reason, accurately parsing the entity causing the power abnormality, the timestamp of the event, and the spatial location of the equipment. Subsequently, the extracted cause of damage is compared with a pre-set disaster feature vocabulary. If the cause is confirmed to fall under the category of environmental meteorological features such as typhoons and thunderstorms or external forces damaging the power grid such as vehicle collisions, the large language model automatically parameterizes the timestamp and spatial location, generating standardized spatiotemporal causal tracing instructions, and proactively invokes a completely independent system separate from the power grid marketing system. Meteorological data servers and other objective environmental databases initiate data cross-comparison; for example, when a refund approval work order describes "On August 15, 2023, localized torrential rain caused flooding, resulting in the total loss of equipment in the power distribution room of an industrial park, and an application is made to reduce the basic electricity fee for that month", the big language model extracts three core entities: "rainstorm and flooding", "August 15, 2023" and "power supply address of an industrial park". Since "rainstorm and flooding" belongs to the category of environmental meteorology, a tracing instruction is immediately generated to call the historical precipitation database of the China Meteorological Administration to verify the actual rainfall level and disaster warning records of the latitude and longitude coordinates of the park during the above time period. Compared to existing technologies that only review the compliance of single-system text approval formats, this implementation method introduces fair third-party data, which is not controlled by internal business personnel, as an anti-fraud audit anchor point. On the other hand, it establishes a hard verification mechanism of causal logic that transcends physical and digital spaces. It can not only directly expose scams that use fabricated force majeure as a pretext to obtain electricity fees, but also effectively sever the collaborative chain of false evidence that uses external natural disasters to cover up internal human-caused illegal electricity use. In the overall solution, it constructs a secure foundation to prevent high-dimensional logic fraud.

[0036] More specifically, the step of cross-referencing spatiotemporal data by calling an objective environment database independent of the marketing system includes the following steps: The spatial location entity or the power supply address of the target user is converted into geographical latitude and longitude coordinates, and the information of the transformer equipment associated with the target user is obtained; The entity representing the cause of the damage, the geographical latitude and longitude coordinates, the equipment information of the transformer area, and the timestamp entity are encapsulated into request parameters and sent to the meteorological data server and the power distribution automation server, which serve as the objective environment database. If the results returned by the meteorological data server and the power distribution automation server are received, and the results indicate that no severe weather event occurred within the time period corresponding to the timestamp entity, or that there is no corresponding lightning trip or grounding fault alarm log in the transformer equipment information associated with the target user, then the large language model is used to determine that there is a cross-system data contradiction, and a spatiotemporal correlation anomaly conclusion with the highest risk level is generated.

[0037] In power fraud scenarios, colluding individuals or unauthorized users can easily manipulate unstructured business transaction text within the marketing system, but they have no ability to tamper with the historical weather database of the China Meteorological Administration, nor can they forge the protection action logs of the underlying hardware measurement and control terminals of the distribution automation system. In the specific implementation process, when performing spatiotemporal data cross-comparison, the first step is to call the geographic information system interface to parse the spatial entities extracted from the refund text or the power supply address recorded in the target user's profile into precise geographic latitude and longitude coordinates. Simultaneously, it penetrates the power grid topology database to obtain the identification of the distribution transformer area and feeder equipment directly connected to the user. Subsequently, the extracted entities indicating damage causes such as lightning strikes or rainstorms, the parsed latitude and longitude coordinates, the transformer area equipment identification, and the time stamp entities accurate to the hour are uniformly encapsulated into structured request parameters, and a two-way concurrent call is initiated: one side requests the historical micro-meteorological conditions of the corresponding spatiotemporal grid from the meteorological data server, which serves as the external objective environmental database, and the other side requests the distribution automation system, which serves as the internal production and operation database. The system retrieves historical operating logs of the corresponding transformer area equipment. For example, when a marketing work order initiates a refund approval based on the reason that "the transformer was burned by lightning on July 20, resulting in over-metering of electricity," the system converts the factory address into coordinates and locates it to "10kV line 05 transformer." It then verifies with the meteorological data server that there are no records of thunderstorms at this coordinate during this period and checks with the distribution automation system to confirm that neither the transformer nor the feeder has generated zero-sequence overcurrent or tripping ground alarms. Based on this, the big data model extracts double negative evidence of "no meteorological conditions" and "no physical alarms," ​​determines that there is a fundamental cross-system data contradiction between the text description and the physical spatiotemporal truth, and directly generates a spatiotemporal correlation anomaly report of electricity fraud with the highest risk level. Compared to existing technologies that rely solely on the surface-level logical consistency of marketing documents for review, this implementation method extends the anti-fraud defense from the financial approval process to the physical power grid operation and the causal relationship between natural weather conditions. It effectively ends the fraudulent chain of fabricating natural disasters to obtain electricity fees and significantly reduces the verification costs and integrity risks of manual on-site investigations through deterministic collision of multi-source heterogeneous data, playing a core anti-counterfeiting role in the overall invention.

[0038] More specifically, after acquiring the multi-source heterogeneous power business data of the target user and before performing semantic analysis on the unstructured business flow text using a large language model, a data space mapping processing step is also included: Using the target user's electricity account number as the primary key, a unique mapping relationship is established between the user's account number and the physical electricity meter asset number, spatial geographic coordinates, and the topology node of the connected distribution transformer through a ledger mapping table between the marketing system, equipment asset system, and preset geographic information system.

[0039] In the IT architecture of large power companies, marketing systems, asset systems, and GIS power distribution systems often suffer from severe data fragmentation and "one item, multiple names" phenomena due to different construction periods. If large language models are used directly to process the free and unstructured business flow text without mapping, the models are prone to logical illusions or failure to align and verify the main body due to unclear cross-system entity references. In practical implementation, after extracting massive amounts of multi-source heterogeneous power business data and before initiating deep semantic discrimination using a large language model, a data spatial mapping preprocessing program is forcibly inserted. Using the target user's unique electricity account number as the core primary key, the program cross-domain calls the customer files of the marketing system, the physical ledgers of the equipment asset system, and the topology elements of the preset geographic information system. Through the underlying ledger mapping table, a four-dimensional unique mapping relationship of "account number - physical electricity meter asset number - spatial geographic coordinates - connected distribution transformer topology node" is accurately constructed. For example, when a high-risk user "account number 10086" is identified, the spatial mapping program immediately binds it to the smart meter "asset number 2023X001", the precise coordinates of 113.5 degrees east longitude and 28.2 degrees north latitude, and the topology node of "10kV line 05 distribution transformer" to form an inseparable digital twin identity card. Compared to existing technologies that rely on fragmented data across business systems and superficial verification based solely on data fields within a single system, this implementation method completely breaks down data silos across systems during the model computation phase. It directly anchors the previously volatile, plain text-based business flow entities to the real physical space and power grid topology, effectively avoiding verification deviations caused by incorrect account relationships or ledger shifts.

[0040] More specifically, the method for generating the electricity fraud detection report, which includes qualitative analysis of abnormal states, includes the following steps: Based on the cross-validation results of the large language model, conflicting policy citations, missing segments that failed the ledger verification, and spatiotemporal contradictory evidence chains that failed the objective environment database verification are structurally assembled to generate a natural language detection report containing the reasoning chain for restoring abnormal data.

[0041] In traditional machine learning-based fraud prevention and control systems, models typically output a "fraud probability score" between 0 and 1. This black-box prediction lacks feasibility in actual power system operations because business personnel cannot use a "95% suspected fraud" score to punish users or hold internal employees accountable; they must rely on manual review of all documents to find conclusive evidence. This solution completely abandons the probability scoring mechanism, instead utilizing the natural language generation and logical arrangement capabilities of a large language model to mimic the case-handling logic of experienced human auditors, forcibly connecting all previously collected objective evidence fragments. It doesn't merely tell users the conclusion, but demonstrates the entire deduction process: using policy documents to point out rule boundaries, using missing physical ledgers to reveal behavioral loopholes, and using meteorological and automated spatiotemporal truth values ​​to provide physical false evidence. This design stitches isolated anomalous data points into a structured narrative with exclusivity and causal coherence, fundamentally solving the pain point of AI risk control models struggling to provide evidence in rigorous judicial investigation scenarios.

[0042] In practical implementation, after completing multi-dimensional cross-validation, the text generation capabilities of the large language model are invoked to structurally align and assemble the original text of conflicting electricity price policy clauses generated during the semantic prediction stage, the missing fragments of equipment asset ledgers that were not found during the cross-domain SQL retrieval stage, and the spatiotemporal inconsistencies in meteorological data and the objective truth of power grid operation returned during the external objective environment interface call stage. This automatically generates a natural language detection report containing a complete abnormal data restoration and reasoning chain. For example, when an electricity account applies for a large refund on the grounds that "the transformer was burned by lightning on August 15th" and the system determines that the claim is questionable, the large language model will extract and splice the following evidence chain to generate a report: "Risk characterization: Extremely high risk of fraudulent refund." The reasoning chain is as follows: 1. Policy conflict: The 'electroplating process' recorded in the survey text violates the scope of application of agricultural drainage electricity pricing; 2. Missing ledgers: The asset system did not retrieve any old form return or new form requisition work orders associated with this account number; 3. Spatiotemporal contradiction: The meteorological database confirms that there were no thunderstorm records at this coordinate on August 15, and the SCADA system did not have a corresponding trip alarm. Compared to existing technologies where traditional risk control systems only output black-box warning scores or rigid error codes, this implementation method translates the complex, high-dimensional data collision results from multiple modalities and cross-systems into a highly interpretable natural language case file with legal audit logic, breaking down the cognitive barriers of human-computer interaction and reducing the time required for secondary penetration verification and manual evidence collection by auditors.

[0043] A system for identifying electricity fraud in electricity marketing using a large language model includes: at least one processor and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 9.

[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for identifying electricity fraud in the power marketing field using a large language model, characterized in that, Includes the following steps: Acquire multi-source heterogeneous power business data of the target user. The multi-source heterogeneous power business data includes at least structured business file data and unstructured business flow text. The unstructured business flow text includes electricity fee refund or reversal approval text containing natural language reasons for refunds, as well as on-site survey record text. When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is started. In response to the discrimination task, a large language model is used to perform semantic analysis and feature extraction on the unstructured business flow text, and cross-dimensional cross-validation is performed. The cross-validation includes: policy applicability verification by comparing the business file data with the electricity price policy vector database, physical ledger authenticity verification by comparing with the equipment asset system, and spatiotemporal correlation verification by comparing with the objective environment database. Based on the results of the above verifications, an electricity fraud detection report containing qualitative information about abnormal states is output.

2. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, The acquisition of multi-source heterogeneous power business data also includes: acquiring the electricity load characteristics and payment behavior data of target users; When it is determined that the target user's business data meets the preset risk triggering conditions, the large language model discrimination task is initiated, which specifically includes the following steps: A pre-trained tree-based machine learning algorithm is used to process the electricity load characteristics and payment behavior data to obtain a basic risk score. When the basic risk score exceeds a preset threshold, it is determined that the target user's business data meets the preset risk triggering condition, triggering a large language model discrimination task for the unstructured business flow text.

3. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, The policy applicability verification, which combines the business file data with the electricity price policy vector database, specifically includes the following steps: Based on the current electricity consumption characteristics of the target user in the business file data, retrieve the corresponding electricity price policy constraint clauses from the electricity price policy vector database; Using the large language model, semantic implication analysis is performed to determine whether the actual operating conditions described in the on-site survey record text meet the applicable conditions of the electricity price policy constraints. The result that does not meet the applicable conditions is used as the anomaly determination result of the policy applicability verification.

4. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 3, characterized in that, The step of retrieving the corresponding electricity price policy constraint clauses from the electricity price policy vector database specifically includes the following steps: The historical and current electricity price implementation rules are obtained and segmented into paragraphs. A high-dimensional semantic vector is generated using a text embedding model and stored in the electricity price policy vector database. Extract industry classification tags and billing strategy identifiers from the business file data as query statements, perform vector similarity retrieval, and recall electricity price policy constraint clauses with the highest relevance that contain scope restrictions and exclusion clauses.

5. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, The verification of the authenticity of the physical ledger compared with the equipment asset system specifically includes the following steps: The verification of the authenticity of the physical ledger compared with the equipment asset system specifically includes the following steps: The large language model is used to extract named entities from the electricity fee refund or reversal approval text to infer the implicit physical equipment change actions. Using the tool call function of the large language model, a corresponding standard database query statement is generated, and a cross-domain query command is sent to the equipment asset system to verify the work order record and material entry and exit record with the metering point corresponding to the target user as the anchor point. If the cross-domain query command returns an empty result, a ledger missing risk flag is generated as an anomaly determination result for the verification of the authenticity of the physical ledger.

6. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, The spatiotemporal correlation verification by comparing with an objective environmental database specifically includes the following steps: Extract the cause of loss entity, timestamp entity, and spatial location entity from the natural language refund reason; When the entity causing the damage falls within the category of environmental meteorological or power grid external force damage characteristics, the large language model is used to automatically generate spatiotemporal causal tracing instructions to call the objective environmental database independent of the marketing system for spatiotemporal data cross-comparison.

7. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 6, characterized in that, The step of cross-referencing spatiotemporal data by calling an objective environment database independent of the marketing system specifically includes the following steps: The spatial location entity or the power supply address of the target user is converted into geographical latitude and longitude coordinates, and the information of the transformer equipment associated with the target user is obtained; The entity representing the cause of the damage, the geographical latitude and longitude coordinates, the equipment information of the transformer area, and the timestamp entity are encapsulated into request parameters and sent to the meteorological data server and the power distribution automation server, which serve as the objective environment database. If the results returned by the meteorological data server and the power distribution automation server are received, and the results indicate that no severe weather event occurred within the time period corresponding to the timestamp entity, or that there is no corresponding lightning trip or grounding fault alarm log in the transformer equipment information associated with the target user, then the large language model is used to determine that there is a cross-system data contradiction, and a spatiotemporal correlation anomaly conclusion with the highest risk level is generated.

8. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, After acquiring the multi-source heterogeneous power business data of the target user and before performing semantic analysis on the unstructured business flow text using a large language model, a data space mapping processing step is also included: Using the target user's electricity account number as the primary key, a unique mapping relationship is established between the user's account number and the physical electricity meter asset number, spatial geographic coordinates, and the topology node of the connected distribution transformer through a ledger mapping table between the marketing system, equipment asset system, and preset geographic information system.

9. The method for identifying electricity fraud in electricity marketing using a large language model as described in claim 1, characterized in that, The method for generating the electricity fraud detection report, which includes qualitative analysis of abnormal conditions, includes the following steps: Based on the cross-validation results of the large language model, conflicting policy citations, missing segments that failed the ledger verification, and spatiotemporal contradictory evidence chains that failed the objective environment database verification are structurally assembled to generate a natural language detection report containing the reasoning chain for restoring abnormal data.

10. A system for identifying electricity fraud in electricity marketing using a large language model, characterized in that, include: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 9.