Electricity price data analysis method and system
By automatically analyzing and classifying electricity price data, using large language models and graph attention networks, the problem of high development and maintenance costs of electricity price data analysis systems is solved, and the API format changes of different electricity price providers are flexible to adapt to the changes in the analysis formats and improve analysis efficiency and accuracy.
Patent Information
- Application Number
- CN202510483728.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
The existing electricity price data analysis method requires the development of analysis logic for each electricity price provider separately, with high development and maintenance costs and unable to automatically adapt to changes in the electricity price provider API format, resulting in poor system flexibility and adaptability.
By analyzing and classifying electricity price data, obtaining the feature representation vector of the target field, and inputting a preset electricity price data analysis model, using a large language model and a graph attention network for automatic analysis, combining dynamic decision boundaries and loss function optimization model to adapt to different formats.
It realizes automatic parsing of electricity price data, reduces the need for manual writing and maintenance of parsers, reduces development and maintenance costs, improves system flexibility and adaptability, and enhances analysis accuracy and efficiency.
Smart Images

Figure CN120337905A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to data analysis and processing, and particularly relates to a method and system for parsing electricity price data. Background Art
[0002] Existing solutions for parsing electricity price data usually include manually writing parsers or using a rule - engine - based approach to parse the electricity price data of each electricity price provider one by one. This method is relatively effective when dealing with a small number of APIs with fixed formats, but when facing multiple electricity price providers and various complex - format APIs, it is necessary to develop separate parsing logics for each electricity price provider, resulting in extremely high development and maintenance costs. In addition, as electricity price providers continuously update their API formats, existing parsers need to be frequently adjusted and cannot automatically adapt to changes, leading to poor flexibility and adaptability of the system. Summary of the Invention
[0003] Embodiments of this application provide a method and system for parsing electricity price data, which can solve the technical problems in existing electricity price data parsing methods, such as the need to develop separate parsing logics for each electricity price provider, resulting in extremely high development and maintenance costs, and that as electricity price providers continuously update their API formats, existing parsers need to be frequently adjusted and cannot automatically adapt to changes, leading to poor flexibility and adaptability of the electricity price data parsing system.
[0004] In a first aspect, embodiments of this application provide a method for parsing electricity price data, including:
[0005] Parsing and classifying the obtained electricity price data to obtain at least one target field;
[0006] Obtaining a feature representation vector for each target field;
[0007] Inputting the feature representation vector of each target field into a preset electricity price data parsing model to obtain the parsing result of each target field.
[0008] In a possible implementation manner of the first aspect, the parsing and classifying the obtained electricity price data to obtain at least one target field includes:
[0009] Obtaining the electricity price data through an API interface;
[0010] Using a large - language model to parse the electricity price data to obtain initial target fields;
[0011] Classifying the initial target fields to obtain at least one of the target fields.
[0012] In a possible implementation manner of the first aspect, the obtaining a feature representation vector for each target field includes:
[0013] Construct a semantic graph based on the target field;
[0014] Obtain the initial feature representation vector of each word in the semantic graph;
[0015] Based on the graph attention network, perform an aggregation operation on the initial feature representation vectors to obtain the feature representation vectors of each target field.
[0016] In a possible implementation manner of the first aspect, after obtaining the parsing result of each target field, further:
[0017] Perform normalization processing on the parsing result according to the preset electricity price data output format to obtain a normalized parsing result.
[0018] In a possible implementation manner of the first aspect, the training process of the electricity price data parsing model includes:
[0019] Determine positive sample pairs and negative sample pairs based on the feature representation vectors of each target field;
[0020] Perform contrastive learning on the positive sample pairs and the negative sample pairs to obtain updated feature representation vectors;
[0021] Determine the center of the category based on the updated feature representation vectors;
[0022] Calculate the distance between each sample pair and the center of the category;
[0023] Based on the distance and the dynamic decision boundary radius of the category, determine the probability that the sample belongs to each category.
[0024] In a possible implementation manner of the first aspect, the loss function L of the electricity price data parsing model intra is:
[0025] where N is the number of samples in one round of training iteration, x i is the feature representation vector of the i-th sample, y i is the category to which the sample i belongs, is the feature representation vector of the center of the category y to which the sample i belongs i of, is the dynamic decision boundary radius of the category y to which the sample i belongs i of.
[0026] In a second aspect, an embodiment of the present application provides an electricity price data parsing system, including:
[0027] A parsing and classification module, configured to parse and classify the obtained electricity price data to obtain at least one target field;
[0028] A feature acquisition module, configured to acquire a feature representation vector of each target field;
[0029] An analysis result determination module, configured to input the feature representation vector of each target field into a preset electricity price data analysis model to obtain an analysis result of each target field.
[0030] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electricity price data analysis method described in any one of the above first aspects is implemented.
[0031] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the electricity price data analysis method described in any one of the above first aspects is implemented.
[0032] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a computer device, the computer device is enabled to execute the electricity price data analysis method described in any one of the above first aspects.
[0033] In the embodiment of the present application, by automatically analyzing and classifying electricity price data, key information in the electricity price data, that is, at least one target field, is identified and extracted. For each target field, its feature representation vector is acquired. The feature representation vector of the target field is input into a preset electricity price data analysis model, which can process these vectors and output an analysis result. It does not depend on a specific data format, so it can adapt to various electricity price data formats. Through the automatic analysis process, the need to manually write and maintain parsers for different electricity price providers or different data formats is reduced, thereby reducing the development and maintenance costs.
[0034] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 is a schematic flowchart of an electricity price data analysis method provided by an embodiment of the present application;
[0037] Figure 2 It is a schematic flowchart of a method for parsing electricity price data provided by an embodiment of the present application;
[0038] Figure 3 It is a schematic structural diagram of an electricity price data parsing system provided by an embodiment of the present application;
[0039] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0040] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0041] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0042] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0043] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0044] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0045] References to "one embodiment" or "some embodiments" in the description of this application mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear in different places in this specification, do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants mean "including but not limited to", unless otherwise specifically emphasized.
[0046] In a photovoltaic and energy storage system, the electricity price not only affects the user's electricity consumption cost, but also affects the economic benefits and operation strategies of the photovoltaic and energy storage system. To provide users with more convenient and flexible electricity price options, developers need to access the electricity price data of major electricity price providers. These electricity price data are usually provided through an Application Programming Interface (API). However, the format of these electricity price data is customized by the electricity price providers, and the formats of the electricity price data provided by different electricity price providers are not the same.
[0047] Figure 1 The schematic flowchart of the electricity price data parsing method provided by an embodiment of this application is shown.
[0048] S101, parse and classify the obtained electricity price data to obtain at least one target field.
[0049] Among them, the electricity price data is structured / semi-structured data generated during the operation of the power market. However, due to different data recording and reporting habits of different electricity price providers, there are differences in the format and content of the electricity price data. Or, the electricity price data is provided through APIs, and the formats of these APIs may vary depending on the electricity price provider, including different structured data such as JSON and XML. It is necessary to parse and uniformly standardize the electricity price data to adapt to different application scenarios.
[0050] Among them, the target field refers to the key information field extracted from the original electricity price data, and these fields contain important data required for further analysis, processing, or decision-making.
[0051] Among them, the electricity price data usually includes the following electricity price data fields:
[0052] (1) Electricity price
[0053] Field name: The field representing the electricity price, and common ones are price, cost, tariff, etc.
[0054] Field value: The specific value of the electricity price, usually in currency per kilowatt-hour (e.g., 0.12 means $0.12 per kilowatt-hour).
[0055] (2) Time
[0056] Field name: Such as time, date, timestamp, etc.
[0057] Time format: The time formats of different providers may be different. ISO format: 2025-02-19T12:00:00Z (international standard format, with time zone); European format: 19 / 02 / 2025 (day / month / year); US format: 02 / 19 / 2025 (month / day / year).
[0058] (3) Currency unit
[0059] Field name: Such as currency, unit, currency_code, etc.
[0060] Field value: Common currency units include US dollar (USD), euro (EUR), Chinese yuan (CNY), etc.
[0061] (4) Other fields
[0062] The electricity price data may also include other auxiliary information, such as:
[0063] Time period: Such as peak period (peak) and off-peak period (off-peak).
[0064] Additional information: Such as tax rate, discount, etc.
[0065] Among them, the target fields may include: electricity price, time, and currency unit. The target fields are a subset of the electricity price data fields.
[0066] S102, obtain the feature representation vector of each target field.
[0067] In the embodiment of the present application, each target field is converted into a feature representation vector, which can capture the key features of the field and is used in the subsequent electricity price data parsing model.
[0068] Among them, the feature representation vector can be obtained in various ways. For example, natural language processing techniques, word embeddings, or other machine learning feature engineering methods are used to extract and learn the features of the target fields.
[0069] S103, input the feature representation vector of each target field into a preset electricity price data parsing model to obtain the parsing result of each target field.
[0070] Among them, in the embodiments of the present application, the above-mentioned feature representation vector is provided as input to a preset electricity price data parsing model. The parsing model processes the input feature representation vector according to the patterns and relationships learned during its training process, and outputs the parsing results of each target field.
[0071] Among them, the parsing result of the target field can not only identify which category the field belongs to (for example, time, electricity price, currency unit, etc.), and optionally, can also output the specific value of the field. This dual output makes the parsing result have both classification information and detailed data content, thus providing richer information for subsequent data processing and analysis. The following are specific output examples:
[0072] Electricity price (Price):
[0073] Type recognition (i.e., the field name above): Electricity price
[0074] Field value: 0.12 USD / kWh
[0075] Application: Can be directly used for calculating costs or comparing with other electricity prices.
[0076] In the embodiments of the present application, by automatically parsing and classifying electricity price data, key information in the electricity price data is identified and extracted, that is, at least one target field. For each target field, its feature representation vector is obtained. The feature representation vector of the target field is input into a preset electricity price data parsing model, and this model can process these vectors and output parsing results. It does not depend on a specific data format, so it can adapt to multiple electricity price data formats. Through the automated parsing process, the need to manually write and maintain parsers for different electricity price providers or different data formats is reduced, thereby reducing development and maintenance costs.
[0077] In an alternative embodiment, S101 parses and classifies the obtained electricity price data to obtain at least one target field, including:
[0078] Step a1, obtaining the electricity price data through an API interface.
[0079] In the embodiments of the present application, electricity price data is automatically obtained through predefined API interfaces. These APIs may be provided by different electricity price providers. The diverse API interfaces from different electricity price providers may each have a unique data format and structure. For example, some APIs may return data in JSON format, while others may return data in XML format. These differences require the electricity price data parsing system to be able to flexibly identify and process data with different structures.
[0080] Step a2, using a large language model to parse the electricity price data to obtain initial target fields.
[0081] In this application, a large language model is applied to parse electricity price data. These models can identify and classify keyword fields in electricity price data, such as electricity price, time, currency unit, etc., according to the learned patterns.
[0082] Step a3: Classify the initial target fields to obtain at least one of the target fields.
[0083] In this step, the extracted initial target fields are classified to ensure that each field is correctly identified and categorized.
[0084] In the embodiment of this application, electricity price data is parsed and initial target fields are extracted. Then, the initial target fields are classified to ensure that each field is correctly identified and categorized. This method not only improves the efficiency and accuracy of data processing, but also enhances the flexibility and adaptability of the electricity price data parsing system, enabling it to handle electricity price data in different formats and structures.
[0085] In an alternative embodiment, a large language model is used to perform a preliminary parse of the electricity price data. Through its natural language processing capabilities, the large language model identifies possible time, electricity price, and currency unit fields and classifies them. During this process, the large language model generates multiple sub-Prompts to simulate the collaboration of multiple agents. This process of multi-agent collaboration simulates the scenario of multiple "agents" (which can be understood as different "experts") working together. Each "agent" has its own responsibilities, but they cooperate with each other to finally obtain an accurate parsing result. Specifically:
[0086] Language understanding agent: Analyze the field content to determine which fields may be prices, time, or currency units.
[0087] Knowledge graph agent: Use knowledge in the power field (such as peak-valley electricity price rules, the correspondence between currency symbols and countries) to verify and refine the preliminary parsing results.
[0088] Decision-making and planning agent: Integrate the outputs of the previous two agents, determine the final field mapping, and trigger more parsing strategies in complex situations.
[0089] Although these "agents" are functionally independent, under the unified management of the Prompt generator, they can be executed sequentially or in parallel. For example, first let the language understanding agent analyze the fields, then the knowledge graph agent verifies, and finally the decision-making and planning agent makes the final decision. Or, in some complex situations, multiple "agents" can work simultaneously, such as analyzing and verifying multiple fields at the same time.
[0090] In the embodiments of the present application, by simulating "multi-agent" collaboration inside the model, the complexity and resource consumption of actually deploying multiple physical models are avoided. Through the division of labor and cooperation of different "agents", it is possible to better handle complex electricity price data formats and improve the accuracy and flexibility of parsing.
[0091] In an alternative embodiment, S102 obtaining the feature representation vectors of each target field includes:
[0092] Step b1, constructing a semantic graph based on the target field.
[0093] In the embodiments of the present application, the WordNet is used to obtain the synonyms of each target field and construct a multi-layer semantic graph. Among them, the semantic graph is a graph structure, where the nodes represent words and the edges represent the semantic relationships between words.
[0094] In the embodiments of the present application, WordNet is a lexical database that organizes words according to semantic relationships. Each word has one or more "synonym sets", and the words in these synonym sets are semantically similar. Through WordNet, the synonyms of a word can be found. Suppose there is a target field "price", and its synonyms are found through WordNet. The synonyms of "price" may include: "cost", "tariff", "rate", etc. In the present application, not only the field name and its synonyms will be added, but also the related words (associated words) of these synonyms will be further searched to construct a multi-layer semantic graph.
[0095] For example:
[0096] Target field: price;
[0097] Synonyms: cost, tariff;
[0098] Associated words: The associated words of "cost" may be "expense", "budget"; the associated words of "tariff" may be "fee", "charge".
[0099] Step b2, obtaining the initial feature representation vectors of each word in the semantic graph.
[0100] In the embodiments of the present application, based on the word embedding method word2vec, the initial feature representation vectors of each word in the semantic graph are obtained.
[0101] Among them, word2vec maps words to vectors in a high-dimensional space. These vectors can capture the semantic features of words, so that words with similar semantics are closer in the vector space.
[0102] For example:
[0103] The initial feature representation vector of price: [0.1, 0.2, 0.3]
[0104] The initial feature representation vector of cost: [0.15, 0.25, 0.35]
[0105] The initial feature representation vector of tariff: [0.2, 0.3, 0.4]
[0106] Step b3, based on the graph attention network, perform an aggregation operation on the initial feature representation vectors to obtain the feature representation vectors of each target field.
[0107] Among them, the graph attention network (Graph Attention Networks, GAT) is a subcategory of graph neural networks, which can effectively learn the relationships between words of the same type, so as to effectively fuse information.
[0108] In the embodiments of the present application, a multi-layer GCN is used to gradually aggregate the information in the semantic graph.
[0109] First, the outer-layer GCN learns the relationships between related words and synonyms, and updates the feature representation vectors of synonyms.
[0110] For example:
[0111] The initial feature representation vector of cost: [0.15, 0.25, 0.35];
[0112] The feature representation vector of expense: [0.1, 0.2, 0.3];
[0113] The feature representation vector of budget: [0.12, 0.22, 0.32];
[0114] After aggregation by the outer-layer GCN, the new feature representation vector of cost may become: [0.14, 0.24, 0.34].
[0115] Then, the inner-layer GCN learns the relationships between synonyms and target fields, and updates the feature representation vectors of target fields.
[0116] For example:
[0117] The initial feature representation vector of price: [0.1, 0.2, 0.3];
[0118] The new feature representation vector of cost: [0.14, 0.24, 0.34];
[0119] The new feature representation vector of tariff: [0.21, 0.31, 0.41];
[0120] After aggregation by the inner - layer GCN, the final feature representation vector of price may become: [0.15, 0.25, 0.35]
[0121] For ease of understanding, it is described here in combination with Figure 2 The target field is price. Based on WordNet, synonyms and related words of the target field price are expanded to obtain electricity price (synonym), load (related word), and PV power generation (related word), and a multi - layer semantic graph is constructed to obtain the initial feature representation vectors of each word in the semantic graph. Then, based on GCN, the feature representation vector of the target field is updated to obtain the feature representation vector of the target field.
[0122] In the embodiments of the present application, the multi - layer structure of the semantic graph provides a panoramic view of semantic associations, while the aggregation mechanism of GCN effectively integrates local and global semantic features. Through this combined method, the model can capture richer semantic relationships, thereby obtaining the feature representation vector of the target field more accurately.
[0123] In the embodiments of the present application, assume that a new field rate is encountered when parsing electricity price data. It can be determined whether it is related to a known field (such as price) through the feature representation vector aggregated by the semantic graph and GCN. By comparing the feature representation vector of rate with that of price, if their distance is relatively close in the vector space, it can be determined that rate also represents the electricity price. Through this combined method, the model can represent the semantic features of the target field more accurately, thereby better handling complex semantic tasks.
[0124] Specifically: Use WordNet to query the synonyms and related words of the field rate. For example, the synonyms of rate may include: price, tariff, cost; the related words of rate may include: fee, charge, expense.
[0125] According to the query results, construct a multi - layer semantic graph of rate. This graph includes the target field rate, its synonyms and related words. Generate initial feature representation vectors for rate and its synonyms and related words. For example: The initial feature representation vector of rate: [0.2, 0.3, 0.4]; the initial feature representation vector of price: [0.1, 0.2, 0.3]; the initial feature representation vector of tariff: [0.2, 0.3, 0.4]; the initial feature representation vector of cost: [0.15, 0.25, 0.35]; the initial feature representation vector of fee: [0.25, 0.35, 0.45]; the initial feature representation vector of charge: [0.3, 0.4, 0.5]; the initial feature representation vector of expense: [0.1, 0.2, 0.3].
[0126] Aggregate the nodes in the semantic graph through multiple layers of GCN to update the feature representation vectors of each node. The outer layer of GCN aggregates the relationships between related words and synonyms to update the feature representation vectors of synonyms.
[0127] The updated feature representation vectors:
[0128] price: [0.12, 0.22, 0.32] (combining the information of cost and fee); tariff: [0.22, 0.32, 0.42] (combining the information of charge and expense).
[0129] The inner layer of GCN aggregates the relationships between synonyms and target fields to update the feature representation vectors of target fields.
[0130] The final feature representation vectors:
[0131] rate: [0.18, 0.28, 0.38] (combining the information of price and tariff)
[0132] Compare and analyze the final feature representation vector of rate with the feature representation vector of a known field (such as price).
[0133] For example, the feature representation vector of rate: [0.18, 0.28, 0.38]; the feature representation vector of price: [0.15, 0.25, 0.35]. Then, the semantic similarity between them can be judged by calculating the distance between the two vectors (such as the Euclidean distance). If the distance is relatively close, it can be considered that rate and price are semantically related.
[0134] It should be noted that the final feature representation vector of rate can also be compared and analyzed with the feature representation vector of a known field (such as price) through other technical means to judge whether they are semantically related and are of the same type of field.
[0135] In an optional embodiment, after S103 obtains the parsing result of each target field, it also:
[0136] Standardize the parsing result according to the preset output format of electricity price data to obtain a standardized parsing result.
[0137] In the embodiments of the present application, a standard output format of electricity price data needs to be predefined in advance. Each parsing result is converted to conform to the output format of electricity price data to ensure data consistency and availability, and facilitate user understanding and use.
[0138] Optionally, verify whether the converted data meets the requirements of the electricity price data output format, including checking the integrity, correctness and consistency of the data. For data that does not meet the electricity price data output format, error processing is performed. The verified and corrected data is output as a standardized parsing result to facilitate further processing and analysis, such as being directly used for data analysis, report generation, etc.
[0139] In an optional embodiment, the training process of the electricity price data parsing model includes:
[0140] Step c1, determining positive sample pairs and negative sample pairs based on the feature representation vector of each target field.
[0141] In the embodiment of the present application, the same type of field pairs (semantically similar target field pairs) are used as positive sample pairs. Different types of field pairs (semantically dissimilar target field pairs) are used as negative sample pairs. For example, if the fields price and cost both represent price information, they constitute a positive sample pair. The fields price (indicating price) and date (indicating date) constitute a negative sample pair.
[0142] Step c2: performing comparative learning on the positive sample pair and the negative sample pair to obtain an updated feature representation vector.
[0143] In the embodiment of the present application, a contrastive learning method is used to update the feature representation vector, which shortens the distance between positive sample pairs and pushes the distance between negative sample pairs in the feature representation space, thereby making the relationship between samples clearer and providing a better basis for subsequent classification tasks.
[0144] Step c3, determining the center of the category based on the updated feature representation vector.
[0145] In the embodiment of the present application, the method for determining the center of a category includes but is not limited to:
[0146] 1) Mean method: The coordinates of the category center are obtained by averaging the coordinates of all sample points in the same category.
[0147] 2) Median method: Determine the category center by finding the median of all sample points within the category.
[0148] 3) Centroid method: The category center is calculated by considering the weighted average of each feature dimension.
[0149] 4) Weighted average method: By assigning different weights to each data point, the weighted average is calculated as the category center.
[0150] 5) K-means clustering: By iteratively adjusting the center position of the category, the distance from each sample point to the center of its category is minimized.
[0151] 6) Density clustering: Classes are divided according to the density between sample points, and the class center can be defined as the central position of the area with the highest sample point density.
[0152] Step c4, calculate the distance of each sample to the center of the said class.
[0153] In the embodiments of the present application, a suitable distance metric method can be selected to calculate the distance between each sample and the class center. The distance metric methods include but are not limited to: Euclidean distance, Manhattan distance, cosine similarity, Mahalanobis distance.
[0154] Step c5, based on the said distance and the dynamic decision boundary radius of the class, determine the probability that the sample belongs to each class.
[0155] In the embodiments of the present application, in a multi-classification problem, when the sample is represented by a high-dimensional vector, it may occur that the sample is closer to the center of a certain class but actually belongs to another class. To solve this problem, the present application introduces the concept of a dynamic decision boundary, that is, a radius is set for each class (i.e., the above-mentioned dynamic decision boundary radius) to determine whether the sample belongs to the class.
[0156] The formula for calculating the probability that the sample of the present application belongs to each class is:
[0157]
[0158] where p(y = i|x) represents the probability that the sample x belongs to class i; d i represents the distance from the sample to the center of class i; r i represents the dynamic decision boundary radius of class i; C represents the total number of classes; d j represents the distance from the sample x to the center of class j; r j represents the dynamic decision boundary radius of class j; exp represents the exponential function, which is used to calculate the exponential part of the probability.
[0159] This formula combines two factors:
[0160] The distance from the sample to the class center: The smaller the distance, the greater the classification probability.
[0161] The ratio of the sample to the radius: When the sample distance is close to the radius of the class, the weight will be amplified; when the sample distance far exceeds the radius of the class, the weight will be significantly reduced.
[0162] Therefore, when the distance from a sample to a certain class center exceeds the radius of that class, even if the distance to that class is relatively small, the model will not easily select it because the weight of that class will decrease. This method makes the classification more flexible and accurate by dynamically adjusting the decision boundary.
[0163] To ensure that samples are better distributed within the dynamic decision boundaries of their respective classes, a concept of dynamic decision boundary is introduced in the embodiments of this application, with the radius serving as the controller of the decision boundary. The core idea of this method is to consider the radius of each class as an optimizable parameter and adjust it according to the feedback of the loss function L intra where the loss function L of the electricity price data parsing model intra is:
[0164] where N is the number of samples in one round of training iteration, x i is the feature representation vector of the i-th sample, y i is the class to which sample i belongs, is the feature representation vector of the center of the class y to which sample i belongs i , and is the radius of the dynamic decision boundary of the class y to which sample i belongs. When the distance d i of the sample is less than , the model will consider that the sample is reasonably distributed and the loss is small. If the distance d i of the sample is greater than the radius, the loss will increase significantly, forcing the model to adjust the radius of the class to better enclose the sample distribution of that class. When the distance d of the sample is greater than the radius, the loss will increase significantly, forcing the model to adjust the radius of the class to better enclose the sample distribution of that class. i If the distance of the sample from
[0165] During training, the radius of each class is adjusted according to the feedback of the loss function. If the distance from the sample to the class center exceeds the class radius, the loss is increased, forcing the model to adjust the radius to better enclose the sample distribution. The radius is adjusted through an optimization algorithm (such as gradient descent) to minimize the loss function.
[0166] Based on the above method, by ensuring that samples are within their respective decision boundaries, classification errors are reduced and classification accuracy is improved. The dynamic decision boundary enables the model to adapt to data distributions with different densities and shapes, enhancing the adaptability of the model. By adjusting the radius, the model can better understand and process the distribution of samples, thereby optimizing the sample distribution. Based on the above method, the electricity price data parsing model can more flexibly adapt to the sample distributions of different classes, thereby improving the accuracy and efficiency of parsing. This method is not only applicable to electricity price data parsing but can also be extended to other application scenarios that need to process diverse data distributions.
[0167] Optionally, after obtaining the parsing results of each target field, continuously adjust the weight distribution between the generation rules of the large language model and the electricity price data parsing model (which can be a random forest model) according to the accuracy of the parsing results of the target fields. If the generation rules of the large language model perform better in certain cases, adjust and increase its weight; if the electricity price data parsing model is more accurate in other cases, adjust and increase its weight. Repeat the steps of evaluation and adjustment, and continuously iterate and optimize the model. In each iteration, retrain the model according to the latest weight distribution and evaluate the performance on the validation set. Enable the electricity price data parsing model to dynamically adjust its strategy according to the actual parsing results, so as to maintain efficient and accurate parsing capabilities when facing diverse and changing electricity price data.
[0168] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0169] Corresponding to the electricity price data parsing method described in the above embodiments, Figure 2 The structural block diagram of the electricity price data parsing system provided by the embodiments of the present application is shown. For the convenience of description, only the parts related to the embodiments of the present application are shown.
[0170] Referring to Figure 3 , the electricity price data parsing system includes:
[0171] A parsing and classification module, configured to parse and classify the obtained electricity price data to obtain at least one target field;
[0172] A feature acquisition module, configured to acquire the feature representation vector of each target field;
[0173] A parsing result determination module, configured to input the feature representation vector of each target field into a preset electricity price data parsing model to obtain the parsing result of each target field.
[0174] In a possible implementation manner, the parsing and classification module is configured to:
[0175] Obtain the electricity price data through the API interface;
[0176] Use the large language model to parse the electricity price data to obtain the initial target fields;
[0177] Classify the initial target fields to obtain at least one of the target fields.
[0178] In a possible implementation manner, the feature acquisition module is configured to:
[0179] Construct a semantic graph based on the target field;
[0180] Obtain the initial feature representation vectors of each word in the semantic graph;
[0181] Based on the graph attention network, perform an aggregation operation on the initial feature representation vectors to obtain the feature representation vectors of each target field.
[0182] In a possible implementation, the electricity price data parsing system further includes:
[0183] A normalization module, configured to perform normalization processing on the parsing result according to a preset output format of the electricity price data to obtain a normalized parsing result.
[0184] In a possible implementation, the training process of the electricity price data parsing model includes:
[0185] Determine positive sample pairs and negative sample pairs based on the feature representation vectors of each target field;
[0186] Perform contrastive learning on the positive sample pairs and the negative sample pairs to obtain updated feature representation vectors;
[0187] Determine the centers of the categories based on the updated feature representation vectors;
[0188] Calculate the distance of each sample pair from the center of the category;
[0189] Based on the distance and the dynamic decision boundary radius of the category, determine the probability that the sample belongs to each category.
[0190] In a possible implementation, the loss function L of the electricity price data parsing model intra is:
[0191] where is the number of samples in one round of training iteration, x i is the feature representation vector of the i-th sample, y i is the category to which the sample i belongs, is the feature representation vector of the center of the category y to which the sample i belongs i and is the dynamic decision boundary radius of the category y to which the sample i belongs. is the feature representation vector of the center of the category y to which the sample i belongs i and is the dynamic decision boundary radius of the category y to which the sample i belongs.
[0192] It should be noted that for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiments of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0193] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0194] An embodiment of this application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the foregoing method embodiments are implemented.
[0195] An embodiment of this application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in each of the foregoing method embodiments can be implemented.
[0196] An embodiment of this application provides a computer program product. When the computer program product runs on a computer device, the computer device is caused to execute the steps in each of the foregoing method embodiments.
[0197] Figure 3 FIG. is a schematic structural diagram of a computer device provided in an embodiment of this application. As Figure 3 shown, the computer device in this embodiment includes: at least one processor 20 ( Figure 3 only one is shown in the figure), a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20. When the processor 20 executes the computer program 22, the steps in any of the foregoing method embodiments for parsing electricity price data are implemented.
[0198] The computer device may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art can understand that Figure 3This is only an example of a computer device and does not limit the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0199] The so-called processor 20 may be a central processing unit (CPU). The processor 20 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0200] In some embodiments, the memory 21 may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 21 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory 21 may also include both the internal storage unit and the external storage device of the computer device. The memory 21 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 21 may also be used to temporarily store data that has been output or will be output.
[0201] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / computer equipment, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, USB flash drive, mobile hard disk, magnetic disk, or optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be electrical carrier signals and telecommunication signals.
[0202] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0203] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0204] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0205] The unit described as a separating component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0206] The foregoing embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the same; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for parsing electricity price data, characterized in that, Including: Parsing and classifying the obtained electricity price data to obtain at least one target field; Obtaining the feature representation vector of each target field; Inputting the feature representation vector of each target field into a preset electricity price data parsing model to obtain the parsing result of each target field.
2. The electricity price data parsing method according to claim 1, wherein The parsing and classifying the obtained electricity price data to obtain at least one target field includes: Obtaining the electricity price data through the API interface; Parsing the electricity price data using a large language model to obtain the initial target field; Classifying the initial target field to obtain at least one of the target fields.
3. The electricity price data parsing method according to claim 2, wherein The obtaining the feature representation vector of each target field includes: Constructing a semantic graph based on the target field; Obtaining the initial feature representation vector of each word in the semantic graph; Performing an aggregation operation on the initial feature representation vector based on a graph attention network to obtain the feature representation vector of each target field.
4. The electricity price data parsing method according to claim 1, characterized in that After obtaining the parsing result of each target field, further: Performing a normalization process on the parsing result according to a preset electricity price data output format to obtain a normalized parsing result.
5. The electricity price data parsing method according to any one of claims 1 to 4, characterized in that The training process of the electricity price data parsing model includes: Determining positive sample pairs and negative sample pairs based on the feature representation vector of each target field; Performing contrastive learning on the positive sample pairs and the negative sample pairs to obtain updated feature representation vectors; Determining the center of the category based on the updated feature representation vector; Calculating the distance of each sample pair from the center of the category; Determining the probability of each sample belonging to each category based on the distance and the dynamic decision boundary radius of the category.
6. The electricity price data parsing method according to claim 5, wherein The loss function L of the electricity price data parsing model intra is as follows: Among them, N is the number of samples in one round of training iteration, x i is the feature representation vector of the i-th sample, y i is the category to which sample i belongs, is the feature representation vector of the center of the category y to which sample i belongs i ; is the dynamic decision boundary radius of the category y to which sample i belongs i .
7. A electricity price data parsing system, characterized in that, Including: A parsing and classification module for parsing and classifying the obtained electricity price data to obtain at least one target field; A feature acquisition module for obtaining the feature representation vector of each target field; A parsing result determination module for inputting the feature representation vector of each target field into a preset electricity price data parsing model to obtain the parsing result of each target field.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, When the computer program product runs on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 6.