Data aggregation method and system based on energy big data
By classifying and prioritizing energy data, and using mapping vectors to indicate the association relationship between data and archive requirements, the problem of low accuracy in energy big data processing systems is solved, and more accurate data aggregation and archiving is achieved.
Patent Information
- Application Number
- CN202510177793.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing energy big data processing systems have low accuracy when processing unstructured data and analyzing the correlation between different energy data, resulting in inaccurate data analysis and archive results.
By obtaining modal information of energy data, classifying and determining categories, determining priority based on archive requirements and categories, and using mapping vectors to indicate the association relationship between data and archive requirements, and finally performing structured processing.
It improves the data aggregation accuracy of energy big data, can quantify the importance of different modes and types of energy data in the energy industry, and improves the accuracy of data archiving.
Smart Images

Figure CN119646276B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of energy data analysis, and particularly to a data aggregation method and system based on energy big data. Background Art
[0002] At present, with the development of big data technology, more and more enterprises and organizations tend to use data analysis methods to assist and supervise the process nodes and production processes of traditional manufacturing industries, so as to improve production efficiency. In the energy industry, energy big data processing systems are usually used to analyze and archive structured data. However, this method has insufficient processing ability for unstructured data and lacks the analysis of the correlation between different energy data, resulting in low accuracy of data analysis and data archiving results. Summary of the Invention
[0003] In view of the above, it is necessary to propose a data aggregation method and system based on energy big data to solve the technical problem of low accuracy of data aggregation based on energy big data.
[0004] This application provides a data aggregation method based on energy big data, which is applied to an electronic device. The method includes: obtaining energy data to be archived and obtaining modal information corresponding to the energy data; classifying the energy data according to the modal information to obtain a category of the energy data, where the category is used to represent the position of the energy data in the energy industry; determining an archiving priority of the energy data according to pre-stored archiving requirements and the category, where the archiving priority is used to indicate the importance of the energy data in the energy industry; determining a mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modal information, where the mapping vector is used to indicate the association relationship between the energy data and the archiving requirements; and archiving the energy data according to the mapping vector.
[0005] In some embodiments, the classifying the energy data according to the modal information to obtain a category of the energy data includes: encoding the energy data to obtain a first encoding; encoding the modal information to obtain a second encoding; determining an encoding vector corresponding to the energy data according to the first encoding and the second encoding; and inputting the encoding vector into a pre-trained first model to obtain a category vector of the energy data, where the category vector is used to indicate the category of the energy data.
[0006] In some embodiments, the method further includes training the first model, and the training of the first model includes: encoding pre-acquired sample energy data and corresponding sample modality information to obtain a sample encoding; determining the confidence of the sample encoding according to the sample energy data and the sample modality information; based on the sample encoding, determining a corresponding predicted category vector according to a pre-constructed initial classification model; determining a first loss value of the initial classification model according to the label category vector corresponding to the sample encoding, the predicted category vector, and the confidence; continuously updating the initial classification model based on the backpropagation algorithm until the first loss value meets a preset condition, and stopping updating the initial classification model to obtain the first model.
[0007] In some embodiments, the determining the confidence of the sample encoding according to the sample energy data and the sample modality information includes: performing semantic analysis on the energy data to obtain semantic information of the energy data; determining the matching degree between the semantic information and the sample modality information; and determining the confidence of the sample encoding according to the matching degree.
[0008] In some embodiments, the determining the first loss value of the initial classification model according to the label category vector corresponding to the sample encoding, the predicted category vector, and the confidence includes: ; where Loss1 represents the first loss value of the initial classification model; S represents the confidence; i represents the index of the category; n represents the number of categories; A i represents the probability value corresponding to the i-th category in the label category vector; B i represents the probability value corresponding to the i-th category in the predicted category vector.
[0009] In some embodiments, the determining the mapping vector between the energy data and the filing requirement according to the category, the priority, and the modality information includes: generating prompt data based on the category, the modality information, and a pre-stored prompt template; inputting the prompt data into a pre-trained generative model to obtain an initial mapping vector between the energy data and the filing requirement; and determining the mapping vector between the energy data and the filing requirement based on the priority and the initial mapping vector.
[0010] In some embodiments, the method further includes training the generative model, and the training of the generative model includes: generating sample prompt data based on pre-acquired label categories, sample modality information, and pre-stored prompt templates; inputting the sample prompt data into a pre-constructed initial generative model to obtain a prediction mapping vector between the energy data and the archiving requirements; determining a second loss value of the initial generative model according to the label mapping vector corresponding to the sample modality information and the prediction mapping vector; updating the initial generative model based on the backpropagation algorithm until the second loss value meets a preset condition, and stopping updating the initial generative model to obtain the generative model.
[0011] In some embodiments, the determining the second loss value of the initial generative model according to the label mapping vector corresponding to the sample modality information and the prediction mapping vector includes: ; where Loss2 represents the second loss value of the initial generative model; j represents the index of the dimension in the label mapping vector; m represents the number of dimensions in the label mapping vector; C j represents the mapping value corresponding to the j-th dimension in the label mapping vector; D j represents the mapping value corresponding to the j-th category in the prediction mapping vector.
[0012] An embodiment of the present application further provides a data aggregation system based on energy big data. The system includes an electronic device and a server; the electronic device is configured to obtain energy data to be archived and obtain the modality information corresponding to the energy data; the electronic device is configured to classify the energy data according to the modality information to obtain the category of the energy data; where the category is used to characterize the position of the energy data in the energy industry; the electronic device is configured to determine the archiving priority of the energy data according to the pre-stored archiving requirements and the category, and the archiving priority is used to indicate the importance of the energy data in the energy industry; the electronic device is configured to determine a mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modality information; where the mapping vector is used to indicate the association relationship between the energy data and the archiving requirements; the electronic device is configured to archive the energy data according to the mapping vector.
[0013] In some embodiments, the electronic device is further configured to encode the energy data to obtain a first encoding; the electronic device is configured to encode the modality information to obtain a second encoding; the electronic device is configured to determine an encoding vector corresponding to the energy data according to the first encoding and the second encoding; the electronic device is configured to input the encoding vector into a pre-trained first model to obtain a class vector of the energy data; wherein, the class vector is used to indicate the class of the energy data.
[0014] As can be seen from the above technical solutions, in the embodiments of the present application, the energy data to be archived is classified according to the modality information of the energy data to obtain the class of the energy data, so as to characterize the position of the energy data to be archived in the energy industry based on the quantization data, providing support for subsequent data archiving. Then, the archiving priority of the energy data is determined according to the archiving requirements and the class to indicate the importance degree of the energy data in the energy industry, and the association relationship between the energy data and the archiving requirements can be corrected based on the archiving priority. Then, according to the class, priority, and modality information, a mapping vector between the energy data and the archiving requirements is determined, so as to indicate the association relationship between the energy data and the archiving requirements based on the quantization data. Finally, the energy data is archived according to the mapping vector. In this way, the importance degree of different modalities and different types of energy data can be determined during the data archiving process, and the energy data can be structurally processed based on the mapping vector, which can improve the accuracy of data aggregation based on energy big data. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 FIG. is an application scenario diagram of a data aggregation method based on energy big data provided by an embodiment of the present application.
[0016] Figure 2 FIG. is a flowchart of a data aggregation method based on energy big data provided by an embodiment of the present application.
[0017] Figure 3 FIG. is a flowchart of a method for training a first model provided by an embodiment of the present application.
[0018] Figure 4 FIG. is a flowchart of a method for training a second model provided by an embodiment of the present application.
[0019] Figure 5 FIG. is a functional module diagram of a data aggregation system based on energy big data provided by an embodiment of the present application.
[0020] Figure 6 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to more clearly understand the purpose, features and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. In the following description, many specific details are set forth in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0022] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0024] The embodiment of the present application provides a data aggregation method based on energy big data, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0025] The electronic device can be any electronic product that can perform human-computer interaction with customers. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet protocol television (IPTV), a smart wearable device, etc.
[0026] The electronic device may further include a network device and / or a client device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on Cloud Computing.
[0027] The network where the electronic device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, Virtual Private Network (VPN), etc.
[0028] As Figure 1 Shown is an application scenario diagram of a data aggregation method based on energy big data provided by an embodiment of the present application. A data aggregation method based on energy big data provided by the present application can be applied to the electronic device 100. Among them, the electronic device 100 is communicatively connected to the server 200. Among them, the server 200 is used to store the energy data to be archived and the corresponding modality information. The electronic device 100 is used to obtain the energy data to be archived and obtain the modality information corresponding to the energy data; classify the energy data according to the modality information to obtain the category of the energy data; where the category is used to represent the position of the energy data in the energy industry. The electronic device 100 is further used to determine the archiving priority of the energy data according to the pre-stored archiving requirements and the category, and the archiving priority is used to indicate the importance of the energy data in the energy industry; determine the mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modality information; where the mapping vector is used to indicate the association relationship between the energy data and the archiving requirements; archive the energy data according to the mapping vector.
[0029] As Figure 2 Shown is a flowchart of a data aggregation method based on energy big data provided by an embodiment of the present application. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted. A data aggregation method based on energy big data provided by an embodiment of the present application includes the following steps.
[0030] S20, obtain the energy data to be archived and obtain the modality information corresponding to the energy data.
[0031] In an embodiment of the present application, the energy data to be archived is used to record and describe information in multiple fields of the energy industry. Specifically, the energy data can be energy production and supply data, such as the production volume, import and export volume, reserve volume, etc. of various energy sources. The energy production data can be used to characterize the extraction and production of energy resources, and the supply data can be used to characterize the supply capacity and stability of energy. The energy data can also be energy consumption and demand data, such as the energy consumption volume, energy consumption structure, energy demand forecast, etc. of each department and industry. The energy consumption data can be used to evaluate the energy utilization efficiency, guide energy conservation and optimize the energy structure. The energy data can also be energy price data, such as the market price, price trend, price index, etc. of energy. The energy price data has important reference significance for the operation of the energy market and the business decisions of energy enterprises. The energy data can also be energy emission data, such as the carbon dioxide emission volume, pollutant emission volume, etc. of energy. The energy emission data is an important indicator for measuring the impact of the energy industry on the environment and evaluating environmental policies. The energy data can also be energy policy and technology data, such as the energy policies and regulations of various countries, the research and development and application status of energy technologies, etc. The energy policy and technology data can help analyze the policy environment of the energy market and the impact of technological progress. In addition, from the perspective of data types, energy data is usually collected and recorded through energy equipment, sensors, measuring instruments, etc. It includes data on aspects such as the production, transmission, distribution, and use of energy. These data usually have the characteristics of time-series data, that is, they are arranged in chronological order and can record and reflect the characteristics and trends that change over time. From the application scenario perspective, energy data can be used in multiple aspects such as energy production optimization, energy consumption monitoring, energy demand forecasting, energy conservation and emission reduction monitoring, energy market analysis, intelligent energy management systems, energy supply chain optimization, and energy efficiency evaluation.
[0032] In an embodiment of the present application, energy data of multiple modalities refers to data obtained for the same description object in the energy industry from different fields or perspectives. Among them, the energy data of multiple modalities covers various types, and each type of data can be regarded as a modality. Specifically, the modality information corresponding to the energy data can be text modality, including various energy-related reports, documents, records, etc., such as energy policy documents, equipment operation logs, energy transaction records, etc., which can provide background information, rule descriptions, and historical data in the energy field. The modality information can also be image modality. In the energy field, image data usually comes from surveillance cameras, drone shots, satellite images, etc. For example, the energy data of image modality can be energy equipment images taken of power lines, oil and gas pipelines, wind turbines, etc., which are used to monitor the equipment status and identify potential safety hazards. The modality information can also be video modality. The energy data of video modality can be a continuous sequence of images, which is used to provide relatively rich image information in the time dimension and can be used to monitor equipment operation, personnel operations, etc., as well as for the recording and retrospective analysis of security incidents. The modality information can also be audio modality, including equipment sound monitoring, environmental noise analysis, etc. For example, by analyzing the operating sound of the equipment, it can be determined whether the equipment is operating normally or has a fault. The modality information can also be sensor modality. The energy data of sensor modality includes measurement values of various physical quantities such as temperature, pressure, flow rate, vibration, etc., which can be used to monitor the operating status of the equipment in real time and warn of potential faults.
[0033] S21, classify the energy data according to the modality information to obtain the category of the energy data; wherein, the category is used to characterize the position of the energy data in the energy industry.
[0034] In an embodiment of the present application, since the energy data to be archived corresponds to multiple modalities and the energy data to be archived includes structured and unstructured data, in order to determine the quantitative relationship between each energy data during the archiving of energy data and facilitate ensuring the accuracy of data retrieval results or data mining results in subsequent data retrieval and data mining processes, the energy data can be classified according to the energy data and the corresponding modality information first to determine the position of the energy data in the energy industry and provide support for quantitative data for the archiving of energy data. Exemplarily, the category of the energy data can be upstream, midstream or downstream. When the category of the energy data is upstream, it indicates that the energy data is generated by the upstream industry in the energy industry; when the category of the energy data is downstream, it indicates that the energy data is generated by the downstream industry in the energy industry.
[0035] Exemplarily, the modal information corresponding to the energy production data includes, but is not limited to, sensor data (such as temperature, pressure, flow rate, etc.), image data (such as the operating status of energy equipment), sound data (such as abnormal sounds during the operation of energy equipment), etc. Since the energy production data comes from the source of the energy industry (such as coal mining, oil refining, wind power generation, solar power generation, etc.), the energy production data is usually located in the upstream position of the energy industry and can reflect the energy production capacity. By analyzing the energy production data, the efficiency, cost, and potential risk points of energy production can be obtained.
[0036] In an embodiment of the present application, when performing data aggregation based on energy big data, different modalities and different types of energy data can be integrated to provide data support for real-time monitoring of energy production and prediction of demand changes.
[0037] In order to indicate the position of the energy data in the energy industry according to the category of the energy data, the category of the energy data can be determined based on a pre-trained first model. Specifically, the energy data is classified according to the modal information, and the category of the energy data obtained includes: encoding the energy data to obtain a first encoding; encoding the modal information to obtain a second encoding; determining an encoding vector corresponding to the energy data according to the first encoding and the second encoding; inputting the encoding vector into the pre-trained first model to obtain a category vector of the energy data; wherein, the category vector is used to indicate the category of the energy data.
[0038] Specifically, the energy data and the modality information can be encoded respectively according to a preset encoding algorithm to obtain a first encoding corresponding to the energy data and a second encoding corresponding to the modality information. The preset encoding algorithm can be a bag-of-words model, or a term frequency-inverse document frequency algorithm, or a word embedding algorithm, which is not limited in this application. The first encoding and the second encoding can be concatenated to obtain an encoding vector corresponding to the energy data, so as to ensure that the encoding vector represents the complete information in the energy data and the modality information; alternatively, the concatenated first encoding and second encoding can be input into a pre-trained autoencoder model, and the vector output by the autoencoder model can be used as the encoding vector corresponding to the energy data. In this way, the dimensions of the first encoding and the second encoding can be reduced, and it can also be ensured that the encoding vector contains the key information in the energy data and the modality information, thereby improving the efficiency of classifying the energy data while ensuring the subsequent classification accuracy. Among them, the first model can be a pre-trained model with data classification function, and the input data of the first model is the encoding vector corresponding to the energy data, and the output data of the first model is the category vector corresponding to the energy data. The category vector has numerical values in multiple dimensions, each dimension corresponding to a category, and the numerical value of any one dimension is used to represent the probability that the energy data belongs to the category corresponding to that dimension. Specifically, the category corresponding to the highest numerical value in the category vector can be determined as the category corresponding to the energy data.
[0039] In an embodiment of the present application, in order to improve the accuracy of classifying energy data, the first model can be trained according to the pre-acquired sample data. The sample data includes sample energy data and corresponding sample modality information, and the sample data also includes corresponding label category vectors. The label category vector is used to indicate the category of the sample energy data. Specifically, for the specific method of training the first model according to the sample data, please refer to Figure 3 the corresponding detailed description.
[0040] S22. Determine the archiving priority of the energy data according to the pre-stored archiving requirements and the category, and the archiving priority is used to indicate the importance of the energy data in the energy industry.
[0041] In an embodiment of the present application, determining the archiving priority of the energy data according to the pre-stored archiving requirements and the category of the energy data includes: performing semantic analysis on the archiving requirements to obtain a semantic encoding corresponding to the archiving requirements; determining the similarity between the semantic encoding and the category vector; and determining the archiving priority of the energy data according to the similarity. The archiving requirement can be the requirement to retain multi-modal or multi-dimensional data for a long time in the data management process. The archiving requirement is usually based on multiple aspects, including energy industry cost control, energy system performance optimization, energy data security and compliance, and future data access and utilization, etc.
[0042] S23. Determine a mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modality information, where the mapping vector is used to indicate the association between the energy data and the archiving requirements.
[0043] In an embodiment of the present application, determining the mapping vector between the energy data and the archiving requirements according to the category of the energy data, the priority of the energy data, and the modality information corresponding to the energy data includes: generating prompt data based on the category, the modality information, and a pre-stored prompt template; inputting the prompt data into a pre-trained generative model to obtain an initial mapping vector between the energy data and the archiving requirements; and determining the mapping vector between the energy data and the archiving requirements based on the priority and the initial mapping vector.
[0044] The pre-stored prompt template is used to format the category and the modality information of the energy data. Exemplarily, the content of the prompt template may be "Task prompt: Please determine the mapping relationship between the energy data and the archiving requirements according to the following category information and modality information; Output prompt: Please output a mapping vector to represent the above mapping relationship; Category information:; Modality information:; Archiving requirements:; Priority:". The generative model may be a GPT model or an LLM model, and the present application does not limit this.
[0045] In an embodiment of the present application, the input data of the generative model is the prompt data, and the output data of the generative model is the initial mapping vector, which is used to represent the association between the energy data predicted by the generative model and the archiving requirements. To improve the accuracy of the association, the initial mapping vector may be corrected according to the priority corresponding to the energy data. Specifically, the product of the priority and the initial mapping vector may be determined as the mapping vector corresponding to the energy data.
[0046] S24. Archive the energy data according to the mapping vector.
[0047] In an embodiment of the present application, for any archiving requirement, the information in the archiving requirement (such as text information, image information, etc.) may be confirmed as the first node; for any energy data and the corresponding modality information, the energy data and the modality information may be confirmed as the second node; and the mapping vector may be determined as the edge weight between the first node and the second node. In this way, the energy data can be archived in a structured manner, thereby improving the accuracy of data aggregation based on energy big data.
[0048] As can be seen from the above technical solutions, in the embodiments of the present application, the energy data to be archived is classified according to the modal information of the energy data to obtain the category of the energy data, so as to characterize the position of the energy data to be archived in the energy industry based on the quantitative data, providing support for subsequent data archiving. Then, the archiving priority of the energy data is determined according to the archiving requirements and the category to indicate the importance of the energy data in the energy industry, and the association relationship between the energy data and the archiving requirements can be corrected based on the archiving priority. Then, according to the category, priority, and modal information, the mapping vector between the energy data and the archiving requirements is determined, so as to indicate the association relationship between the energy data and the archiving requirements based on the quantitative data. Finally, the energy data is archived according to the mapping vector. In this way, the importance of different modal and different types of energy data can be determined during the data archiving process, and the energy data can be structured based on the mapping vector, which can improve the accuracy of data aggregation based on energy big data.
[0049] As Figure 3 shown, it is a flowchart of a method for training a first model provided by an embodiment of the present application. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted. The method for training a first model provided by the embodiments of the present application includes the following steps.
[0050] S30. Encode the pre-acquired sample energy data and the corresponding sample modal information to obtain a sample code.
[0051] In an embodiment of the present application, the method for encoding the pre-acquired sample energy data and the corresponding sample modal information is the same as the method for encoding the energy data and the modal information in step S21, and will not be described in detail here.
[0052] S31. Determine the confidence level of the sample code according to the sample energy data and the sample modal information.
[0053] In an embodiment of the present application, determining the confidence level of a sample code based on sample energy data and sample modal information includes: performing semantic analysis on the energy data to obtain the semantic information of the energy data; determining the matching degree between the semantic information and the sample modal information; and determining the confidence level of the sample code according to the matching degree. Among them, the semantic information of the energy data can be quantitative data, which is used to characterize the industrial information contained in the energy data. For example, the industrial information in the energy data that the semantic information can characterize, where the industrial information can characterize the energy supply volume and resource reserve situation, can also characterize the energy consumption of different industries, and can also characterize the energy price trend, etc. Specifically, a natural language processing model can be used to analyze the energy data to obtain the semantic information of the energy data, and the matching degree between the semantic information and the modal information can be determined according to a preset matching degree calculation method. Among them, the preset matching degree calculation method can be the Euclidean distance algorithm, can also be the cosine similarity algorithm, and can also be the Hamming distance algorithm. The present application does not make any limitations in this regard.
[0054] In an embodiment of the present application, the confidence level of the sample code can be determined according to the matching degree between the semantic information and the modal information. Among them, the higher the matching degree between the semantic information and the modal information, the higher the similarity between the information contained in the energy data and its corresponding sample modal information, and the higher the confidence level of the sample energy code corresponding to the sample energy data; the lower the matching degree between the semantic information and the modal information, the lower the similarity between the information contained in the energy data and its corresponding sample modal information, and the lower the confidence level of the sample energy code corresponding to the sample energy data. Specifically, the matching degree between the semantic information and the sample modal information can be determined as the confidence level of the corresponding sample code.
[0055] S32. Based on the sample code, determine the corresponding predicted category vector according to the pre-constructed initial classification model.
[0056] In an embodiment of the present application, the sample code can be input into the pre-constructed initial classification model to obtain the predicted category vector output by the initial classification model. Among them, the predicted category vector has numerical values in multiple dimensions, and each dimension corresponds to a category, and the numerical value of any one dimension is used to characterize the probability that the sample energy data corresponding to the sample code belongs to the category corresponding to this dimension. Specifically, the category corresponding to the highest numerical value in the predicted category vector can be determined as the category corresponding to the sample energy data. Among them, the initial classification model can be a multi-layer perceptron model, can also be a convolutional neural network model, and can also be a recurrent neural network model. The present application does not make any limitations in this regard.
[0057] S33. Determine the first loss value of the initial classification model according to the label category vector corresponding to the sample code, the predicted category vector, and the confidence level.
[0058] In an embodiment of the present application, in order to quantitatively evaluate the performance of the initial classification model, the first loss value of the initial classification model can be determined according to the label category vector, prediction category vector, and confidence corresponding to the sample encoding. And based on the first loss value, the performance of the initial classification model is evaluated, so as to provide data support for subsequent updating of the initial classification model. Specifically, the method for determining the first loss value satisfies the following relational expression:
[0059] ;
[0060] where Loss1 represents the first loss value of the initial classification model; S represents the confidence; i represents the index of the category; n represents the number of categories; A i represents the probability value corresponding to the i-th category in the label category vector; B i represents the probability value corresponding to the i-th category in the prediction category vector.
[0061] In an embodiment of the present application, the higher the first loss value, the higher the degree of difference between the prediction category vector output by the initial classification model and the label category vector, and the lower the accuracy of the output result of the initial classification model.
[0062] S34, continuously update the initial classification model based on the backpropagation algorithm until the first loss value meets the preset condition, and then stop updating the initial classification model to obtain the first model.
[0063] In an embodiment of the present application, the initial classification model can be continuously updated based on the backpropagation algorithm until the first loss value meets the preset condition, and then stop updating the initial classification model to obtain the first model for classifying energy data. Among them, the preset condition can be that the first loss value is less than or equal to a preset termination threshold. When the first loss value is less than or equal to the preset termination threshold, it indicates that the degree of difference between the prediction category vector output by the initial classification model and the label category vector is small, and the accuracy of the output result of the initial classification model is high. Exemplarily, the termination threshold can be 0.1.
[0064] As Figure 4 shown, it is a flowchart of a method for training a second model provided by an embodiment of the present application. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted. The method for training a second model provided by the embodiment of the present application includes the following steps.
[0065] S40, generate sample prompt data based on the pre-acquired label category, sample modality information, and pre-stored prompt templates.
[0066] In an embodiment of the present application, the method for generating sample prompt data based on the pre-acquired tag categories, sample modality information, and the pre-stored prompt templates is the same as the method for obtaining prompt data in step S23, and will not be elaborated here.
[0067] S41. Input the sample prompt data into the pre-constructed initial generative model to obtain a predicted mapping vector between the energy data and the archiving requirements.
[0068] In an embodiment of the present application, the sample prompt data can be input into the pre-constructed initial generative model to obtain a predicted mapping vector output by the initial generative model. Among them, the predicted mapping vector is used to quantitatively represent the correlation between the energy data predicted by the initial generative model and the archiving requirements. Among them, the initial generative model can be a GPT model, or an LLM, etc., and the present application does not make any limitations in this regard.
[0069] S42. Determine a second loss value of the initial generative model according to the tag mapping vector corresponding to the sample modality information and the predicted mapping vector.
[0070] In an embodiment of the present application, in order to quantitatively evaluate the performance of the initial generative model, a second loss value of the initial generative model can be determined according to the tag mapping vector corresponding to the sample modality information and the predicted mapping vector. And evaluate the performance of the initial generative model based on the second loss value, so as to provide data support for subsequent updating of the initial generative model. Specifically, the method for determining the second loss value satisfies the following relational expression:
[0071] ;
[0072] Among them, Loss2 represents the second loss value of the initial generative model; j represents the index of the dimension in the tag mapping vector; m represents the number of dimensions in the tag mapping vector; C j represents the mapping value corresponding to the j-th dimension in the tag mapping vector; D j represents the mapping value corresponding to the j-th category in the predicted mapping vector.
[0073] In an embodiment of the present application, the higher the second loss value, the higher the degree of difference between the predicted mapping vector output by the initial generative model and the tag mapping vector, and the lower the accuracy of the output result of the initial generative model.
[0074] S43. Update the initial generative model based on the backpropagation algorithm until the second loss value meets the preset conditions, and stop updating the initial generative model to obtain the generative model.
[0075] In an embodiment of the present application, the initial generative model can be continuously updated based on the backpropagation algorithm until the second loss value meets a preset condition, at which point the update of the initial generative model stops, and a second model for predicting the correlation between energy data and archiving requirements is obtained. Herein, the preset condition may be that the second loss value is less than or equal to a preset termination threshold. When the second loss value is less than or equal to the preset termination threshold, it indicates that the difference between the predicted mapping vector output by the initial generative model and the label mapping vector is small, and thus the accuracy of the output result of the initial generative model is high. Exemplarily, the termination threshold may be 0.1.
[0076] Please refer to Figure 5 , Figure 5 FIG. is a functional block diagram of a data aggregation system based on energy big data provided by an embodiment of the present application. The data aggregation system 500 based on energy big data includes an electronic device 100 and a server 200. The modules / units referred to in the present application refer to a series of computer-readable instruction segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0077] The electronic device 100 is configured to obtain energy data to be archived from the server 200, and obtain modal information corresponding to the energy data.
[0078] The electronic device 100 is configured to classify the energy data according to the modal information to obtain a category of the energy data; wherein, the category is used to characterize the position of the energy data in the energy industry.
[0079] The electronic device 100 is configured to determine an archiving priority of the energy data according to pre-stored archiving requirements and the category, and the archiving priority is used to indicate the importance of the energy data in the energy industry.
[0080] The electronic device 100 is configured to determine a mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modal information; wherein, the mapping vector is used to indicate the correlation between the energy data and the archiving requirements.
[0081] The electronic device 100 is configured to archive the energy data according to the mapping vector.
[0082] Please refer to Figure 6, which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 includes a memory 12 and a processor 13. The memory 12 is used to store computer-readable instructions, and the processor 13 is used to execute the computer-readable instructions stored in the memory to implement a data aggregation method based on energy big data described in any of the above embodiments.
[0083] In an embodiment of the present application, the electronic device 100 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as a data aggregation program based on energy big data.
[0084] Figure 6 Only the electronic device 100 with a memory 12 and a processor 13 is shown. Those skilled in the art can understand that Figure 6 the shown structure does not constitute a limitation on the electronic device 100, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0085] Combined with Figure 2 , the memory 12 in the electronic device 100 stores multiple computer-readable instructions to implement the data aggregation method based on energy big data. The processor 13 can execute the multiple instructions to achieve: obtaining energy data to be archived, and obtaining the modality information corresponding to the energy data; classifying the energy data according to the modality information to obtain the category of the energy data, where the category is used to characterize the position of the energy data in the energy industry; determining the archiving priority of the energy data according to the pre-stored archiving requirements and the category, where the archiving priority is used to indicate the importance of the energy data in the energy industry; determining the mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modality information, where the mapping vector is used to indicate the association relationship between the energy data and the archiving requirements; and archiving the energy data according to the mapping vector.
[0086] Specifically, for the specific implementation method of the above instructions by the processor 13, reference can be made to Figure 2 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0087] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 100 and does not constitute a limitation on the electronic device 100. The electronic device 100 can be a bus-type structure or a star-type structure. The electronic device 100 can also include more or fewer other hardware or software than shown, or different component arrangements. For example, the electronic device 100 can also include input / output devices, network access devices, etc.
[0088] It should be noted that the electronic device 100 is only an example. Other existing or future electronic products that can be adapted to this application should also be included within the protection scope of this application and are hereby incorporated by reference.
[0089] Among them, the memory 12 includes at least one type of readable storage medium. The readable storage medium can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. The memory 12 can be an internal storage unit of the electronic device 100 in some embodiments, such as the mobile hard disk of the electronic device 100. The memory 12 can also be an external storage device of the electronic device 100 in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. equipped on the electronic device 100. The memory 12 can be used not only to store application software installed on the electronic device 100 and various types of data, such as the code of a data aggregation program based on energy big data, etc., but also to temporarily store data that has been output or will be output.
[0090] The processor 13 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 100, connecting various components of the entire electronic device 100 through various interfaces and circuits, and by running or executing programs or modules stored in the memory 12 (such as executing a data aggregation program based on energy big data, etc.), and calling data stored in the memory 12, to execute various functions of the electronic device 100 and process data.
[0091] The processor 13 executes the operating system of the electronic device 100 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned various embodiments of the data aggregation method based on energy big data, for example Figure 2 the steps shown.
[0092] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 100.
[0093] The integrated unit implemented in the form of a software functional module may be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor (Processor) to execute a part of the method for data aggregation based on energy big data described in each embodiment of the present application.
[0094] If the module / unit integrated in the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it may also be completed by a computer program instructing relevant hardware devices. The computer program may be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented.
[0095] Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory, and other memories, etc.
[0096] Furthermore, the computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0097] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, in Figure 6It is represented by only one arrow in the figure, but it does not mean that there is only one bus or one type of bus. The bus is set to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0098] An embodiment of the present application also provides a computer-readable storage medium (not shown in the figure), in which computer-readable instructions are stored, and the computer-readable instructions are executed by a processor in an electronic device to implement the data aggregation method based on energy big data described in any of the above embodiments.
[0099] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0100] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0101] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0102] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not represent any specific order.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A data aggregation method based on energy big data, applied to an electronic device, characterized in that, The method includes: Obtaining the energy data to be archived and obtaining the modality information corresponding to the energy data; Classify the energy data according to the modal information to obtain the category of the energy data; wherein, the category is used to represent the position of the energy data in the energy industry; including: encoding the energy data to obtain a first encoding; encoding the modal information to obtain a second encoding; determining an encoding vector corresponding to the energy data according to the first encoding and the second encoding; inputting the encoding vector into a pre-trained first model to obtain a category vector of the energy data; wherein, training the first model includes: encoding pre-obtained sample energy data and corresponding sample modal information to obtain a sample encoding; determining a confidence level of the sample encoding according to the sample energy data and the sample modal information, the confidence level being used to indicate the matching degree between the semantic information of the energy data and the sample modal information; determining a corresponding predicted category vector based on the sample encoding according to a pre-constructed initial classification model; determining a first loss value of the initial classification model according to the label category vector corresponding to the sample encoding, the predicted category vector, and the confidence level; continuously updating the initial classification model based on the backpropagation algorithm until the first loss value meets a preset condition, and stopping updating the initial classification model to obtain the first model; wherein, determining the first loss value of the initial classification model includes: ; wherein, Loss1 represents the first loss value of the initial classification model; S represents the confidence level; i represents the index of the category; n represents the number of categories; A i represents the probability value corresponding to the i-th category in the label category vector; B i represents the probability value corresponding to the i-th category in the predicted category vector; wherein, determining the confidence level of the sample encoding according to the sample energy data and the sample modal information includes: performing semantic analysis on the energy data to obtain the semantic information of the energy data; determining the matching degree between the semantic information and the sample modal information; determining the confidence level of the sample encoding according to the matching degree; Determining the archiving priority of the energy data according to the pre-stored archiving requirements and the category, where the archiving priority is used to indicate the importance of the energy data in the energy industry; wherein, the archiving requirements are used to indicate the requirement for long-term retention of multi-modal or multi-dimensional data in the data management process; wherein, the archiving requirements include energy industry cost control, energy system performance optimization, energy data security and compliance, and future data access and utilization; Determining a mapping vector between the energy data and the archiving requirements according to the category, the archiving priority, and the modality information; wherein, the mapping vector is used to indicate the association relationship between the energy data and the archiving requirements; Archiving the energy data according to the mapping vector.
2. The data aggregation method based on energy big data according to claim 1, wherein The determining a mapping vector between the energy data and the archiving requirements according to the category, the priority, and the modality information includes: Generating prompt data based on the category, the modality information, and the pre-stored prompt template; Inputting the prompt data into a pre-trained generative model to obtain an initial mapping vector between the energy data and the archiving requirements; Determining a mapping vector between the energy data and the archiving requirements based on the priority and the initial mapping vector.
3. The data aggregation method based on energy big data according to claim 2, wherein, The method further includes training the generative model, and the training the generative model includes: Generating sample prompt data based on the pre-obtained label categories, sample modality information, and the pre-stored prompt template; Inputting the sample prompt data into a pre-constructed initial generative model to obtain a predicted mapping vector between the energy data and the archiving requirements; Determining a second loss value of the initial generative model according to the label mapping vector corresponding to the sample modality information and the predicted mapping vector; Updating the initial generative model based on the backpropagation algorithm until the second loss value meets a preset condition, and stopping updating the initial generative model to obtain the generative model.
4. The data aggregation method based on energy big data according to claim 3, characterized in that, The determining a second loss value of the initial generative model according to the label mapping vector corresponding to the sample modality information and the predicted mapping vector includes: ; Among them, Loss2 represents the second loss value of the initial generative model; j represents the index of the dimension in the label mapping vector; m represents the number of dimensions in the label mapping vector; C j represents the mapping value corresponding to the j-th dimension in the label mapping vector; D j represents the mapping value corresponding to the j-th category in the predicted mapping vector.
5. A data aggregation system based on energy big data, characterized in that, The system is used to implement the method according to any one of claims 1 to 4, and the system includes an electronic device and a server; The electronic device is used to obtain the energy data to be archived from the server and obtain the modality information corresponding to the energy data; The electronic device is used to classify the energy data according to the modality information to obtain the category of the energy data; wherein, the category is used to characterize the position of the energy data in the energy industry; The electronic device is used to determine the archiving priority of the energy data according to the pre-stored archiving requirements and the category, where the archiving priority is used to indicate the importance of the energy data in the energy industry; The electronic device is configured to determine a mapping vector between the energy data and the archiving requirement according to the category, the archiving priority, and the modality information; wherein the mapping vector is used to indicate the association relationship between the energy data and the archiving requirement. The electronic device is configured to archive the energy data according to the mapping vector.
6. The data aggregation system based on energy big data according to claim 5, wherein, The electronic device is further configured to encode the energy data to obtain a first encoding. The electronic device is configured to encode the modality information to obtain a second encoding. The electronic device is configured to determine an encoding vector corresponding to the energy data according to the first encoding and the second encoding. The electronic device is configured to input the encoding vector into a pre-trained first model to obtain a category vector of the energy data; wherein the category vector is used to indicate the category of the energy data.
Citation Information
Patent Citations
Data center data backup disaster recovery intelligent management and control platform and method
CN118245285A
Energy power data analysis method and system and medium
CN118469126A