An automated metadata extraction method for the energy sector

Generate prompt words through the first language model and extract the second metadata of energy data using the second language model, the problem of inefficient metadata extraction in the energy field is solved, and efficient and accurate metadata extraction is achieved to adapt to the needs of different energy types.

CN119719176BActive Publication Date: 2025-08-12BIG DATA CENT OF STATE GRID CORP OF CHINA +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411879783.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-08-12
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Metadata extraction is inefficient and low adaptability in the energy field, and existing automation tools are difficult to meet the metadata extraction needs in specific fields.

Method used

By using the first language model to generate prompt words corresponding to the energy data type and input them into the second language model, extracting refined and professional second metadata to ensure that the data structure corresponds to the energy type.

Benefits of technology

It improves the efficiency and accuracy of metadata extraction, meets the metadata needs of different energy types, improves adaptability and flexibility, and avoids the inefficiency and high cost problems of manual extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719176B_ABST
    Figure CN119719176B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and storage medium for automatic metadata extraction for the energy field, belonging to the field of energy data management, including: determining first metadata and parsed data of energy data; inputting the first metadata into a first language model to obtain multiple prompt words of the energy data; inputting the multiple prompt words and parsed data of the energy data into a second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data. While adapting to the diversity and complexity of energy data, the present application meets the metadata extraction needs of different energy types and improves the adaptability and flexibility of metadata extraction. It solves the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of energy data management, and specifically to an automatic metadata extraction method, device, equipment and storage medium for the energy field. Background Art

[0002] In the process of building a digital energy network in the energy sector, automated metadata extraction is key to achieving efficient data management and utilization. With the coexistence of various data sources and standards, the distributed, heterogeneous, and cross-subject complexity of energy data is particularly prominent. For example, the metadata content varies significantly across different energy sectors (such as solar energy and hydrogen energy).

[0003] At present, metadata extraction is mainly divided into two categories: one is to obtain metadata manually, and the other is to automatically generate metadata with the help of technical means, usually relying on metadata extraction tools.

[0004] However, manual extraction methods have obvious inefficiencies, and metadata generated with the help of existing automated tools can only extract limited types of metadata information, making it difficult to meet users' needs for metadata extraction in specific fields. Summary of the Invention

[0005] This application aims to provide a method, device, equipment and storage medium for automatic metadata extraction in the energy field, at least to solve the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata.

[0006] In a first aspect, the present application discloses an automatic metadata extraction method for the energy field, comprising:

[0007] Determining first metadata of energy data and determining parsed data of the energy data; the first metadata having a preset data structure;

[0008] Inputting the first metadata into a first language model to obtain a plurality of prompt words for the energy data; the generated plurality of prompt words correspond to energy types corresponding to the energy data;

[0009] The plurality of prompt words and parsed data of the energy data are input into a second language model to obtain second metadata of the energy data, so that a data structure of the second metadata corresponds to an energy type corresponding to the energy data.

[0010] In a second aspect, the present application also discloses an automatic metadata extraction device for the energy field, comprising:

[0011] a data extraction module, configured to determine first metadata of energy data and to determine parsed data of the energy data; the first metadata having a preset data structure;

[0012] a prompt word generation module, configured to input the first metadata into a first language model to obtain a plurality of prompt words for the energy data; the generated plurality of prompt words corresponding to energy types corresponding to the energy data;

[0013] The metadata generation module is configured to input the plurality of prompt words and parsed data of the energy data into a second language model to obtain second metadata of the energy data, so that a data structure of the second metadata corresponds to an energy type corresponding to the energy data.

[0014] In a third aspect, an embodiment of the present application further discloses an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] In summary, in the embodiment of the present application, the prompt words generated by the first language model can accurately reflect the specific types and characteristics of energy data. Through these prompt words, the context and domain characteristics of energy data can be accurately identified and understood, and the adaptability of metadata extraction is improved; then, under the guidance of the prompt words, the second language model extracts more detailed and professional metadata information from the parsed data, so that the extracted second metadata not only has high accuracy, but also can cover the professional metadata requirements of various sub-scenarios in the energy field. Therefore, the method based on the embodiment of the present application avoids the inefficiency and high cost problems of manual extraction and greatly improves the efficiency of metadata extraction. At the same time, the use of the language model ensures that the extracted metadata has high accuracy and consistency. While adapting to the diversity and complexity of energy data, it meets the metadata extraction requirements of different energy types and improves the adaptability and flexibility of metadata extraction. It solves the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In the attached figure:

[0018] Figure 1 This is a flowchart of the steps of an automatic metadata extraction method for the energy field provided by an embodiment of the present application;

[0019] Figure 2 This is a flowchart of another method for automatically extracting metadata for the energy field provided by an embodiment of the present application;

[0020] Figure 3 This is the initial construction process of the second language model in the embodiment of the present application;

[0021] Figure 4 It is a program logic diagram provided according to an embodiment of the present application;

[0022] Figure 5 This is a block diagram of an automatic metadata extraction device for the energy field provided by an embodiment of the present application;

[0023] Figure 6 is a block diagram of an electronic device according to an embodiment of the present application;

[0024] Figure 7 This is a block diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0027] Figure 1 This is an automatic metadata extraction method for the energy field provided in an embodiment of the present application.

[0028] The method may include the following steps:

[0029] Step 101: Determine first metadata of energy data and determine parsed data of the energy data.

[0030] The first metadata has a preset data structure.

[0031] In some embodiments of the present application, to ensure that energy data is structured and standardized during subsequent processing, first metadata for the energy data is determined, along with parsed data for the energy data. The first metadata has a predefined data structure. This predefined data structure is typically pre-defined metadata that may not be tailored to specific needs. This provides effective data support for model input in subsequent steps.

[0032] In a specific example, a user might need to process data from the coal industry. First, the user collects and organizes this data, including the raw data and associated metadata. This process provides a reliable foundation for subsequent metadata extraction.

[0033] Step 102: Input the first metadata into a first language model to obtain a plurality of prompt words of energy data.

[0034] The generated multiple prompt words correspond to the energy types corresponding to the energy data.

[0035] In some embodiments of the present application, in order to utilize the language understanding capabilities of the first language model to generate prompt words that match the specific type of energy data, the first metadata is input into the first language model, and the model is processed to generate multiple prompt words. These prompt words match the energy type corresponding to the energy data, such as role definitions and task descriptions for energy data, as well as user-described metadata field requirements, or input and output samples. Prompt words are important information that guides the model in metadata extraction, helping the model to better understand and parse the content of energy data. This lays the foundation for accurate metadata extraction by the second language model.

[0036] In a specific example, we'll process data from the solar industry. This data is converted into structured primary metadata and then fed into a primary language model. The model processes the input metadata and generates multiple prompts, such as "solar radiation intensity," "photovoltaic cell efficiency," and "installation angle." These prompts correspond to the specific types and characteristics of solar data. Through this process, users can obtain a set of prompts that accurately reflect the characteristics of solar data, providing a reliable foundation for subsequent metadata extraction.

[0037] Step 103 : Input the plurality of prompt words and the parsed data of the energy data into a second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data.

[0038] In some embodiments of this application, to leverage the powerful processing capabilities of a second language model to extract more detailed and specialized metadata from energy data, multiple generated prompt words and parsed data are fed into the second language model. The second language model then parses and extracts secondary metadata for the energy data based on this input information. These prompt words guide the model in understanding the context and domain characteristics of the energy data. This enables the acquisition of more detailed and specialized secondary metadata corresponding to the energy data, thereby improving the accuracy and adaptability of the data.

[0039] In a specific example, a user processes data from the hydrogen energy industry. The user inputs structured primary metadata and generated prompt words into a second language model. The model uses the prompt words to guide data parsing and extract secondary metadata from the hydrogen energy data, such as "hydrogen purity," "storage pressure," and "transportation temperature." This extracted secondary metadata corresponds to the specific types and characteristics of hydrogen energy, ensuring data accuracy and professionalism. Through this process, users can obtain high-quality metadata information for further data management and utilization.

[0040] In summary, in the embodiment of the present application, the prompt words generated by the first language model can accurately reflect the specific types and characteristics of energy data. Through these prompt words, the context and domain characteristics of energy data can be accurately identified and understood, and the adaptability of metadata extraction is improved; then, under the guidance of the prompt words, the second language model extracts more detailed and professional metadata information from the parsed data, so that the extracted second metadata not only has high accuracy, but also can cover the professional metadata requirements of various sub-scenarios in the energy field. Therefore, the method based on the embodiment of the present application avoids the inefficiency and high cost problems of manual extraction and greatly improves the efficiency of metadata extraction. At the same time, the use of the language model ensures that the extracted metadata has high accuracy and consistency. While adapting to the diversity and complexity of energy data, it meets the metadata extraction requirements of different energy types and improves the adaptability and flexibility of metadata extraction. It solves the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata.

[0041] Figure 2 This is another method for automatically extracting metadata for the energy field provided in an embodiment of the present application.

[0042] The method may include the following steps:

[0043] Step 201: Determine first metadata of energy data and determine parsed data of the energy data.

[0044] The first metadata has a preset data structure.

[0045] The method shown in this step has been described in step 101 and will not be repeated here.

[0046] Optionally, in order to determine the first metadata of the energy data, step 201 includes the following sub-steps:

[0047] Sub-step 2011: extracting a first description text of the energy data from the energy data.

[0048] In some embodiments of this application, to ensure that subsequent metadata extraction is based on detailed and accurate descriptive text, first descriptive text related to the energy data is extracted. This descriptive text contains key information and features of the energy data, facilitating subsequent model processing. The descriptive text generated in this step provides fundamental information support for subsequent steps, ensuring the comprehensiveness and accuracy of metadata extraction.

[0049] In a specific example, a user needs to process a batch of data from the coal industry. The user extracts relevant descriptive text from this data, such as coal type name, calorific value, and ash content. This descriptive text contains the main characteristics and important information of the coal data. Through this process, the user obtains a set of detailed descriptive text, providing the necessary foundational data for subsequent metadata extraction.

[0050] Sub-step 2012: When the amount of information in the first description text is less than or equal to a preset information threshold, information enhancement is performed on the first description text to update the first description text.

[0051] In some embodiments of the present application, in order to ensure that the description text contains sufficient information so that it can provide sufficient data support for metadata extraction, the description text will be enhanced when it is detected that the information content of the first description text is less than or equal to a preset information threshold. Information enhancement includes adding or supplementing more relevant data information to make the description text more detailed and comprehensive. The information threshold refers to the standard used to measure whether the description text contains sufficient information. The clear technical effect brought about by executing this step is that the updated first description text contains more information, thereby improving the accuracy and completeness of metadata extraction in subsequent steps.

[0052] In one specific example, a user discovered that the description text extracted from wind energy data was insufficiently informative, perhaps containing only two data points: "wind speed" and "wind direction." The user enhanced this description text, adding more relevant information such as "wind force level" and "wind energy conversion efficiency." This process resulted in an updated description text with more information, providing a more comprehensive data foundation for subsequent metadata extraction. Ultimately, the updated description text ensured the accuracy and completeness of the metadata extraction process.

[0053] Sub-step 2013: When the information amount of the first description text is greater than the information amount threshold, keyword extraction is performed on the description text to obtain multiple keywords in the description text.

[0054] In some embodiments of the present application, in order to effectively extract key information from a large amount of information and facilitate subsequent data processing and analysis, keyword extraction is performed on the first description text when it is detected that the amount of information in the first description text exceeds a preset information threshold. Keyword extraction refers to identifying and extracting key words from the text that best represent the text content. The information threshold is a criterion for measuring whether the description text contains too much information. The keywords generated in this way can effectively summarize the main content of the description text, facilitating subsequent metadata packaging.

[0055] In one specific example, a user processed a batch of solar data description text and discovered that it contained a large amount of information, such as "solar radiation intensity," "photovoltaic panel efficiency," "installation angle," and "weather conditions." Because the amount of information exceeded a preset threshold, the user performed keyword extraction on the description text, identifying and extracting keywords such as "radiation intensity," "panel efficiency," and "installation angle." This process enabled the user to obtain a set of keywords that effectively summarized the main content of the description text, providing a concise and accurate information foundation for subsequent metadata packaging.

[0056] Sub-step 2014: encapsulate the obtained multiple keywords according to the data structure to obtain first metadata.

[0057] In some embodiments of this application, to structure the extracted key information and facilitate subsequent data processing and analysis, the previously obtained multiple keywords are encapsulated according to a predefined data structure to obtain first metadata. The metadata data structure, also known as the metadata schema, refers to a method of organizing and storing data to ensure that the data can be processed and used consistently and accurately. This generates first metadata that meets preset standards, thereby improving data consistency and integrity.

[0058] In a specific example, a user is processing a batch of wind energy data. They have previously extracted multiple keywords from the descriptive text, such as "wind speed," "wind direction," and "power generation." The user then encapsulates these keywords in a defined data structure, such as a table or JSON format, to generate primary metadata. This process yields a structured set of primary metadata, providing a reliable data foundation for subsequent metadata extraction and analysis.

[0059] Optionally, in order to determine the parsed data of the energy data, step 201 includes the following sub-steps:

[0060] In sub-step 2015 , when the energy data is structured data, the text data extracted from the energy data is determined as parsed data.

[0061] In some embodiments of this application, to directly utilize text data within structured data as parsed data and simplify the processing flow, if the energy data is confirmed to be structured data, the extracted text data will be determined as parsed data. Structured data refers to data organized and stored according to predefined models and formats. This allows existing text data to be directly utilized as parsed data, reducing additional data processing and improving processing efficiency and accuracy.

[0062] In a specific example, a user needs to process a set of structured data from the power industry. This data includes predefined fields such as power plant name, installed capacity, and power generation. From this structured data, the user extracts the corresponding text data, such as "Power Plant Name: XX Power Plant," "Installed Capacity: 500MW," and "Power Generation: 300GWh," and directly identifies this as parsed data. This process simplifies data processing and quickly obtains the required parsed data.

[0063] Sub-step 2016 , when the energy data includes unstructured data, performing data parsing on the unstructured data to obtain a second description text of the energy data, and encapsulating the second description text with the unstructured data in the energy data as parsed data.

[0064] In some embodiments of the present application, in order to ensure that unstructured data can be effectively parsed and utilized, thereby improving the overall integrity and availability of the data, when the energy data contains unstructured data, data parsing of the unstructured data is performed to obtain a second description text. Unstructured data refers to free-format data that is not organized according to a predefined model. Data parsing is the process of converting these free-format data into more structured and understandable text. Subsequently, the obtained second description text is encapsulated with the original unstructured data as parsed data. The clear technical effect brought about by executing this step is that the consistency and integrity of the data are enhanced, ensuring the effective utilization of unstructured data in metadata extraction.

[0065] In one specific example, a user processed a batch of unstructured data from the oil industry, including free-form daily operational records and maintenance reports. The user parsed this unstructured data, extracting key descriptive text such as "maintenance time," "equipment status," and "operator." This extracted descriptive text was packaged with the original unstructured data to generate parsed data. This process enabled the user to convert unstructured data into structured information, facilitating subsequent metadata extraction and analysis.

[0066] Step 202: Input the first metadata into a first language model to obtain a plurality of prompt words of energy data.

[0067] The generated multiple prompt words correspond to the energy types corresponding to the energy data.

[0068] The method shown in this step has been described in step 102 and will not be repeated here.

[0069] Step 203 : Input the plurality of prompt words and the parsed data of the energy data into a second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data.

[0070] The method shown in this step has been described in step 103 and will not be repeated here.

[0071] Step 204 : In response to the adjustment operation on the second metadata, the first language model and / or the second language model are adjusted to update the second metadata according to the adjusted first language model and / or the second language model.

[0072] In some embodiments of the present application, to ensure the accuracy and adaptability of the second metadata, dynamic adjustments can be made based on actual needs. The first language model and / or the second language model are adjusted in response to user adjustments to the second metadata. These adjustments can be optimized based on user feedback and needs to ensure that the metadata generated by the model better meets the user's actual needs. By adjusting model parameters and training data, the accuracy and flexibility of metadata extraction can be further improved. The clear technical effect of executing this step is that the updated second metadata can better reflect the specific characteristics of energy data and user needs, thereby improving the applicability and accuracy of the data.

[0073] In a specific example, a user may find that the extracted metadata for solar data is inaccurate or not as expected, for example, the "photovoltaic panel type" information is insufficiently detailed. The user adjusts the secondary metadata and provides feedback to the system. Based on this feedback, the system optimizes and adjusts the parameters and training data of the first and / or second language models. This adjustment process results in more detailed and accurate secondary metadata, such as the addition of detailed classification information for "photovoltaic panel type." Ultimately, the updated metadata better meets user needs and enhances the data's practicality and credibility.

[0074] Optionally, step 204 includes the following sub-steps:

[0075] Sub-step 2041 : obtaining adjusted data of the second metadata in response to the adjustment operation on the second metadata.

[0076] In some embodiments of the present application, in order to obtain user feedback on modifications to the second metadata and further optimize the model, relevant adjustment data is generated in response to user adjustments to the second metadata. Adjustment data refers to specific information reflecting user modifications or optimizations to the metadata. This generated adjustment data provides foundational data for subsequent model adjustments, ensuring that the model better meets user needs.

[0077] In one specific example, a user discovered that some metadata in the extracted natural gas industry data was insufficiently detailed, such as missing information on "gas purity." The user adjusted this metadata, adding detailed "gas purity" information. In response, the system generated adjusted data containing the newly added "gas purity" information. This process ensured that all critical information was included in the metadata, providing accurate feedback for subsequent model optimization. Ultimately, the generated adjusted data ensured that the model better adapted to actual needs and improved the accuracy and completeness of the metadata.

[0078] Sub-step 2042: adjusting the first language model and / or the second language model according to the adjustment data of the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model.

[0079] In some embodiments of this application, to ensure adaptive model optimization and improvement, thereby more accurately extracting metadata information, at least one of the first and second language models is adjusted based on the obtained second metadata adjustment data. This adjustment process may involve retraining the model or modifying model parameters to reflect user feedback and needs. Through this adjustment, the model can better adapt to and interpret energy data. This updated model generates more accurate and adaptable second metadata, further improving metadata extraction accuracy and user satisfaction.

[0080] In a specific example, a user adjusted the extracted metadata for the natural gas industry, adding new data fields such as "gas density" and "transmission pressure." Based on this adjusted data, the system adjusted at least one of the first and second language models, for example by retraining the model or adjusting parameter settings. Through this adjustment process, the metadata regenerated by the model more accurately reflects the user's needs, including the new data fields.

[0081] Optionally, for the adjustment process of the first language model, sub-step 2042 includes the following sub-steps:

[0082] Sub-step 20421: Determine the adjustment data of the first metadata according to the adjustment data of the second metadata.

[0083] In some embodiments of the present application, to ensure that the optimization and training of the first language model are based on the latest user needs and data adjustments, the corresponding adjustment data of the first metadata is determined based on the adjustment data of the second metadata. Adjustment data refers to the specific information after the user modifies or optimizes the metadata. In this way, the first language model can more accurately reflect user needs during subsequent training. The adjustment data of the first metadata generated in this way will provide a foundation for subsequent model training and optimization.

[0084] In a specific example, a user adjusted the secondary metadata extracted from the natural gas industry, adding new data fields such as "gas purity" and "pipeline pressure." Based on these adjustments, the system identified the relevant primary metadata adjustments, including updated field descriptions and relationships. This ensures that the training data for the primary language model incorporates the latest requirements and adjustments.

[0085] In sub-step 20422 , the adjusted data of the first metadata is used as first training data to train the first language model to obtain a trained first language model, and the process returns to step 202 .

[0086] In some embodiments of the present application, to optimize the first language model through continuous training so that it can better adapt to the user's adjusted needs, the determined first metadata adjustment data is used as the first training data to retrain the first language model. The process of step 202 is then repeated, where the first metadata is input into the first language model to obtain multiple prompt words for energy data. Training refers to further learning and optimizing the model using new data, aiming to improve its performance and accuracy. In this way, the first language model is updated and better understands and handles metadata extraction tasks. This results in a trained first language model that can more accurately generate prompt words, thereby improving the effectiveness of metadata extraction in subsequent steps.

[0087] In a specific example, a user modified data for the natural gas industry, adding new metadata fields such as "gas composition" and "transmission conditions." The system used this modified metadata as training data to retrain the first language model. Through this process, users were able to enable the first language model to better learn and understand the new data structure and field requirements. Ultimately, the trained first language model was able to generate more accurate and adaptable prompts, helping to improve the performance of the second language model in metadata extraction, further enhancing the accuracy and comprehensiveness of metadata extraction.

[0088] Optionally, for the adjustment process of the second language model, sub-step 2042 includes the following sub-steps:

[0089] Sub-step 20423: updating the plurality of prompt words according to the adjustment data of the second metadata.

[0090] In some embodiments of this application, to ensure prompts accurately reflect the latest user needs and data adjustments, multiple prompts are updated based on the adjusted data in the second metadata. Updating prompts means regenerating or modifying them based on the adjusted data to better adapt them to the new data content and extraction requirements. Prompts are important information that guides the model during metadata extraction, helping the model understand and parse data. These updated prompts provide more accurate information support for subsequent model training and metadata extraction.

[0091] In a specific example, a user adjusted data for the natural gas industry, adding new metadata fields such as "gas composition" and "transmission conditions." Based on these adjustments, the system updated the corresponding prompts, generating new prompts such as "composition analysis," "transmission pressure," and "storage conditions." This process ensures that prompts reflect the latest requirements and data content, providing an accurate information foundation for subsequent model training.

[0092] In sub-step 20424 , the updated prompt word is used as the second training data to train the second language model to obtain a trained second language model, and then the process returns to step 203 .

[0093] In some embodiments of the present application, to continuously optimize the second language model so that it can better process and parse the latest data requirements, the updated prompt words are used as second training data to retrain the second language model. The process of step 203 is then repeated, whereby the multiple prompt words and parsed data of the energy data are input into the second language model to obtain the second metadata of the energy data. The training process involves further learning and optimizing the model using the updated prompt word data to ensure that the model more accurately reflects the latest data content and extraction requirements. Through this continuous training, the second language model can maintain efficient and accurate metadata extraction capabilities. The more accurate and efficient second language model thus obtained improves the overall effectiveness of metadata extraction.

[0094] In a specific example, a user adjusted data for the natural gas industry, updating prompt terms such as "gas composition analysis" and "pressure conditions." The system used these updated prompt terms as new training data to retrain the second language model. Through this process, the user ensured that the model could learn and adapt to the new prompt terms, improving its understanding and parsing capabilities of the data. Ultimately, the trained second language model was able to more accurately and efficiently extract specialized metadata for the natural gas industry, providing a more reliable information foundation for subsequent data applications and management.

[0095] Step 205: Store the second metadata together with the energy data.

[0096] In some embodiments of this application, to ensure that all data can be centrally managed and effectively utilized, the extracted secondary metadata is stored in the system alongside the original energy data. This storage method ensures data integrity and traceability. By storing both types of data together, users can easily conduct comprehensive queries and analyses. The clear technical effect of executing this step is enhanced data consistency and integrity, and improved data management efficiency.

[0097] In one specific example, users have extracted detailed metadata from hydrogen industry data. They store this metadata alongside the corresponding raw hydrogen data in a data warehouse. This process ensures that all hydrogen-related data is centralized in a single storage system, facilitating future data query, analysis, and application. Ultimately, this shared storage approach not only ensures data integrity and consistency but also improves the efficiency of data management and utilization.

[0098] Optionally, step 205 includes the following sub-steps:

[0099] Sub-step 2051 : generating identification data of the energy data according to the second metadata, so that the identification data corresponds one-to-one with the energy data.

[0100] In some embodiments of the present application, to ensure that each energy data item has a unique identifier for subsequent data management and tracing, corresponding identification data is generated based on the second metadata to ensure a one-to-one correspondence between the identification data and the energy data item. Identification data refers to specific information that uniquely identifies each energy data item. The clear technical effect of executing this step is the generation of identification data that uniquely identifies each energy data item, improving data management and retrieval efficiency.

[0101] In a specific example, a user processed a batch of power industry data and extracted secondary metadata, such as power generation and installed capacity. The system then generated unique identifiers based on this metadata, using, for example, a UUID (Universally Unique Identifier) or other identification methods. This unique identifier ensures that each power data record has a unique identity. This process enables users to easily manage and trace each data record.

[0102] Sub-step 2052: storing the second metadata in a metadata database, storing the energy data in an energy database, and establishing a mapping relationship between the metadata database and the energy database based on the identification data.

[0103] In some embodiments of the present application, to achieve separate storage and management of metadata and energy data, while simultaneously establishing a link between the two through identification data for efficient data retrieval and management, the extracted second metadata is stored in a metadata database, the corresponding energy data is stored in an energy database, and a mapping relationship is established between the two databases based on the generated identification data. The mapping relationship refers to the use of identification data to connect the metadata in the metadata database with the actual data in the energy database. This enables efficient management of metadata and energy data, enhancing data traceability and query efficiency.

[0104] In a specific example, a user processes data from a batch of photovoltaic power plants. Extracted metadata, such as "power plant name" and "installed capacity," is stored in a metadata database, while actual energy data, such as "real-time power generation" and "equipment status," is stored in an energy database. The system uses unique identifiers, such as UUIDs, to establish a mapping relationship between the metadata in the metadata database and the energy data in the energy database. This process enables users to conveniently query and manage photovoltaic power plant data. Ultimately, the established mapping ensures efficient data retrieval and management, while also improving data integrity and traceability.

[0105] Optionally, in some embodiments of the present application, a fine-tuned language model is used as the second language model.

[0106] Fine-tuned language models can acquire more accurate language understanding and information extraction capabilities within a specific domain through further training on data from that domain. This enables the model to more accurately identify and parse relevant metadata information when processing energy data. At the same time, the fine-tuning process enables the model to adapt to the specific needs and characteristics of the energy sector, thereby improving the adaptability and accuracy of metadata extraction. This targeted optimization can maintain high extraction efficiency and accuracy even when energy data is complex and diverse. In addition, the use of fine-tuned language models can also reduce the difficulty of processing unstructured data while ensuring data processing performance. By fine-tuning the model, it can better handle terminology and data structures unique to the energy sector, thereby improving the overall performance of data parsing and processing. Moreover, because the fine-tuned language model has been optimized on datasets in related fields, it performs better than a general language model that has not been fine-tuned when performing metadata extraction tasks. This advantage ensures that the model can provide high-quality metadata extraction results across different energy types and complex data environments. Therefore, choosing to use a fine-tuned language model as the second language model not only improves the accuracy and efficiency of metadata extraction, but also enhances the model's adaptability to the specific needs of the energy field, thereby ensuring the efficiency, accuracy and reliability of the metadata extraction process.

[0107] like Figure 3 As shown in FIG, a specific process of obtaining an initial fine-tuning model when the fine-tuning model is used as the second language model in an embodiment of the present application:

[0108] R1: Data Acquisition. The data acquisition process involves collecting relevant data from various energy databases or data sources. Data acquisition is performed to gather foundational data for model training. This step provides a large amount of initial data for fine-tuning the model, providing foundational data support for subsequent processing and training.

[0109] R2: Data Processing. This process involves removing noise, handling missing values, and standardizing the data format. Data processing is performed to clean and organize the collected data to ensure data quality and consistency. This processing step yields clean and structured data, ensuring that the data quality meets the requirements for model training.

[0110] R3: Data Annotation. The data annotation process involves manually or automatically annotating the data, noting its key features and attributes. Data annotation is performed to label the data so that the model can learn how to extract meaningful metadata from the data. This labeled data can serve as effective training samples for the model, improving its learning and accuracy.

[0111] R4: Data Augmentation. The process of data augmentation involves performing various transformations on existing data, such as data expansion, noise addition, and data augmentation techniques. Data augmentation is performed to expand the diversity of the training dataset and improve the model's generalization capabilities. Data augmentation can enrich the variety of training data and improve the model's performance in different scenarios.

[0112] R5: Model Fine-Tuning. Model fine-tuning involves retraining the pre-trained model using processed and annotated data. Model fine-tuning is performed to further train the base model on domain-specific data to improve its performance within that domain. Fine-tuning allows the model to better adapt to the data characteristics of a specific domain, improving its professionalism and accuracy.

[0113] R6: Model Generation. The model generation process involves saving the fine-tuned model and deploying it to the system. Model generation is performed to complete the construction of the fine-tuned model so that it can be applied to actual metadata extraction tasks. The resulting fine-tuned model can serve as a second language model for metadata extraction from energy data.

[0114] like Figure 4 As shown, it is the program execution logic under the method disclosed in the embodiment of the present application, which specifically completes the entire process of storing energy data based on metadata by executing the following steps:

[0115] S1 User uploads data: Users begin uploading energy data that needs to be packaged into digital objects. This data can include various types of energy data such as coal, electricity, oil, and gas.

[0116] S2 Data Analysis:

[0117] After the data is uploaded, the system parses the user data and converts it into a standard format that meets the model input requirements. This step ensures that the data can be processed and understood correctly.

[0118] in:

[0119] S2.1 Metadata format upload: Users upload the metadata format (schema), describe the specific requirements for metadata extraction, and specify the metadata fields that need to be extracted.

[0120] S2.2 Calling the First Language Model to Parse Metadata Format: The system calls the first language model to parse the metadata format uploaded by the user. The parsed format generates standard prompt words, including role definition, task description, metadata field requirements for user description, input and output samples, etc.

[0121] S3 data parsing: The system performs detailed parsing of the uploaded data and converts the data into a structured format for subsequent processing.

[0122] S4 Prompt word generation: Based on the parsed data and metadata format, a set of prompt words is generated. These prompt words are used to guide the second language model to extract metadata.

[0123] S5 calls the second language model to generate metadata: The system passes the prompt word and parsed data to the second language model. The large model parses the data based on the prompt word and generates structured metadata that meets user needs and industry standards.

[0124] S6 User Review: Users review the generated metadata to ensure its accuracy and completeness. If any issues are found, users can provide suggestions for adjustments.

[0125] S7 User Adjustment: Users adjust metadata based on actual needs to ensure it meets expectations. Adjustment information is fed back to the system for model optimization.

[0126] S8 Digital Object Packaging: Finally, the adjusted and reviewed metadata enters the digital object packaging program to ensure the uniqueness and traceability of the data.

[0127] in:

[0128] S8.1 The identity resolution system generates a unique identifier: The system calls the identity resolution system to generate a unique identifier for each digital object to ensure the uniqueness of the digital object.

[0129] S8.2 Identification-associated metadata is stored in the registry: After the generated metadata is associated with the unique identifier, it is stored in the system's built-in metadata registry in a standardized format, supporting efficient indexing and query functions.

[0130] S8.3 Identify and associate entities and store them in the warehouse: The system associates the corresponding data entity with a unique identifier and stores it in the system's built-in data warehouse to achieve complete encapsulation of energy digital objects and support long-term storage and management.

[0131] In summary, in the embodiment of the present application, the prompt words generated by the first language model can accurately reflect the specific types and characteristics of energy data. Through these prompt words, the context and domain characteristics of energy data can be accurately identified and understood, and the adaptability of metadata extraction is improved; then, under the guidance of the prompt words, the second language model extracts more detailed and professional metadata information from the parsed data, so that the extracted second metadata not only has high accuracy, but also can cover the professional metadata requirements of various sub-scenarios in the energy field. Therefore, the method based on the embodiment of the present application avoids the inefficiency and high cost problems of manual extraction and greatly improves the efficiency of metadata extraction. At the same time, the use of the language model ensures that the extracted metadata has high accuracy and consistency. While adapting to the diversity and complexity of energy data, it meets the metadata extraction requirements of different energy types and improves the adaptability and flexibility of metadata extraction. It solves the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata.

[0132] refer to Figure 5 , which shows an automatic metadata extraction device 30 for the energy field provided by an embodiment of the present application, including:

[0133] The data extraction module 301 is used to determine first metadata of the energy data and determine parsed data of the energy data; the first metadata has a preset data structure;

[0134] A prompt word generation module 302 is configured to input the first metadata into a first language model to obtain a plurality of prompt words for energy data; the generated plurality of prompt words correspond to energy types corresponding to the energy data;

[0135] The metadata generation module 303 is configured to input multiple prompt words and parsed data of the energy data into a second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data.

[0136] Optionally, the data extraction module 301 includes:

[0137] A description extraction submodule, configured to extract a first description text of the energy data from the energy data;

[0138] An information enhancement submodule, configured to enhance the information of the first description text to update the first description text when the amount of information in the first description text is less than or equal to a preset information amount threshold;

[0139] A keyword submodule is used to extract keywords from the description text to obtain multiple keywords in the description text when the information amount of the first description text is greater than the information amount threshold;

[0140] The first metadata submodule is configured to encapsulate the obtained multiple keywords according to a data structure to obtain first metadata.

[0141] Optionally, the data extraction module 301 includes:

[0142] a first parsing submodule, configured to, when the energy data is structured data, determine text data extracted from the energy data as parsed data;

[0143] The second parsing submodule is used to parse the unstructured data to obtain a second description text of the energy data when the energy data contains unstructured data, and to encapsulate the second description text with the unstructured data in the energy data as parsed data.

[0144] Optionally, the device 30 further includes:

[0145] The model training module is configured to adjust the first language model and / or the second language model in response to the adjustment operation on the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model.

[0146] Optional model training modules include:

[0147] a second metadata adjustment submodule, configured to obtain adjustment data of the second metadata in response to an adjustment operation on the second metadata;

[0148] The model training submodule is configured to adjust the first language model and / or the second language model according to the adjustment data of the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model.

[0149] Optional model training submodules include:

[0150] a first adjustment unit, configured to determine the adjustment data of the first metadata according to the adjustment data of the second metadata;

[0151] The first training unit is configured to train the first language model using the adjusted data of the first metadata as first training data to obtain a trained first language model, and return to the step of inputting the first metadata into the first language model to obtain a plurality of prompt words for energy data.

[0152] Optionally, the model training submodule includes:

[0153] a second adjusting unit, configured to update the plurality of prompt words according to the adjustment data of the second metadata;

[0154] The second training unit is configured to train the second language model using the updated prompt words as second training data to obtain a trained second language model, and return to the step of inputting multiple prompt words and parsed data of the energy data into the second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data.

[0155] Optionally, the device 30 further includes:

[0156] The data storage module is used to store the second metadata together with the energy data.

[0157] Optionally, the data storage module includes:

[0158] an identification generating submodule, configured to generate identification data of the energy data according to the second metadata, so that the identification data corresponds one-to-one with the energy data;

[0159] The data storage submodule is configured to store the second metadata in a metadata database, store the energy data in an energy database, and establish a mapping relationship between the metadata database and the energy database according to the identification data.

[0160] In summary, in the embodiment of the present application, the prompt words generated by the first language model can accurately reflect the specific types and characteristics of energy data. Through these prompt words, the context and domain characteristics of energy data can be accurately identified and understood, and the adaptability of metadata extraction is improved; then, under the guidance of the prompt words, the second language model extracts more detailed and professional metadata information from the parsed data, so that the extracted second metadata not only has high accuracy, but also can cover the professional metadata requirements of various sub-scenarios in the energy field. Therefore, the method based on the embodiment of the present application avoids the inefficiency and high cost problems of manual extraction and greatly improves the efficiency of metadata extraction. At the same time, the use of the language model ensures that the extracted metadata has high accuracy and consistency. While adapting to the diversity and complexity of energy data, it meets the metadata extraction requirements of different energy types and improves the adaptability and flexibility of metadata extraction. It solves the problems of low efficiency of metadata extraction in the energy field and low adaptability of the extracted metadata.

[0161] Reference Figure 6 , electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output (I / O) interface 512 , a sensor component 514 , and a communication component 516 .

[0162] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 502 may include one or more modules to facilitate interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate interaction between the multimedia component 508 and the processing component 502.

[0163] The memory 504 is used to store various types of data to support operations on the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0164] The power supply assembly 506 provides power to the various components of the electronic device 500. The power supply assembly 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 500.

[0165] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the demarcation of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the electronic device 500 is in an operating mode, such as a capture mode or a multimedia mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.

[0166] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that receives external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, or a voice recognition mode. The received audio signals may be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 also includes a speaker for outputting audio signals.

[0167] The input / output I / O interface 512 provides an interface between the processing component 502 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0168] The sensor assembly 514 includes one or more sensors for providing various aspects of status assessment for the electronic device 500. For example, the sensor assembly 514 can detect the open / closed state of the electronic device 500, the relative positioning of components, such as the display and keypad of the electronic device 500. The sensor assembly 514 can also detect changes in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and temperature changes of the electronic device 500. The sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0169] The communication component 516 is used to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0170] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.

[0171] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by the processor 520 of the electronic device 500 to perform the above method. For example, the non-transitory storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0172] Figure 7 FIG2 is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server.

[0173] Reference Figure 7 The electronic device 600 includes a processing component 622, which further includes one or more processors, and a memory resource represented by a memory 632 for storing instructions executable by the processing component 622, such as an application. The application stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute the instructions to perform the method provided in the embodiments of the present application.

[0174] The electronic device 600 may further include a power supply component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0175] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0176] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for automatically extracting metadata in the energy field, characterized by: include: determining first metadata of energy data, and determining parsed data of the energy data; The first metadata has a preset data structure; Inputting the first metadata into a first language model to obtain a plurality of prompt words for the energy data; The generated plurality of prompt words correspond to the energy types corresponding to the energy data; Inputting the plurality of prompt words and the parsed data of the energy data into a second language model to obtain second metadata of the energy data, such that a data structure of the second metadata corresponds to an energy type corresponding to the energy data; The determining of the first metadata of the energy data includes: extracting a first description text of the energy data from the energy data; When the amount of information in the first description text is less than or equal to a preset information amount threshold, performing information enhancement on the first description text to update the first description text; When the information amount of the first description text is greater than the information amount threshold, performing keyword extraction on the description text to obtain a plurality of keywords in the description text; Encapsulating the obtained multiple keywords according to the data structure to obtain the first metadata; The method further comprises: In response to the adjustment operation on the second metadata, the first language model and / or the second language model are adjusted to update the second metadata according to the adjusted first language model and / or the second language model.

2. The method according to claim 1, wherein In response to the adjustment operation on the second metadata, adjusting the first language model and / or the second language model to update the second metadata according to the adjusted first language model and / or the second language model includes: In response to an adjustment operation on the second metadata, obtaining adjustment data of the second metadata; The first language model and / or the second language model are adjusted according to the adjustment data of the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model.

3. The method according to claim 2, wherein The adjusting the first language model and / or the second language model according to the adjustment data of the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model, includes: determining the adjustment data of the first metadata according to the adjustment data of the second metadata; The first language model is trained using the adjusted data of the first metadata as first training data to obtain the trained first language model, and the process returns to the step of inputting the first metadata into the first language model to obtain multiple prompt words for the energy data.

4. The method according to claim 2, wherein The adjusting the first language model and / or the second language model according to the adjustment data of the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model, includes: updating the plurality of prompt words according to the adjustment data of the second metadata; The updated prompt words are used as second training data to train the second language model to obtain a trained second language model, and the process returns to the step of inputting the multiple prompt words and parsed data of the energy data into the second language model to obtain second metadata of the energy data, so that the data structure of the second metadata corresponds to the energy type corresponding to the energy data.

5. The method according to claim 1, wherein The method further comprises: The second metadata is stored together with the energy data.

6. An automatic metadata extraction device for the energy field, characterized in that: include: a data extraction module, configured to determine first metadata of energy data and determine parsed data of the energy data; The first metadata has a preset data structure; a prompt word generation module, configured to input the first metadata into a first language model to obtain a plurality of prompt words for the energy data; The generated plurality of prompt words correspond to the energy types corresponding to the energy data; a metadata generation module, configured to input the plurality of prompt words and parsed data of the energy data into a second language model to obtain second metadata of the energy data, such that a data structure of the second metadata corresponds to an energy type corresponding to the energy data; The data extraction module includes: a description extraction submodule, configured to extract a first description text of the energy data from the energy data; an information enhancement submodule, configured to, when the amount of information in the first description text is less than or equal to a preset information amount threshold, perform information enhancement on the first description text to update the first description text; a keyword submodule, configured to extract keywords from the description text when the amount of information in the first description text is greater than the information amount threshold, so as to obtain a plurality of keywords in the description text; A first metadata submodule, configured to encapsulate the obtained multiple keywords according to the data structure to obtain the first metadata; The automatic metadata extraction device for the energy field also includes: A model training module is used to adjust the first language model and / or the second language model in response to the adjustment operation on the second metadata, so as to update the second metadata according to the adjusted first language model and / or the second language model.

7. An electronic device, characterized in that: include: a processor, a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Metadata management method and device, computer equipment and storage medium

    CN116541752A

  • Data resource metadata semantic retrieval method and system based on large model

    CN118349690A