Enterprise carbon asset digitalization method and device based on multi-modal knowledge graph
By building a multimodal knowledge graph and using graph neural network model for feature dissemination, the problem of multi-source heterogeneous data integration in carbon asset management is solved, and the accuracy, intelligence and real-time improvement of carbon asset management is achieved.
Patent Information
- Application Number
- CN202510498436.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the field of carbon asset management, it is difficult for the existing technology to effectively integrate and integrate multi-source heterogeneous multimodal data, resulting in insufficient data silos and information extraction, affecting the accuracy and intelligence of carbon asset management.
By obtaining structured data, unstructured data and sensor data, preprocessing and entity relationship extraction, building a multimodal knowledge graph, and using graph neural network models for feature propagation and embedding operations, realizing the deep fusion of multimodal data.
It realizes panoramic data fusion and dynamic monitoring of corporate carbon assets, improves the accuracy, intelligence and real-time nature of carbon asset management, reduces model errors, and enhances the competitiveness of enterprises in the carbon market.
Smart Images

Figure CN120012905A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph technology, and in particular to a method and device for digitizing enterprise carbon assets based on a multimodal knowledge graph. Background Art
[0002] In the context of global response to climate change, countries have successively launched carbon emission trading systems and carbon markets to promote the realization of carbon emission reduction targets through market mechanisms. As important participants in the carbon market, enterprises are faced with complex issues such as carbon emission monitoring, accounting, trading and compliance.
[0003] Despite the continuous development of related technologies, the following bottlenecks still exist in the field of carbon asset management: data are multi-source and heterogeneous, making integration difficult; the multimodal data involved in carbon asset management is complex; and existing technologies face great challenges in data fusion and consistency. Summary of the invention
[0004] The present invention provides a method and device for digitizing enterprise carbon assets based on a multimodal knowledge graph, so as to solve the defects in the prior art.
[0005] In a first aspect, the present invention provides a method for digitizing enterprise carbon assets based on a multimodal knowledge graph, comprising: acquiring multimodal data related to the enterprise's carbon asset management; the multimodal data includes structured data, unstructured data and sensor data; the structured data includes numerical data, the unstructured data includes text data, and the sensor data includes time series data; preprocessing the multimodal data to form a multimodal data set; the multimodal data set includes a structured data set, a text data set and sensor time series data; extracting entity-entity relationships from the structured data set and the text data set, and embedding the sensor time series data into the entity to construct a knowledge graph including multimodal information; performing an embedding operation on the constructed knowledge graph to map the knowledge graph to a lower dimensional vector space.
[0006] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, entities and entity relationships are extracted from structured data sets and text data sets, and sensor time series data are embedded into entities to construct a knowledge graph including multimodal information, including: extracting entities and entity relationships based on the structural relationships of the structured data set, extracting entities and entity relationships from the text data set based on a pre-trained BERT model to construct entity nodes of the knowledge graph and the relationships between entity nodes; mapping sensor entity data to feature vectors of entity nodes to achieve the fusion of information of different modalities into a unified knowledge graph.
[0007] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, sensor entity data is mapped to a feature vector of an entity node, specifically: Specifically: ; in, is the feature representation of entity e, s For Entity e Sensor timing data, x For Entity e Basic information, t For Entity e time information.
[0008] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, an embedding operation is performed on the constructed knowledge graph, including: embedding the constructed knowledge graph with the loss function of the GraphSAGE model as the goal; The loss function of the GraphSAGE model L Specifically: ; in, , , Represent the embedding vectors of the head entity, entity relation, and tail entity respectively, is the distance function, represents a triple, h Represents the head entity, r Represents entity relationships, t Represents the tail entity, K Represents a set of triples.
[0009] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, after embedding the constructed knowledge graph and mapping the knowledge graph to a lower-dimensional vector space, it also includes: using a graph neural network model to propagate features of the knowledge graph.
[0010] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, a graph neural network model is used to propagate features of the knowledge graph, specifically: ; in, Indicates The feature representation of layer nodes, Indicates The feature representation of layer nodes, represents the activation function, are the learnable parameters of the model, Representation Node i Neighborhood set of , For Node and neighbor nodes j The connection weights between Representation Node j In the l The feature representation of the layer.
[0011] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, structured data is preprocessed, including: detecting and removing outliers on the structured data; and filling missing values on the structured data.
[0012] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, unstructured data is preprocessed, including: using a natural language processing method to extract key text data related to the enterprise's carbon asset management from a text data set.
[0013] According to a method for digitizing enterprise carbon assets based on a multimodal knowledge graph provided by the present invention, sensor data is preprocessed, including: smoothing the sensor data using a sliding average method; and dynamically reducing the noise of the sensor data using a Kalman filtering method.
[0014] In a second aspect, the present invention further provides a device for digitizing enterprise carbon assets based on a multimodal knowledge graph, comprising: The first processing module is used to obtain multimodal data related to the carbon asset management of the enterprise; the multimodal data includes structured data, unstructured data and sensor data; the structured data includes numerical data, the unstructured data includes text data, and the sensor data includes time series data; A second processing module is used to pre-process the multimodal data to form a multimodal data set; the multimodal data set includes a structured data set, a text data set and sensor time series data; The third processing module is used to extract entity-entity relationships from structured data sets and text data sets, and embed sensor time series data into entities to build a knowledge graph including multimodal information; The fourth processing module is used to perform embedding operations on the constructed knowledge graph and map the knowledge graph to a lower-dimensional vector space.
[0015] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for digitizing enterprise carbon assets based on a multimodal knowledge graph as described above are implemented.
[0016] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for digitizing enterprise carbon assets based on a multimodal knowledge graph.
[0017] The enterprise carbon asset digitization method and device based on multimodal knowledge graph provided by the present invention significantly improves the accuracy and intelligence level of carbon asset management, and has the following beneficial effects compared with the prior art: (1) The present invention introduces multimodal data collection, integrates structured data (carbon emissions, carbon trading records), unstructured data (policy texts, carbon disclosure reports) and time-series sensor data, and constructs a multi-source data fusion mechanism; it effectively solves the problem of data islands in carbon asset management, comprehensively covers all aspects of carbon asset management, and realizes panoramic data fusion and dynamic monitoring of corporate carbon assets.
[0018] (2) The present invention ensures data consistency and reliability through a multi-step preprocessing process including outlier detection, missing value filling, smoothing and noise reduction (sliding average and Kalman filtering); significantly improves the quality and stability of multi-source heterogeneous data, reduces the noise and uncertainty in carbon emission monitoring data, reduces model errors, and improves the accuracy of carbon asset valuation.
[0019] (3) Based on natural language processing (NLP) models such as BERT, the present invention automatically extracts key information such as carbon emission quotas, carbon taxes, and trading rules from policy and regulatory texts, generates entity-relationship-entity triples, and constructs a multimodal knowledge graph of corporate carbon assets, thus achieving structured expression of unstructured data. This makes up for the shortcomings of traditional methods in extracting insufficient information from policy and regulatory texts, strengthens the binding force and dynamic adaptability of carbon market rules on corporate carbon asset management, and makes carbon asset valuation more timely and compliant.
[0020] The present invention comprehensively improves the accuracy, intelligence and real-time performance of carbon asset management, realizes the deep integration of multimodal data and dynamic carbon asset valuation, and not only solves the problems of data silos, insufficient information extraction and untimely risk warning in traditional carbon asset management, but also significantly enhances the competitiveness of enterprises in the carbon market, providing a solid technical guarantee for the green and low-carbon development of enterprises. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 It is a flow chart of the enterprise carbon asset digitization method based on multimodal knowledge graph provided by the present invention; Figure 2 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] It should be noted that, in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0025] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0026] Combine the following Figure 1-Figure 2 The present invention describes a method and device for digitizing enterprise carbon assets based on a multimodal knowledge graph.
[0027] Figure 1 is a flow chart of the enterprise carbon asset digitization method based on the multimodal knowledge graph provided by the present invention, such as Figure 1 As shown, including but not limited to the following steps: Step 101: Acquire multimodal data related to the enterprise's carbon asset management.
[0028] Carbon asset management involves multi-source heterogeneous data, which come from different dimensions and formats and have complex multimodal characteristics. Effectively collecting and preprocessing these data is the basis for building a digital neural network model for carbon assets.
[0029] The main sources of multimodal data include: (1) Structured data Structured data includes numerical data such as corporate carbon emissions, carbon trading records, and electricity consumption. This type of data is usually stored in databases or tables for direct reading and processing.
[0030] For example: ; in, Represents a collection of structured data, For the structured data items, .
[0031] (2) Unstructured data Such as policy and regulatory texts, corporate carbon disclosure reports, environmental news, etc., this type of data usually exists in the form of natural language text, and useful information needs to be extracted through natural language processing (NLP) technology.
[0032] For example: ; Represents a collection of unstructured text data, Indicates Text data, .
[0033] (3) Sensor data Carbon emission monitoring equipment on corporate production equipment and sensors such as smart meters can provide continuous real-time monitoring data on carbon emissions, usually in the form of time series data.
[0034] For example: ; Indicates time Sensor data at the moment, Indicates d The sensor readings, .
[0035] Step 102: Preprocess the multimodal data to form a multimodal data set.
[0036] Since multimodal data are not in uniform formats and have varying quality, preprocessing is crucial. This paper adopts a multi-step preprocessing process, including data cleaning, format conversion, missing value filling and feature extraction, to ensure data quality and consistency.
[0037] (1) Preprocessing of structured data (1.1) Outlier detection and removal For structured data, a standardized method is used to detect outliers, and outliers are eliminated using the triple standard deviation principle: ; in, Represents a numerical data. represents the mean of all numerical data, Represents the standard deviation of all numerical data. If >3, then judge As outliers, they are removed.
[0038] (1.2) Missing value filling For small proportions of missing data, the mean imputation method can be used to fill in the missing values, that is, the average value of multiple non-missing values is used to fill in the missing values.
[0039] (2) Unstructured data feature extraction Using natural language processing (NLP) methods, key text data related to the company's carbon asset management, such as carbon emission quotas, penalty clauses, etc.
[0040] (3) Sensor data processing Sensor data has time series characteristics and needs to be processed using smoothing and noise reduction methods: Optionally, the present invention uses a sliding average method to smooth the sensor data; and, The Kalman filter method is used to dynamically reduce the noise of sensor data.
[0041] (4) Optionally, since the numerical ranges of different modal data vary greatly, the numerical data with large numerical range differences can be standardized to ensure the convergence speed and accuracy during the model training process.
[0042] The standardization method may be a maximum-minimum normalization method.
[0043] After the above data cleaning and preprocessing, the present invention can obtain a set of high-quality multimodal data sets, including structured data sets, text data sets and sensor time series data. These data will be used as input to enter the next stage of multimodal knowledge graph construction.
[0044] Step 103: Extract entity-entity relationships from the structured dataset and the text dataset, and embed the sensor time series data into the entities to build a knowledge graph including multimodal information.
[0045] Multimodal knowledge graph maps multimodal information (such as text, sensor time series data, table data, etc.) from different data sources into a unified representation space, and stores and displays it in the form of a graph structure. This method can effectively depict the complex relationship between corporate carbon assets, provide rich contextual information and structured features for neural network models, and improve the intelligence and automation level of carbon asset management.
[0046] In the present invention, the purpose of constructing a multimodal knowledge graph of corporate carbon assets is to comprehensively reflect the dynamic changes and potential correlations of carbon assets in different business scenarios by integrating multi-source data such as corporate carbon emissions, equipment energy efficiency, carbon trading records, policies and regulations.
[0047] Define the core entities and relationships in enterprise carbon asset management: Core entity E = { }, where E represents an entity set, including the following main entities: enterprises, production equipment, carbon emission sources, policies and regulations, and carbon markets.
[0048] Relationship definition: The relationship between entities is defined as a set: R = { The main relationships include: equipment-emission source relationship (equipment generates carbon emissions), enterprise-equipment relationship (equipment belongs to the enterprise), enterprise-market relationship (enterprise participates in the carbon trading market), and policy constraint relationship (carbon emissions are subject to policy constraints).
[0049] Entities and relations are extracted from structured and unstructured data to generate triple representations : ; in, is the head entity, is the tail entity, Indicates the entity relationship between two entities. For example: (equipment 1, emission source, carbon emission source A).
[0050] As an optional embodiment, step 103 further includes the following steps: (1) Extract entities and entity relationships based on the structural relationships of structured data sets, and extract entities and entity relationships from text data sets based on the pre-trained BERT model to construct entity nodes and relationships between entity nodes of the knowledge graph.
[0051] For structured data sets, there are often data with clear formats and organizations, such as databases, tables, etc. The process of extracting entities and relationships from structured data sets is relatively straightforward because it already contains clear definitions and associations. Entities and their relationships can be directly extracted from structured data using SQL queries or ETL tools.
[0052] The relationship of structured data can be determined through foreign key constraints or through joins between two or more tables. For example, a company's "Carbon Emission Record" can be linked to the "Carbon Emission Source" table to extract the relationship between the company and its carbon emission sources.
[0053] For text datasets, the pre-trained BERT model is used for named entity recognition (NER) and relationship extraction; the BERT model is used to identify entities in the text, such as clauses in policies and regulations, company names, etc.; based on contextual understanding, the relationship between entities is extracted from the text. For example, the impact of a policy on a specific company or industry can be identified.
[0054] (2) Map the sensor entity data into the feature vector of the entity node to integrate information from different modalities into a unified knowledge graph.
[0055] Specifically: ; in, is the feature representation of entity e, s For Entity e Sensor timing data, x For Entity e Basic information, t For Entity e time information.
[0056] Step 104: Perform an embedding operation on the constructed knowledge graph to map the knowledge graph to a lower-dimensional vector space.
[0057] The constructed knowledge graph needs to be embedded to map the graph into a low-dimensional vector space for use in subsequent neural network models. Optionally, embedding operations are performed on the constructed knowledge graph, including: With the goal of minimizing the loss function of the GraphSAGE model, embedding operations are performed on the constructed knowledge graph; The loss function of the GraphSAGE model L Specifically: ; in, , , Represent the embedding vectors of the head entity, entity relation, and tail entity respectively, is the distance function, represents a triple, h Represents the head entity, r Represents entity relationships, t Represents the tail entity, K Represents a set of triples.
[0058] Based on the contents of the above embodiments, as an optional embodiment, the enterprise carbon asset digitization method based on a multimodal knowledge graph provided by the present invention, after embedding the constructed knowledge graph and mapping the knowledge graph to a lower-dimensional vector space, also includes: using a graph neural network model to propagate features of the knowledge graph to enhance the relationship representation between nodes.
[0059] Specifically: ; in, Indicates The feature representation of layer nodes, Indicates The feature representation of layer nodes, represents the activation function, are the learnable parameters of the model, Representation Node i Neighborhood set of , For Node and neighbor nodes j The connection weights between Representation Node j In the l The feature representation of the layer.
[0060] It should be noted that when constructing a knowledge graph, entities are mapped to nodes in the graph (i.e., entity nodes), which means there is a one-to-one mapping relationship between entities and nodes.
[0061] When an entity is converted to a node, all its related information (such as attributes and relationships with other entities) should be retained in the graph structure as much as possible. This can be achieved through node feature vectors as well as edges in the graph. As new entities are added or the attributes of existing entities change, the corresponding nodes and their features will be updated to maintain the consistency and accuracy of the knowledge graph.
[0062] Nodes are the specific manifestations of entities in the graph structure. There is a direct mapping relationship between the two, and nodes not only inherit all the attributes of the entity, but also establish connections with other nodes in the graph structure through edges. This conversion process from entities to nodes enables the present invention to use the powerful capabilities of graph neural networks to perform complex analysis and reasoning.
[0063] After the knowledge graph is constructed and embedded, a knowledge graph containing multi-source information is generated, including all entity information and the relationships between entities.
[0064] On the other hand, the present invention also provides a device for digitizing enterprise carbon assets based on a multimodal knowledge graph, the device comprising: The first processing module is used to obtain multimodal data related to the carbon asset management of the enterprise; the multimodal data includes structured data, unstructured data and sensor data; the structured data includes numerical data, the unstructured data includes text data, and the sensor data includes time series data; A second processing module is used to pre-process the multimodal data to form a multimodal data set; the multimodal data set includes a structured data set, a text data set and sensor time series data; The third processing module is used to extract entity-entity relationships from structured data sets and text data sets, and embed sensor time series data into entities to build a knowledge graph including multimodal information; The fourth processing module is used to perform embedding operations on the constructed knowledge graph and map the knowledge graph to a lower-dimensional vector space.
[0065] It should be noted that the enterprise carbon asset digitization device based on multimodal knowledge graph provided in an embodiment of the present invention can execute the enterprise carbon asset digitization method based on multimodal knowledge graph described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0066] The enterprise carbon asset digitization method and device based on multimodal knowledge graph provided by the present invention significantly improves the accuracy and intelligence level of carbon asset management, and has the following beneficial effects compared with the prior art: (1) The present invention introduces multimodal data collection, integrates structured data (carbon emissions, carbon trading records), unstructured data (policy texts, carbon disclosure reports) and time-series sensor data, and constructs a multi-source data fusion mechanism; it effectively solves the problem of data islands in carbon asset management, comprehensively covers all aspects of carbon asset management, and realizes panoramic data fusion and dynamic monitoring of corporate carbon assets.
[0067] (2) The present invention ensures data consistency and reliability through a multi-step preprocessing process including outlier detection, missing value filling, smoothing and noise reduction (sliding average and Kalman filtering); significantly improves the quality and stability of multi-source heterogeneous data, reduces the noise and uncertainty in carbon emission monitoring data, reduces model errors, and improves the accuracy of carbon asset valuation.
[0068] (3) Based on natural language processing (NLP) models such as BERT, the present invention automatically extracts key information such as carbon emission quotas, carbon taxes, and trading rules from policy and regulatory texts, generates entity-relationship-entity triples, and constructs a multimodal knowledge graph of corporate carbon assets, thus achieving structured expression of unstructured data. This makes up for the shortcomings of traditional methods in extracting insufficient information from policy and regulatory texts, strengthens the binding force and dynamic adaptability of carbon market rules on corporate carbon asset management, and makes carbon asset valuation more timely and compliant.
[0069] The technical means of the present invention comprehensively improves the accuracy, intelligence and real-time performance of carbon asset management, realizes the deep integration of multimodal data and dynamic carbon asset valuation, and not only solves the problems of data silos, insufficient information extraction and untimely risk warning in traditional carbon asset management, but also significantly enhances the competitiveness of enterprises in the carbon market, providing a solid technical guarantee for the green and low-carbon development of enterprises.
[0070] Figure 2 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 2 As shown, the electronic device may include: a processor 210, a communication interface 220, a memory 230 and a communication bus 240, wherein the processor 210, the communication interface 220 and the memory 230 communicate with each other through the communication bus 240. The processor 210 may call the logic instructions in the memory 230 to execute the enterprise carbon asset digitization method based on the multimodal knowledge graph.
[0071] In addition, the logic instructions in the above-mentioned memory 230 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0072] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the enterprise carbon asset digitization method based on a multimodal knowledge graph provided in the above-mentioned embodiments.
[0073] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the enterprise carbon asset digitization method based on a multimodal knowledge graph provided in the above-mentioned embodiments.
[0074] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for digitizing enterprise carbon assets based on multimodal knowledge graph, characterized in that: include: Obtain multimodal data related to the company’s carbon asset management; Multimodal data includes structured data, unstructured data, and sensor data; Structured data includes numerical data, unstructured data includes text data, and sensor data includes time series data; Preprocess the multimodal data to form a multimodal data set; The multimodal data set includes a structured data set, a text data set, and sensor time series data; Extract entity-entity relationships from structured and text datasets, and embed sensor time series data into entities to build a knowledge graph that includes multimodal information. The constructed knowledge graph is embedded to map the knowledge graph into a lower-dimensional vector space.
2. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: Entity-entity relationships are extracted from structured and text datasets, and sensor time series data are embedded into entities to build a knowledge graph that includes multimodal information, including: Extract entities and entity relationships based on the structural relationships of structured data sets, and extract entities and entity relationships from text data sets based on the pre-trained BERT model to construct entity nodes and relationships between entity nodes in the knowledge graph; Map the sensor entity data to the feature vector of the entity node to integrate the information of different modalities into a unified knowledge graph.
3. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 2 is characterized in that: Map the sensor entity data to the feature vector of the entity node, specifically: ; in, is the feature representation of entity e, s For Entity e Sensor timing data, x For Entity e Basic information, t For Entity e time information.
4. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: Embed the constructed knowledge graph, including: With the goal of minimizing the loss function of the GraphSAGE model, embedding operations are performed on the constructed knowledge graph; The loss function of the GraphSAGE model L Specifically: ; in, , , Represent the embedding vectors of the head entity, entity relation, and tail entity respectively, is the distance function, represents a triple, h Represents the head entity, r Represents entity relationships, t Represents the tail entity, K Represents a set of triples.
5. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: After embedding the constructed knowledge graph and mapping it to a lower-dimensional vector space, the following steps are also included: Use graph neural network model to propagate features of knowledge graph.
6. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 5 is characterized in that: The graph neural network model is used to propagate features of the knowledge graph, specifically: ; in, Indicates The feature representation of layer nodes, Indicates The feature representation of layer nodes, represents the activation function, are the learnable parameters of the model, Representation Node i Neighborhood set of , For Node and neighbor nodes j The connection weights between Representation Node j In the l The feature representation of the layer.
7. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: Preprocess structured data, including: Outlier detection and removal for structured data; and, Fill missing values in structured data.
8. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: Preprocess unstructured data, including: Natural language processing methods are used to extract key text data related to the company's carbon asset management from text datasets.
9. The enterprise carbon asset digitization method based on multimodal knowledge graph according to claim 1 is characterized in that: Preprocess the sensor data, including: Smoothing the sensor data using a sliding average; and, The Kalman filter method is used to dynamically reduce the noise of sensor data.
10. An enterprise carbon asset digitization device based on multimodal knowledge graph, characterized in that: include: A first processing module is used to obtain multimodal data related to the carbon asset management of the enterprise; Multimodal data includes structured data, unstructured data, and sensor data; Structured data includes numerical data, unstructured data includes text data, and sensor data includes time series data; A second processing module is used to pre-process the multimodal data to form a multimodal data set; The multimodal data set includes a structured data set, a text data set, and sensor time series data; The third processing module is used to extract entity-entity relationships from structured data sets and text data sets, and embed sensor time series data into entities to build a knowledge graph including multimodal information; The fourth processing module is used to perform embedding operations on the constructed knowledge graph and map the knowledge graph to a lower-dimensional vector space.
Citation Information
Patent Citations
Multi-modal enterprise knowledge graph completion method and device for key information mining
CN116737961A
Energy industry knowledge graph optimizing and updating method based on machine learning
CN118779465A
Carbon asset management method and system based on knowledge graph
CN119067254A
Cited By
Carbon-source-sink spatio-temporal dynamic knowledge graph construction method and system, terminal and storage medium
CN121351978A
A method and system for constructing a spatiotemporal dynamic knowledge graph of carbon sources and sinks, a terminal and a storage medium
CN121351978B
Asset data classification method and device based on knowledge graph
CN121723246A
Carbon emission report generation method, electronic equipment, storage medium and program product
CN122311164A