Space entity alignment method for urban carbon emission data
By using entity encoding and large-model technology methods in carbon emission data processing, the problems of multi-source heterogeneous data and complex spatiotemporal correlation are solved, and efficient alignment of carbon emission sources and optimization support for carbon reduction technology are achieved.
Patent Information
- Application Number
- CN202510041956.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively deal with multi-source heterogeneous carbon emission data, complex spatiotemporal correlations and different data standards, resulting in the limitation of the identification accuracy of carbon emission sources and the accuracy and operability of the optimization path of carbon reduction technology.
Innovative entity coding methods and large model technology are adopted, and spatial address names are processed through BERT encoding, time encoding is processed by time encoding, and carbon emission characteristics are processed by PCA, and the total characteristics of carbon emission entities are formed by combining latitude and longitude coordinates. The CSLS similarity calculation method is used to initially screen entity pairs, and fine matching of multi-dimensional information is performed through large language models, ultimately achieving efficient alignment of carbon emission entities.
It improves the accuracy of the alignment of carbon emission source entities, can effectively utilize multi-dimensional spatio-temporal information, improves the reliability of carbon emission prediction and carbon reduction technology selection, and solves the problems of data heterogeneity and complexity in traditional methods.
Smart Images

Figure CN119961691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon emission reduction, and in particular to a spatial entity alignment method for urban carbon emission data. Background Art
[0002] One form of spatial data is point of interest (POI), which represents geographic entities with spatial and semantic information, such as restaurants, hotels, hospitals, schools, and carbon emission sources. Fusion of point of interest data from different sources (such as the Environmental Protection Agency and the Department of Housing and Urban-Rural Development) helps provide higher quality and more complete carbon emission entity information for location-based applications.
[0003] Geospatial Entity Resolution is the first step in spatial data integration. Its goal is to discover and associate points of interest from different data sources that point to the same entity in the physical world. The matching entity pairs obtained by entity resolution will be integrated into a unified database after data fusion, data cleaning and other operations, thereby obtaining high-quality spatial data. The spatial entity resolution task inputs two spatial data sets from different sources. The data sets are stored in the form of a table. Each row represents a record representing a spatial entity (i.e., a point of interest), and each column represents a spatial attribute containing longitude and latitude information, or text attributes such as name, address, telephone number, type, etc. The goal is to output record pairs from different data sources that point to the same real-world object. The process is usually divided into two steps: blocking and matching. Blocking preliminarily screens out possible matching spatial data based on longitude and latitude positions or text similarity, thereby reducing the number of entity pairs to be compared. Matching determines whether two data point to the same real-world entity by comparing the similarity of each attribute (name, address, longitude, latitude, category, carbon emissions, etc.).
[0004] Carbon emission data has the following characteristics: carbon emission data usually has spatial characteristics, temporal characteristics, category characteristics, carbon emissions and a mixture of Chinese and English; when describing the spatial location of carbon emission entities, it also has an address name focused on a certain county and specific longitude and latitude coordinates; carbon emissions can be subdivided into different fields based on the production end, such as coal, coke, coke oven gas, crude oil, gasoline, kerosene, liquefied petroleum gas and other fields; due to different statistical standards, the entity scope corresponding to carbon emission data is not strictly the same.
[0005] As global emission reduction targets become increasingly urgent, how to accurately identify and align carbon emission entities from different data sources has become an important part of achieving carbon reduction targets. In particular, in carbon emission sources involving different data sources, traditional carbon emission source alignment methods face many challenges due to the heterogeneity of data sources, different data standards, and complex spatiotemporal relationships. Existing technologies are usually unable to effectively handle the alignment of multi-source heterogeneous data, complex spatiotemporal associations, and entities from different original data sources involved in the infrastructure field, thereby limiting the accuracy and operability of the optimization path of carbon reduction technology.
[0006] Therefore, there is an urgent need for a new technical approach to solve this problem, which can efficiently align carbon emission data from different sources with different data standards, improve the accuracy of identifying carbon emission sources, and promote the scientific application of carbon reduction technologies. Summary of the invention
[0007] Purpose of the invention: The purpose of the present invention is to provide a spatial entity alignment method for urban carbon emission data, aiming to accurately identify and align carbon emission entities in different data through innovative entity coding methods and large model technology, thereby providing reliable support for carbon emission prediction and carbon reduction technology selection. Through the precise matching of multi-dimensional spatiotemporal information and entities, the present invention can provide a new technical solution for the accurate identification of carbon emission sources and the optimization of carbon reduction paths.
[0008] Technical solution: To achieve the purpose of the present invention, the technical solution adopted by the present invention is: a spatial entity alignment method for urban carbon emission data, the method comprising the following steps:
[0009] Step 1. Carbon emission entity code:
[0010] In view of the fact that carbon emission data contains both spatial address names and specific longitude and latitude coordinates, BERT encoding is used to process the spatial address names of carbon emission data; time series encoding is used to process the time information of carbon emission entities; in particular, PCA processing is performed on carbon emissions that are subdivided into different fields based on the production end to form carbon emission characteristics. The BERT encoding, time series encoding, carbon emission characteristics, and longitude and latitude coordinates obtained by encoding are connected to form the overall characteristics of the carbon emission entity, making full use of the different attributes of carbon emission data.
[0011] Step 2. Carbon emission entity similarity calculation:
[0012] The CSLS similarity calculation method is used to measure the entity coding similarity between different data sources, so as to preliminarily screen out entity pairs that may point to the same carbon emission source.
[0013] Step 3. Combining multi-dimensional information:
[0014] For the candidate entity pairs that have been initially screened, the large language model is used for further fine matching. Combined with the multi-dimensional information of the entity, such as latitude and longitude, address, category, name, carbon emissions, etc., the reasoning ability of the large model is used for accurate matching to obtain the entity pairs that are most likely to point to the same carbon emission source. In the process of entity alignment, the multi-dimensional information of the carbon emission source is fully utilized, including but not limited to carbon emissions, industry attributes, geographic location, etc. Before submitting the data to the large model, the data is preprocessed. For the name information, according to the multiple names of the data in the data set, the name information that is strictly included is deleted, and the remaining name information is connected; for the industry attribute information, the same method as the name information is used to delete the industry attribute information that is strictly included, and the remaining industry attribute information is connected; for the geographic information, the Euclidean distance of the longitude and latitude of the two entities is calculated. If the information of a region is submitted, the Euclidean distance of the center point of the region is calculated; for the carbon emission characteristics, the sum of carbon emissions based on the production end segmented into different fields is calculated. Submit these processed information to the large model.
[0015] Step 4. Fine matching of candidate entity pairs
[0016] In the process of asking questions, the two-step method of reasoning and rethinking is used to improve the recognition accuracy of the big model. First, let the big model generate the semantic similarity, geographical location similarity, industry attribute similarity, description similarity, carbon emission similarity and total similarity of the carbon emission entity based on the information submitted to it, and then let the big model judge whether the candidate entity pair matches based on the similarity obtained from the first question and answer. Combined with this information, a comprehensive evaluation of the entity can further improve the accuracy of alignment and solve the problems of heterogeneity and complexity in entity alignment. In order to solve the problem that the entity range corresponding to the carbon emission data is not strictly the same, the carbon emission similarity generated by the big model is used. When the carbon emission similarity is extremely low, the big model will reduce the total similarity of the entity pair during the rethinking process, and it is more likely to reject wrong matches. In view of the mixed characteristics of Chinese and English in carbon emission data, English prompt words are designed to ask questions to the big model, and good results have been achieved.
[0017] Step 5. Efficient alignment of candidate entity pairs:
[0018] At the same time, in order to reduce the excessive consumption of computing power of the large model, the 1-10-20 questioning method is adopted, that is, the candidate entity pairs with the greatest possibility of matching are submitted to the large model first, so that the large model can judge whether it is confident enough to believe that the candidate entity pairs are matched successfully, and then the candidate entity pairs with the top 10 similarities are submitted to the large model, so that it can judge whether there is a certain entity pair whose probability of successful matching is much higher than that of the other candidate entity pairs, and then the candidate entity pairs with the top 10 similarities are submitted to the large model, so that it can judge whether there is a certain entity pair whose probability of successful matching is much higher than that of the other candidate entity pairs. If none of the first 20 candidate pairs are matched successfully, the match fails.
[0019] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0020] (1) It can make full use of the different attributes of the data set and improve the accuracy of carbon emission source entity alignment as much as possible according to the number of attributes. It can make maximum use of the attributes of different carbon emission data sets and has a certain universality, filling the gap in this research field.
[0021] (2) It has the feature of zero-shot, does not require extra data for training, and is easy to use. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Example diagram of spatial entity alignment for carbon emission data;
[0023] Figure 2 Flowchart for generating candidate pairs using the encoding module;
[0024] Figure 3 Flowchart for large model inference. DETAILED DESCRIPTION
[0025] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments.
[0026] As shown in the figure, the purpose of the present invention is to provide a spatial entity alignment method for urban carbon emission data, aiming to accurately identify and align carbon emission entities in different data sets through innovative carbon emission entity encoding methods and large model technology, thereby providing reliable support for carbon emission prediction and carbon reduction technology selection. Through the precise matching of multi-dimensional spatiotemporal information and entities, the present invention can provide a new technical solution for the accurate identification of carbon emission sources and the optimization of carbon reduction paths. The specific execution steps are as follows:
[0027] Step 1. Carbon emission entity code:
[0028] Use BERT encoding to process its spatial address name; use time series encoding to process the time information of carbon emission entities; in particular, PCA processing is performed on carbon emissions that are subdivided into different fields based on the production end to form carbon emission characteristics. The BERT encoding, time series encoding, carbon emission characteristics, and longitude and latitude coordinates obtained by encoding are connected to form the overall characteristics of the carbon emission entity.
[0029] The vectors obtained by using these two encodings for the two data sets are W name1 , W time1 , W carbon 1. W pos1 ∈R n1 *c , W name2 , W time2 , W carbon2 , W pos2 ∈R n2*c , where n1 is the number of entities in dataset 1, n2 is the number of entities in dataset 2, c is a variable parameter, indicating the length of the corresponding vector after each entity is encoded, and the subscripts name, time, carbon, and pos represent the encoding of semantic, timing, carbon emissions, and location information, respectively.
[0030] Step 2. Carbon emission entity similarity calculation:
[0031] The CSLS (Cross-domain Similarity Local Scaling) similarity calculation method is used to measure the similarity of entity codes from different sources, so as to preliminarily screen out entity pairs that may point to the same carbon emission source.
[0032] Let W1=(W name 1.W time1 , W carbon 1. W pos1 ), W2=(W name 2.W time2 , W carbon 2. W pos2 )
[0033]
[0034] in:
[0035]
[0036] W=CSLS(W1 i* ,W2 i* ) 1<i<n1,1<j<n2
[0037]
[0038] Blocking is a collection of candidate entity pairs
[0039] Step 3. Combining multi-dimensional information:
[0040] The candidate entity pairs (i, j) that have been initially screened are further matched using the large language model. Combined with the multi-dimensional information of the entity, such as latitude and longitude, address, category, name, etc., the reasoning ability of the large model is used to accurately match and obtain the entity pairs that are most likely to point to the same carbon emission source. In the entity alignment process, the multi-dimensional information of the carbon emission source is fully utilized, including but not limited to carbon emissions, industry attributes, geographic location, etc. Before submitting the data to the large model, the data is pre-processed. For the name information, based on the multiple names of the data in the data set, the name information that is strictly included is deleted, and the remaining name information is connected to obtain the Name i ,Name j ; For industry attribute information, use the same method as name information, delete the industry attribute information that is strictly included, and connect the remaining industry attribute information Category i , Category j ; For geographic information, calculate the Euclidean distance of the longitude and latitude of two entities. If a region is submitted, calculate the Euclidean distance of the center point of the region. i , Distance j ; For the carbon emission characteristics, calculate the total carbon emissions based on the production end segmented into different fields Carbon i , Carbon j . Submit this processed information to the big model.
[0041] Step 4. Fine matching of candidate entity pairs
[0042] In the process of asking questions, the two-step method of reasoning and rethinking is used to improve the recognition accuracy of the large model. First, let the large model generate the semantic similarity, geographical location similarity, industry attribute similarity, description similarity and total similarity of the carbon emission entity based on the information submitted to it, and then let the large model judge whether the candidate entity pair matches based on the similarity obtained in the first question and answer. By combining this information to comprehensively evaluate the entities, the accuracy of alignment can be further improved, and the heterogeneity and complexity problems in entity alignment can be solved. In order to solve the problem that the entity range corresponding to the carbon emission data is not strictly the same, the carbon emission similarity generated by the large model is used. When the carbon emission similarity is extremely low, the large model will reduce the total similarity of the entity pair during the rethinking process, and it is more likely to reject wrong matches. In view of the mixed characteristics of Chinese and English in carbon emission data, English prompts are designed to ask questions to the large model, and good results have been achieved. See the appendix for specific prompts.
[0043] Step 5. Efficient alignment of candidate entity pairs:
[0044] At the same time, in order to reduce the excessive consumption of computing power of the large model, the 1-10-20 questioning method is adopted, that is, the candidate entity pairs with the greatest possibility of matching are first submitted to the large model, allowing the large model to determine whether it is confident enough to believe that the candidate entity pairs are matched successfully, and then the candidate entity pairs with the top 10 similarities are submitted to the large model to let it determine whether there is a certain entity pair whose probability of successful matching is much higher than the other candidate entity pairs, and then the candidate entity pairs with the top 10 similarities are submitted to the large model to let it determine whether there is a certain entity pair whose probability of successful matching is much higher than the other candidate entity pairs. If none of the first 20 candidate pairs are matched successfully, the match fails.
[0045] Wherein, the computer program code of the method of the present invention is as follows:
[0046]
[0047]
[0048]
[0049] The technical means disclosed in the scheme of the present invention are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical schemes composed of any combination of the above-mentioned technical features. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also regarded as the protection scope of the present invention.
Claims
1. A spatial entity alignment method for urban carbon emission data, characterized in that: The method comprises the following steps: Step 1. Carbon emission entity code: In view of the fact that carbon emission data contains both spatial address names and specific longitude and latitude coordinates, BERT encoding is used to process the spatial address names of carbon emission data; time series encoding is used to process the time information of carbon emission entities; PCA processing is performed on carbon emissions that are subdivided into different fields based on the production end to form carbon emission characteristics; the BERT encoding, time series encoding, carbon emission characteristics, and longitude and latitude coordinates obtained by encoding are connected to form the overall characteristics of carbon emission entities; Step 2. Carbon emission entity similarity calculation: The CSLS similarity calculation method is used to measure the similarity of entity encodings in different data, so as to preliminarily screen out entity pairs that may point to the same carbon emission source; Step 3. Combining multi-dimensional information: For the candidate entity pairs initially screened, a large language model is used for further fine matching; combined with the multi-dimensional information of the entity's latitude and longitude, address, category, and name, the reasoning ability of the large model is used for accurate matching to obtain the entity pairs that are most likely to point to the same carbon emission source; in the entity alignment process, full use is made of the multi-dimensional information of the carbon emission source, including but not limited to carbon emissions, industry attributes, geographic location, and carbon emissions; Step 4. Fine matching of candidate entity pairs In the questioning process, the two-step questioning method of reasoning and rethinking is used to improve the recognition accuracy of large models; Step 5. Efficient alignment of candidate entity pairs: In order to reduce the excessive consumption of computing power of large models, the 1-10-20 questioning method is adopted.
2. The spatial entity alignment method for urban carbon emission data according to claim 1 is characterized in that: In step 3, before submitting the data to the big model, the data is preprocessed. For name information, based on multiple names of the data in the data set, the name information that is strictly included is deleted, and the remaining name information is connected; for industry attribute information, the same method as the name information is used to delete the industry attribute information that is strictly included, and the remaining industry attribute information is connected; For geographic information, the Euclidean distance between the longitude and latitude of two entities is calculated. If information about a region is submitted, the Euclidean distance between the center of the region is calculated. For carbon emission characteristics, the total carbon emissions of different fields based on the production end are calculated. Submit this processed information to the big model.
3. The spatial entity alignment method for urban carbon emission data according to claim 1 is characterized in that: In step 4, the method of using reasoning and rethinking in the questioning process is used to improve the accuracy of the large model in identifying carbon emission entity pairs; the specific steps are: First, let the big model generate the semantic similarity, geographical location similarity, industry attribute similarity, description similarity, carbon emission similarity and total similarity of carbon emission entities based on the information submitted to it, and then let the big model judge whether the candidate entity pairs match based on the similarity obtained from the first question and answer; by combining this information to comprehensively evaluate the entities, the accuracy of alignment is further improved, and the heterogeneity and complexity problems in entity alignment are solved; in order to solve the problem that the entity ranges corresponding to carbon emission data are not strictly the same, the carbon emission similarity generated by the big model is used. When the carbon emission similarity is extremely low, the big model will reduce the total similarity of the entity pair during the rethinking process, and is more likely to reject false matches.
4. The spatial entity alignment method for urban carbon emission data according to claim 1 is characterized in that: In step 5, in order to reduce the excessive consumption of computing power of the large model, the 1-10-20 questioning method is used. The specific method is as follows: First, submit the candidate entity pairs with the greatest possibility of matching to the big model, and let the big model determine whether it is confident enough to consider the candidate entity pairs as matched successfully. Then, submit the top 10 candidate entity pairs with the greatest similarity to the big model, and let it determine whether there is an entity pair whose probability of matching successfully is much greater than that of the other candidate entity pairs. Next, submit the top 10 candidate entity pairs with the greatest similarity to the big model, and let it determine whether there is an entity pair whose probability of matching successfully is much greater than that of the other candidate entity pairs. If none of the first 20 candidate pairs are matched successfully, the matching fails.