Building hidden carbon emission prediction method based on vector database

Through vector database technology and fine-tuning of large models, the data processing and prediction problems of the building's implicit carbon emission model under complex conditions are solved, high-precision carbon emission prediction is achieved, and the generalization ability and prediction accuracy of the model are improved.

CN120430441APending Publication Date: 2025-08-05HONG KONG UNIV OF SCI & TECH SHENZHEN-HONG KONG COLLABORATIVE INNOVATION INST (FUTIAN SHENZHEN) +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510297937.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing building implicit carbon emission models are difficult to deal with carbon emission-related data under real and complex conditions, especially when the data types are diverse and formats are different, it is impossible to effectively generalize and high-precision prediction.

Method used

The vector database technology is used to vectorize the original carbon emission data and the carbon emission data to be predicted. A feature dictionary and similarity vector library are generated through the pre-trained architectural implicit carbon model. The large model is fine-tuned using the LoRA method, and the data completion and prediction are combined with cosine similarity and L2 similarity.

Benefits of technology

It realizes high-precision prediction of building carbon emissions, solves the problems of backward data processing and insufficient generalization capabilities in traditional methods, and improves prediction accuracy and ability to adapt to complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430441A_ABST
    Figure CN120430441A_ABST
Patent Text Reader

Abstract

The invention discloses a building hidden carbon emission prediction method based on a vector database, relates to the field of databases, and solves the technical problem that an existing building hidden carbon emission model cannot process carbon emission related data under real and complex conditions. The method comprises the following steps: S11, acquiring original carbon emission data and carbon emission data to be predicted; s12, performing vectorization processing on the original carbon emission data and the to-be-predicted carbon emission data, and converting the original carbon emission data and the to-be-predicted carbon emission data into an original carbon emission vector and a to-be-predicted carbon emission vector; and S13, complementing the feature dictionary according to the original carbon emission vector to obtain a complemented feature dictionary. According to the method, high-precision prediction of building carbon emission is realized, a large model technology based on a vector database is innovatively combined with implicit carbon emission prediction, and an advanced and efficient solution is provided for carbon emission prediction in the building field; the problems of backward data processing, poor generalization ability, insufficient prediction precision and the like in a traditional carbon emission prediction method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of databases, and more particularly to a method for predicting building embodied carbon emissions based on a vector database. Background Art

[0002] When predicting embodied carbon emissions in buildings, the particularity of the industry brings the following two difficulties: 1) The collected historical building data is significantly different from that of newly developed buildings, and general models cannot guarantee generalization; 2) The data of interest are of various types and formats, making it difficult for traditional models to effectively process them.

[0003] The existing patent with publication number CN119378750A discloses a method and device for predicting embodied carbon emissions in buildings based on machine learning. The patent selects the main sources of embodied carbon emissions in buildings as model input, uses grey correlation analysis to further screen important factors, uses early machine learning models such as support vector regression, serial structure ANN and decision tree to build a consumption model, combines carbon emission factors to complete carbon emission prediction, and develops a Python program.

[0004] The existing patent with publication number CN118114335A discloses a building carbon emission prediction method and system based on a database platform. The patent first constructs a carbon emission factor library, manages building parameter information at all stages of the building's life cycle, establishes a carbon emission baseline and evaluates the importance of factors, performs simple regression or trains neural networks to obtain a prediction algorithm, and realizes rapid analysis of new projects.

[0005] Patent publication number CN117390457A discloses a combined offline and online transfer learning method and system for predicting building carbon emissions. This patent first collects building information to establish an offline simulation model and dataset. It then uses dynamic time warping metrics to validate the model and train an online prediction model, ultimately predicting building carbon emissions. This method uses dynamic time warping to analyze similarity between edge and cloud data, improving prediction accuracy and reducing the burden of network training.

[0006] Regarding the difficulty of generalizing the model from past data to newly introduced data, previous methods such as transfer learning and similarity analysis often require the new data to belong to a subset or at least a partial subset of the previous dataset, or to meet more stringent prior assumptions, making it difficult to guarantee the predictive performance of new data under complex conditions.

[0007] Furthermore, due to the diverse input types required to predict building embodied carbon emissions, it is difficult to directly train high-performance models using source data. Existing methods are not specifically optimized for carbon emissions data. Related methods, such as support vector regression and neural networks, require consistent initial data and are unable to process carbon emissions data under complex real-world conditions. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to address the shortcomings of the existing technology and provide a method for predicting building embodied carbon emissions based on a vector database to solve the technical problem that the existing building embodied carbon emission model is unable to process carbon emission related data under real complex conditions.

[0009] The method for predicting building embodied carbon emissions based on a vector database of the present invention is as follows:

[0010] S11: Obtaining original carbon emission data and carbon emission data to be predicted;

[0011] S12: performing vectorization processing on the original carbon emission data and the carbon emission data to be predicted, respectively, to convert them into an original carbon emission vector and a carbon emission vector to be predicted;

[0012] S13: Completing the feature dictionary according to the original carbon emission vector to obtain a completed feature dictionary;

[0013] S14: Inputting the original carbon emission vector into a preset carbon emission vector database to retrieve relevant carbon emission content, and generating a similarity vector library based on the relevant carbon emission content and the completed feature dictionary;

[0014] S15: Inputting the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain a final prediction result.

[0015] As a further improvement, in S12, the method of vectorizing the original carbon emission data and the carbon emission data to be predicted is:

[0016] Converting the categories and labels in the raw carbon emission data into raw carbon emission vectors through an encoder of a pre-trained building embodied carbon model;

[0017] The categories and labels in the carbon emission data to be predicted are converted into the carbon emission vector to be predicted through the encoder of the building embodied carbon model.

[0018] Furthermore, the method for generating the pre-trained building embodied carbon model is:

[0019] The preset corpus is input into the open source big model to generate a big model with the original corpus, and the big model with the original corpus is fine-tuned using the LoRA method to generate a big model of building embodied carbon.

[0020] Furthermore, in S13, the method of completing the feature dictionary using the original carbon emission vector is:

[0021] S21: The original carbon emission vector includes features of d dimensions. The features of the d dimensions are searched in the feature dictionary D respectively, and the d' input features that have not been collected are recorded.

[0022] S22: The d' input features that have not been collected are analyzed using the preset prompt words. Vectorize the d' vector features that are not collected

[0023] S23: Use the semantic search function of the preset vector database to search for d' vector features Perform content enhancement to obtain d' completion features

[0024] S24: The above d' completion features d' vector features and d' input features are stored in the feature dictionary D to generate the completed feature dictionary D[d′ i ]=c i .

[0025] Furthermore, in S23, the semantic search function of the preset vector database is used to search for d' vector features. Perform content enhancement to obtain d' completion features The method is,

[0026] All relevant vectors in the vector database are obtained, cosine similarities of all relevant vectors and vector features are calculated, and when the cosine similarity is greater than a preset cosine similarity threshold, all relevant vectors and vector features are integrated to form a complementary feature.

[0027] Furthermore, the expression for calculating the cosine similarity between the correlation vector and the vector feature is:

[0028]

[0029] Where cos(θ) is the cosine similarity, A is the vector feature, B is all related vectors in the vector database, n is the dimension of the vector, and i is a natural number greater than zero.

[0030] Furthermore, in S14, the method for generating the similarity vector library is:

[0031] Obtaining the carbon emission features in the relevant carbon emission content and k completed features in the completed feature dictionary, and setting a k×m-dimensional similarity vector b, where m represents the dimension into which each of the completed features is compressed;

[0032] When the carbon emission feature is identical to the completed feature, the carbon emission feature is stored in the similarity vector b;

[0033] The similarity vectors b are integrated into a similarity vector library.

[0034] Furthermore, in S15, the method of inputting the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain the final prediction result is:

[0035] The completed feature dictionary and similarity vector library are searched for the completed feature that is most similar to the carbon emission vector to be predicted through L2 similarity, and the most similar completed features are recorded in a preset prediction set, which is used as the final prediction result.

[0036] Furthermore, the method of searching for the most similar completed feature to the carbon emission vector to be predicted in the completed feature dictionary and similarity vector library by L2 similarity is as follows:

[0037] The real-time L2 similarity between the supplementary feature in the similarity vector library and the predicted carbon emission vector is calculated, and when the real-time L2 similarity is greater than a preset L2 similarity threshold, the supplementary feature in the similarity vector library is determined to be the most similar supplementary feature.

[0038] Furthermore, the expression for calculating the real-time L2 similarity between the completed feature and the predicted carbon emission vector is:

[0039]

[0040] Among them, s is the real-time L2 similarity, C is the completed feature in the similarity vector library, and D is the carbon emission vector to be predicted. is the Euclidean distance, n is the dimension of the vector, and i is a natural number greater than zero.

[0041] Beneficial effects

[0042] The advantages of the present invention are:

[0043] The present invention converts original carbon emission data and to-be-predicted carbon emission data into original carbon emission vectors and to-be-predicted carbon emission vectors by vectorization processing respectively, completes a feature dictionary according to the original carbon emission vectors to obtain a completed feature dictionary, inputs the carbon emission vectors into a preset carbon emission vector database and the completed feature dictionary to retrieve relevant carbon emission content, binds the relevant carbon emission content with the original carbon emission data to generate a similarity vector library, inputs the to-be-predicted carbon emission vectors into the completed feature dictionary and the similarity vector library to obtain a final prediction result, realizes high-precision prediction of building carbon emissions, innovatively combines large model technology based on vector database with implicit carbon emission prediction, provides an advanced and efficient solution for carbon emission prediction in the construction field, and solves the problems of backward data processing, poor generalization ability and insufficient prediction accuracy in traditional carbon emission prediction methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of the present invention for constructing a large-scale model for predicting building carbon emissions. DETAILED DESCRIPTION

[0045] The present invention will be further described below in conjunction with the embodiments, but this does not constitute any limitation to the present invention. Any limited number of modifications made by anyone within the scope of the claims of the present invention are still within the scope of the claims of the present invention.

[0046] See Figure 1 The present invention provides a method for predicting building embodied carbon emissions based on a vector database, which includes:

[0047] S11: Obtaining original carbon emission data and carbon emission data to be predicted. Both the original carbon emission data and the carbon emission data to be predicted include a large amount of natural language data in the carbon emission input.

[0048] S12: performing vectorization processing on the original carbon emission data and the carbon emission data to be predicted, respectively, to convert them into an original carbon emission vector and a carbon emission vector to be predicted.

[0049] In S12, the method of vectorizing the original carbon emission data and the carbon emission data to be predicted is as follows:

[0050] The categories and labels in the raw carbon emission data are converted into raw carbon emission vectors through the encoder of the pre-trained building embodied carbon model.

[0051] The encoder of the building embodied carbon model is used to convert the categories and labels in the carbon emission data to be predicted into the carbon emission vector to be predicted.

[0052] The method to generate a pre-trained building embodied carbon model is:

[0053] The preset corpus is input into the open source big model to generate a big model with the original corpus. The big model with the original corpus is fine-tuned using the LoRA method to generate a big model of building embodied carbon.

[0054] Specifically, we designed a large-scale model training architecture specifically for carbon emissions data. To provide the necessary expertise, we selected a corpus of relevant documents, including textbooks on architecture, materials, and chemistry, encyclopedia data, engineering technical documents, papers, and patents. Using this data, we fine-tuned the open-source large-scale model, predicting the next token by inputting tokens corresponding to consecutive text, achieving pre-training based on the original corpus.

[0055] Because fine-tuning large models consumes significant GPU memory and time, we utilize the LoRA (Low-Rank Adaptation) method for fine-tuning. Specifically, unoptimized fine-tuning computes and updates the gradients of all parameters in the network's weight matrix, totaling billions to hundreds of billions of parameters. LoRA's key optimization involves decomposing the weight matrix of the original pre-trained model into a low-rank representation. During training, only the low-rank matrix is updated, leaving the original model parameters unchanged. Only the low-rank matrix is trained and updated to adapt the model to the new task.

[0056] Take the weight matrix A of a 1024×1024 fully connected layer as an example, which contains 2 20 Parameters; if LoRA is used to reduce the rank to 16, the model will optimize two small matrices P∈R 1024×16 and Q∈R 16×1024 , a total of 2 15 Parameters are optimized over 30 times compared to the original solution. During the actual update, the model's weight matrix A is replaced with A+PQ for gradient calculation, while the total dimension remains 1024×1024, and the original model structure remains unchanged. After training, the parameters of A are simply replaced with the values of A+PQ. The main advantage of LoRA is that it significantly reduces the number of trainable parameters, memory usage, and computational overhead, while largely maintaining the model's basic capabilities while enabling task-specific optimizations. This perfectly meets our requirements for natural language input processing capabilities.

[0057] S13: Completing the feature dictionary according to the original carbon emission vector to obtain a completed feature dictionary.

[0058] Traditional algorithms typically perform simple normalization on raw data or directly construct decision trees. These methods suffer from significant performance losses when dealing with complex input conditions and high-precision carbon emissions prediction. Furthermore, overfitting the raw data can cause the model to lose its generalization ability, making it ineffective for predicting new input data.

[0059] To solve this problem, we propose to integrate various types of inputs into a unified form. One intuitive solution is to use a complete natural language representation. In this patent, we process the initial data by vectorization and use a vector database to achieve this goal. First of all, high-quality vectorization needs to rely on an embedding model with strong generalization capabilities. Use a large model encoder pre-trained on massive natural language data to perform preliminary processing on inputs in various formats. To ensure the advancement and reliability of the method, we use leading open source large models and their corresponding encoders (such as Qwen 2.5 and Llama 3) to process carbon emission inputs containing d dimensional features as follows.

[0060] Initialization: Maintain the "feature name"-"completion" dictionary D, which is initially an empty set.

[0061] In S13, the method of completing the feature dictionary using the original carbon emission vector is:

[0062] S21: The original carbon emission vector includes features of d dimensions. The features of d dimensions are retrieved in the feature dictionary D respectively, and the d' input features that have not been collected are recorded.

[0063] S22: Use the preset prompt words to collect d' input features that have not been collected The vectorized features not collected are obtained by the above encoders.

[0064] S23: Use the semantic search function of the preset vector database to search for d' vector features Perform content enhancement to obtain d' completion features

[0065] S24: The above d' completion features d' vector features and d' input features are stored in the feature dictionary D to generate the completed feature dictionary D[d′ i ]=c i .

[0066] We call this method the corpus alignment algorithm. In S21, each feature t i There will also be a corresponding eigenvalue v i In S22, prompt words refer to the use of structured text and other methods to improve simple raw inputs and guide the large model to output the desired results. The encoder is the module that the large model uses to segment the input information and convert it into vectors. In S23, open source vector databases (such as pgvector) can store knowledge related to buildings and carbon emissions and necessary expanded knowledge into the database through encoder vectorization.

[0067] All relevant vectors in the vector database are obtained, and the cosine similarity of all relevant vectors and vector features is calculated. When the cosine similarity is greater than a preset cosine similarity threshold, all relevant vectors and vector features are integrated to form a complementary feature.

[0068] The expression for calculating the cosine similarity between the correlation vector and the vector feature is:

[0069]

[0070] Where cos(θ) is the cosine similarity, A is the vector feature, B is all related vectors in the vector database, n is the dimension of the vector, and i is a natural number greater than zero.

[0071] That is, the cosine similarity between the query vector and all vectors in the database is calculated, and the top k results with the highest similarity are returned.

[0072] For example, for the feature t = "building type," we provide a corresponding prompt such as: "Serving accurate prediction of... embodied carbon emissions..., the keyword is... building type..., including... multi-story, high-rise, super-high-rise, etc." Note that our goal is to enable the large model to accurately locate the keyword, rather than directly instructing it to make predictions. This content is fed into the large model, and its encoder output is processed into a vectorized result a = {3.1, 2.2, -5.0, 4.5, 2.9, -4.1, ...}. The semantic search function of the vector database is invoked, using a as input, and the k = 2 most relevant results are selected:

[0073] Characteristics of embodied carbon emissions of different building types: Multi-story buildings (usually 3-6 floors): relatively less material consumption and lower embodied carbon emissions per unit area; high-rise buildings (usually 7-30 floors): require more reinforced concrete and other materials, and have higher structural requirements; super-high-rise buildings (usually more than 30 floors): due to structural complexity and safety requirements, the material consumption per unit area is the largest, and the embodied carbon emission intensity is the highest.

[0074] In addition to the number of floors, differences in structural requirements will lead to changes in the amount of materials such as steel bars and concrete. Different structural types require different construction techniques and equipment.

[0075] The above two texts will be integrated into the completion feature c corresponding to t and stored in the dictionary D[t] = c, that is, each unique feature will correspond to a completion text.

[0076] S14: Input the original carbon emission vector into a preset carbon emission vector database to retrieve relevant carbon emission content, and generate a similarity vector library based on the relevant carbon emission content and the completed feature dictionary.

[0077] In S14, the method for generating the similarity vector library is:

[0078] Obtain i carbon emission features in the relevant carbon emission content and k completed features in the completed feature dictionary, and set a k×m-dimensional similarity vector b, where m represents the dimension into which each completed feature is compressed;

[0079] When there is a carbon emission feature that is the same as the complement feature, the carbon emission feature is stored in the similarity vector b;

[0080] The similarity vectors b are integrated into a similarity vector library.

[0081] We need to compress and mine information from the processed data. Similarly, for a carbon emission input containing d dimensional features and a maintained dictionary D containing k features, we maintain a k×m-dimensional similarity vector b, where m represents the dimension into which each feature is compressed. For each i-th feature t in dictionary D, if it belongs to one of the d features in the input, the encoder performs an embedding on the combined feature t, the completion D[t], and the feature value v, resulting in an m-dimensional embedding stored in the i-th dimension of the similarity vector b; otherwise, the i-th dimension of the similarity vector b is set to 0.

[0082] Through the above processing, we obtain a similarity vector library V that compresses all inputs into k×m dimensions. We call the above algorithm for implementing similarity vectors from input to k×m dimensions the input embedding algorithm.

[0083] S15: Input the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain the final prediction result.

[0084] In S15, the method of inputting the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain the final prediction result is:

[0085] The L2 similarity is used to search for the most similar completed features to the carbon emission vector to be predicted in the completed feature dictionary and similarity vector library, and the most similar completed features are recorded into the preset prediction set, which is used as the final prediction result.

[0086] New carbon emissions data often include new feature categories. For typical methods, processing these features is equivalent to extrapolating historical data. Machine learning algorithms have high uncertainty and low accuracy in this extrapolation task. Fortunately, after the data undergoes natural language processing in the first step, based on our trained foundational model, new feature categories can be included in a wider corpus, forming interpolation of the results, which greatly improves model stability.

[0087] Specifically, for a new carbon emission input h containing d dimensional features, we will use the corpus alignment algorithm in the first part to expand the new input features, which are stored in the new dictionary F. To achieve prediction, we prepare two parts of preprocessed data for each input:

[0088] In the first step, for each feature t and its corresponding eigenvalue v in the new carbon emission input h, we retrieve the corresponding complementary feature c from the dictionary D or dictionary F and record it in the unified feature / eigenvalue / complement set.

[0089] In the second step, the input is encoded into a k×m dimensional vector using the input embedding algorithm. The L2 similarity is used to find the k most similar historical data inputs and carbon emission values in the similarity vector library V. These inputs are enhanced in the first step, recorded in a unified feature / feature value / completion set, and the carbon emission values are recorded separately.

[0090] The completed feature dictionary and similarity vector library are searched for the completed feature that is most similar to the carbon emission vector to be predicted through L2 similarity, and the most similar completed features are recorded in a preset prediction set, which is used as the final prediction result.

[0091] The method for searching the completed feature dictionary and similarity vector library for the most similar completed feature to the carbon emission vector to be predicted by L2 similarity is:

[0092] The real-time L2 similarity between the completed feature in the similarity vector library and the predicted carbon emission vector is calculated. When the real-time L2 similarity is greater than a preset L2 similarity threshold, the completed feature in the similarity vector library is determined to be the most similar completed feature.

[0093] We write formatted prompt words for the data obtained in the first and second steps. We first explain each feature and its completion, then introduce the value of each feature and carbon emission value of historical data sample by sample, and finally introduce the value of each newly input feature and ask the large model about its carbon emission value. We use the fine-tuned carbon emission large model to achieve prediction.

[0094] Among them, through the first step of feature completion, we solved the problem of difficulty in aligning multiple data sources, aligned the data in natural language, and used large models to achieve effective data processing; through the second step of similarity retrieval, we found the most similar examples to the new input, namely few-shot, to help the large model more accurately predict the embodied carbon emissions of buildings.

[0095] The expression for calculating the real-time L2 similarity between the completed features and the predicted carbon emission vector is:

[0096]

[0097] Among them, s is the real-time L2 similarity, C is the completed feature in the similarity vector library, and D is the carbon emission vector to be predicted. is the Euclidean distance, n is the dimension of the vector, and i is a natural number greater than zero.

[0098] Due to the complex sources of embodied carbon emissions in buildings, simple single-source testing is insufficient. Technical approaches rely on multi-factor analysis across the entire production process, a long-standing challenge. Our proposed model addresses the challenges of traditional carbon emission prediction methods, including backward data processing, poor generalization, and insufficient prediction accuracy.

[0099] The above is only a preferred embodiment of the present invention. It should be pointed out that for those skilled in the art, several modifications and improvements can be made without departing from the structure of the present invention. These modifications and improvements will not affect the effect of the implementation of the present invention and the practicality of the patent.

Claims

1. A method for predicting building embodied carbon emissions based on a vector database, characterized by: The method is, S11: Obtaining original carbon emission data and carbon emission data to be predicted; S12: performing vectorization processing on the original carbon emission data and the carbon emission data to be predicted, respectively, to convert them into an original carbon emission vector and a carbon emission vector to be predicted; S13: Completing the feature dictionary according to the original carbon emission vector to obtain a completed feature dictionary; S14: Inputting the original carbon emission vector into a preset carbon emission vector database to retrieve relevant carbon emission content, and generating a similarity vector library based on the relevant carbon emission content and the completed feature dictionary; S15: Inputting the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain a final prediction result.

2. The method for predicting building embodied carbon emissions based on a vector database according to claim 1 is characterized in that: In S12, the method of vectorizing the original carbon emission data and the carbon emission data to be predicted is as follows: Converting the categories and labels in the raw carbon emission data into raw carbon emission vectors through an encoder of a pre-trained building embodied carbon model; The categories and labels in the carbon emission data to be predicted are converted into the carbon emission vector to be predicted through the encoder of the building embodied carbon model.

3. The method for predicting building embodied carbon emissions based on a vector database according to claim 2 is characterized in that: The method for generating the pre-trained building embodied carbon model is: The preset corpus is input into the open source big model to generate a big model with the original corpus, and the big model with the original corpus is fine-tuned using the LoRA method to generate a big model of building embodied carbon.

4. The method for predicting building embodied carbon emissions based on a vector database according to claim 1 is characterized in that: In S13, the method of completing the feature dictionary using the original carbon emission vector is: S21: The original carbon emission vector includes features of d dimensions. The features of the d dimensions are searched in the feature dictionary D respectively, and the d' input features that have not been collected are recorded. S22: The d' input features that have not been collected are analyzed using the preset prompt words. Vectorize the d' vector features that have not been collected S23: Use the semantic search function of the preset vector database to search for d' vector features Perform content enhancement to obtain d' completion features S24: The above d' completion features d' vector features and d' input features are stored in the feature dictionary D to generate the completed feature dictionary D[d′ i ]=c i .

5. The method for predicting building embodied carbon emissions based on a vector database according to claim 4 is characterized in that: In S23, the semantic search function of the preset vector database is used to search for d' vector features. Perform content enhancement to obtain d' completion features The method is, All relevant vectors in the vector database are obtained, cosine similarities of all relevant vectors and vector features are calculated, and when the cosine similarity is greater than a preset cosine similarity threshold, all relevant vectors and vector features are integrated to form a complementary feature.

6. The method for predicting building embodied carbon emissions based on a vector database according to claim 5 is characterized in that: The expression for calculating the cosine similarity between the correlation vector and the vector feature is: Where cos(θ) is the cosine similarity, A is the vector feature, B is all related vectors in the vector database, n is the dimension of the vector, and i is a natural number greater than zero.

7. The method for predicting building embodied carbon emissions based on a vector database according to claim 4 is characterized in that: In S14, the method for generating the similarity vector library is: Obtaining the carbon emission features in the relevant carbon emission content and k completed features in the completed feature dictionary, and setting a k×m-dimensional similarity vector b, where m represents the dimension into which each of the completed features is compressed; When the carbon emission feature is identical to the completed feature, the carbon emission feature is stored in the similarity vector b; The similarity vectors b are integrated into a similarity vector library.

8. The method for predicting building embodied carbon emissions based on a vector database according to claim 4 is characterized in that: In S15, the method of inputting the carbon emission vector to be predicted into the completed feature dictionary and similarity vector library to obtain the final prediction result is: The completed feature dictionary and similarity vector library are searched for the completed feature that is most similar to the carbon emission vector to be predicted through L2 similarity, the most similar completed feature is recorded into a preset prediction set, and the prediction set is used as the final prediction result.

9. The method for predicting building embodied carbon emissions based on a vector database according to claim 8 is characterized in that: The method for searching the completed feature dictionary and similarity vector library for the most similar completed feature to the carbon emission vector to be predicted by L2 similarity is: The real-time L2 similarity between the supplementary feature in the similarity vector library and the predicted carbon emission vector is calculated, and when the real-time L2 similarity is greater than a preset L2 similarity threshold, the supplementary feature in the similarity vector library is determined to be the most similar supplementary feature.

10. The method for predicting building embodied carbon emissions based on a vector database according to claim 9, characterized in that: The expression for calculating the real-time L2 similarity between the completed feature and the predicted carbon emission vector is: Among them, s is the real-time L2 similarity, C is the completed feature in the similarity vector library, and D is the carbon emission vector to be predicted. is the Euclidean distance, n is the dimension of the vector, and i is a natural number greater than zero.

Citation Information

Patent Citations

  • Offline and online combined transfer learning building carbon emission prediction method and system

    CN117390457A

  • Building carbon emission prediction method and system based on database platform

    CN118114335A

  • Building hidden carbon emission prediction method and device based on machine learning

    CN119378750A