Data resource theme extraction method and system oriented to integrated energy service
Through the BERTopic theme analysis method, combined with BERT, UMAP and HDBSCAN algorithms, the problem of insufficient research on keywords and elements in comprehensive energy services is solved, and key topics and elements are extracted from text data and identify system boundaries and elements.
Patent Information
- Application Number
- CN202510211115.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, there are few researches on keywords and key elements of integrated energy services, and there is a lack of analytical and research paradigms for scientific methods.
The BERTopic theme analysis method is adopted, combined with the BERT language model, UMAP dimensionality reduction technology, HDBSCAN algorithm and c-TF-IDF algorithm, the topics and keywords of comprehensive energy services are extracted from a large amount of text data, and the importance of the topic is identified through preprocessing, dimensionality reduction, clustering and weight calculation.
It realizes automatic extraction of topics and keywords from a large amount of text data, helping users understand the theme structure in the data, and identify the main physical boundaries and key technical elements of the system.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of integrated energy services, and particularly relates to a method and system for extracting data resource themes for integrated energy services. Background Art
[0002] Currently, research on integrated energy services mainly focuses on the coupling of multiple energy systems in integrated energy services, the improvement of energy efficiency and optimal utilization in integrated energy services, and the business models of integrated energy services. There is less research on the keywords and key elements of integrated energy services.
[0003] In the prior art, the keywords and key elements of integrated energy services mainly come from the literature retrieval method, that is, refining and interpreting the keyword phrases in the publicly available literature. This method lacks a research paradigm of using scientific methods to refine scientific problems and analyze and study them. Summary of the Invention
[0004] To solve the problems existing in the prior art, the present invention discloses a method and system for extracting data resource themes for integrated energy services, aiming to automatically extract themes and their keywords from a large amount of text data, judge the importance of themes, and provide a research paradigm for analysis and research based on scientific methods.
[0005] The present invention discloses a method for extracting data resource themes for integrated energy services, and the steps include:
[0006] Loading data resources, setting custom stop words and custom multi-word terms, adjusting the word segmentation result function, preprocessing the data resources, and obtaining the preprocessed text;
[0007] Loading the pre-trained BERT language model, converting the preprocessed text into a high-dimensional vector representation, using the UMAP dimensionality reduction technique to reduce the dimensionality of the high-dimensional vector, and then using the HDBSCAN algorithm to cluster the reduced-dimensional vector;
[0008] For each subset obtained by clustering, by analyzing the center point and frequently occurring words of the subset, using c-TF-IDF to extract the theme words that can represent the subset, obtaining the weight values of different theme words for each subset, and extracting the themes of the data resources based on the weight values.
[0009] Furthermore, the data resources are publicly available data related to integrated energy services, including data released by statistical agencies or government energy departments, data published by energy enterprises, and data collected by integrated energy service platforms.
[0010] Further, the preprocessing uses a word segmentation result function to segment the data resources, splitting the continuous text sequence to obtain words or phrases containing semantics.
[0011] Further, the BERT language model is used to capture the semantic information of word contexts and generate high-dimensional vectors. Its calculation formula is as follows:
[0012]
[0013] where Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector.
[0014] Further, the UMAP dimensionality reduction technique reduces the dimensionality of the high-dimensional vectors by minimizing the following cost function C:
[0015]
[0016] where w i,j is the weight of the high-dimensional vector in the high-dimensional space, and is the corresponding weight of the high-dimensional vector in the low-dimensional space.
[0017] Further, the use of the HDBSCAN algorithm to cluster the dimensionality-reduced vectors is achieved by calculating the core distance between data points, distinguishing different density regions, and thus generating corresponding groups. The calculation formula for the core distance dc between data points is:
[0018] d c (x) = d(x, x k )
[0019] where d(x, x k ) represents the distance from point x to its k-th nearest neighbor point x k , and k is a parameter in the HDBSCAN algorithm related to the density of clustering.
[0020] Further, the clustering process analyzes to obtain the clustering result according to the set minimum cluster size, the number of samples of the core distance, and the stability affecting density estimation.
[0021] Further, the weight calculation formula of c-TF-IDF is:
[0022]
[0023] where f(t, c) is the frequency of occurrence of word t in topic c, n c is the total number of words in topic c, N is the total number of documents, and DF(t) is the number of documents containing word t.
[0024] The present invention also discloses a data resource theme extraction system for integrated energy services, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the foregoing method.
[0025] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing method are implemented.
[0026] Beneficial effects: By inputting a certain scale of representative cases of integrated energy service practices, the present invention extracts themes and their keywords from a large amount of text data, analyzes the distribution of each theme dataset, and judges the importance of the themes, helping users intuitively understand the theme structure in the data, and then refining the main physical boundaries, device contents, and key technical elements of the system.
[0027] The analysis method used in the present invention can extract themes and keywords from large-scale text data, analyze their distribution, and identify the importance of the themes, thereby helping users understand the theme structure in the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic flowchart of the method of the present invention.
[0029] Figure 2 It is the analysis result of the elements of group A in the embodiment of the present invention;
[0030] Figure 3 It is the analysis result of the elements of group B in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0032] The BERTopic topic analysis method is a combination of BERT and Topic Modeling. The former is a pre-training method that has received attention in recent years and plays an important role in the field of natural language processing for effective communication between humans and computers in natural language. It can learn on a large amount of unlabeled data to capture rich language features (keywords, key elements). The latter is an unsupervised learning method that can discover hidden topic structures from a large amount of data, reveal potential connections and semantic hierarchies between data, and has been playing an increasingly important role in natural language processing tasks such as information retrieval, text classification, and even sentiment analysis in recent years. After the two are combined, the advantages of both can be further exerted, which is beneficial to extracting the required keywords and key elements from a representative number of comprehensive energy service practice cases of a certain scale. It can extract topics and their keywords from a large amount of text data, analyze the distribution of each topic dataset, judge the importance of topics, help users intuitively understand the topic structure in the data, and BERTopic is widely used in fields such as text classification, information retrieval, and sentiment analysis.
[0033] The present invention discloses a method for extracting data resource topics for comprehensive energy services. This method uses BERTopic topic analysis, and the involved process is as Figure 1 shown. Preferably, a specific embodiment includes the following steps:
[0034] Step 1: Load the data resource, set custom stop words, custom multi-word terms, adjust the word segmentation result function, preprocess the data resource, and obtain the preprocessed text.
[0035] The data resource includes data released by statistical agencies or government energy departments, data published by energy enterprises, data collected by comprehensive energy service platforms, etc. The custom stop words will be automatically filtered to save storage space and improve search efficiency. The word segmentation result function is used to split the text data into individual words for keyword extraction operations.
[0036] Furthermore, manually adjust the word segmentation result function and optimize the word segmentation algorithm to adapt to the application scenario and requirements of this patent, and ensure the accuracy and applicability of the word segmentation results. In this embodiment, the computer instructions for manually adjusting the word segmentation result function are as follows.
[0037]
[0038] In this embodiment, the custom multi-word terms include: "energy management service", "electrochemical energy storage", "heating and cooling system", "electric vehicle", "V2G charging", "gas-electricity interconnection".
[0039] The data resources selected in this embodiment are randomly selected from hundreds of representative comprehensive energy service practice cases in the "Typical Cases of China's Top 100 Comprehensive Energy Services" from 2021 to 2024 published by China Electric Power Press, and 69 items are selected. Part of the data resource content is shown in the following table.
[0040] Table 1 Content of Data Resources (Partial)
[0041]
[0042]
[0043]
[0044] Step 2: Load the pre-trained BERT language model, convert the preprocessed text into a high-dimensional vector representation, use the UMAP dimensionality reduction technique to reduce the dimensionality of the high-dimensional vector, and then use the HDBSCAN algorithm to cluster the reduced-dimensional vector.
[0045] Input the data resources into the model described in Step 2, and perform model calculations through four steps: text embedding, dimensionality reduction, clustering, and topic extraction. The specific process is as follows:
[0046] (2.1) Text embedding. Use the pre-trained BERT language model to convert the preprocessed text into a high-dimensional vector representation. BERT is a powerful language model that can capture the semantic information of words in context. Its calculation formula is as follows:
[0047]
[0048] Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k is the dimension of the key vector.
[0049] (2.2) Dimensionality reduction. Use the UMAP (Uniform Manifold Approximation and Projection) dimensionality reduction technique to reduce the dimensionality of the high-dimensional vector for subsequent clustering analysis. Denote the high-dimensional vector representation obtained in (2.1) as the target vector. The dimensionality reduction process of UMAP is achieved by minimizing the following cost function C:
[0050]
[0051] Among them, w i,j is the weight of the target vector in the high-dimensional space, is the corresponding weight of the target vector in the low-dimensional space, E is the vector feature included in the target vector, and is the set from i to j.
[0052] In this embodiment, the computer instructions used in this process are as follows:
[0053] umap_model = UMAP(n_neighbors = 5, n_components = 5,
[0054] min_dist = 0.1, metric = 'cosine')
[0055] Among them, umap_model is the formula function of the dimensionality reduction model; n_neighbors = 5 means that when constructing the UMAP model, for each data point, its 5 adjacent values are considered; n_components = 5 means that the dimensionality of the data after dimensionality reduction by the UMAP model is 5; min_dist = 0.1 means that the minimum distance between embedding points is controlled to be 0.1; metric = 'cosine' means that cosine similarity is used as the metric when calculating the distance between data points.
[0056] (2.3) Clustering. Use the HDBSCAN algorithm to cluster the dimensionality-reduced vectors to reduce data noise.
[0057] Using the HDBSCAN algorithm for clustering is to distinguish different density regions by calculating the core distance between data points, thereby generating corresponding groups. Data points with high density are grouped into the same group to form a clustering result, data points with low density are regarded as noise points, and the noise points are separated from the clustering result, making the clustering result clearer and more accurate. The HDBSCAN algorithm uses the following formula to calculate the core distance d between data points c :
[0058] d c (x) = d(x, x k )
[0059] Among them, d(x, x k ) represents the distance from point x to its k-th nearest neighbor point x k , and k is a parameter in the algorithm, usually related to the density of clustering.
[0060] In this embodiment, the computer instructions used in this process are as follows:
[0061] hdbscan_model = HDBSCAN(min_cluster_size = 15,
[0062] min_samples = 2, metric = 'euclidean', prediction_data = True)
[0063] Among them, hdbscan_model is the formula function of the clustering processing model. min_cluster_size = 15 means that the default value of the minimum cluster size is 15; min_samples = 2 means that the number of samples used to calculate the core distance and the stability affecting density estimation is 2; metric refers to the distance metric method, and the default value is euclidean, that is, Euclidean distance; prediction_data refers to whether to save the prediction data.
[0064] In this embodiment, after clustering by the HDBSCAN algorithm, the grouped A class representing the "power supply side - load side" group of the energy system is obtained, which includes vector features such as "A1 comprehensive energy efficiency service", "A2 cooling, heating and power supply multi - energy service", "A3 distributed clean energy service", etc.; the grouped B class representing the "energy storage side" group of the energy system, which also includes vector features such as "B1 comprehensive energy efficiency service", "B2 cooling, heating and power supply multi - energy service", "B3 distributed clean energy service", etc.
[0065] Step 3: For each group obtained by clustering, by analyzing the center point and frequently occurring words of the group, use the c - TF - IDF algorithm to extract the topic words that can represent the group, obtain the weight values of different topic words in each group, and extract the topic of the data resource based on the weight values.
[0066] For each group, by analyzing its center point and frequently occurring words, use the c - TF - IDF algorithm (a variant of TermFrequency - Inverse Document Frequency) to extract the topic words that can represent the cluster. The weight calculation formula of the c - TF - IDF algorithm is:
[0067]
[0068] Among them, f(t, x) is the frequency of occurrence of the word t in the topic c, n c is the total number of words in the topic c, N is the total number of documents, and DF(t) is the number of documents containing the word t.
[0069] In this embodiment, the computer instructions used in the above process are as follows:
[0070]
[0071]
[0072] In the embodiment of the present invention, after the A - class grouping obtained in step 2 undergoes the topic extraction in step 3, the obtained keywords and their corresponding weight values are as Figure 2As shown. It can be seen from the analysis results that the keywords are ranked by weight as "photovoltaic", "V2G charging", "energy", "heating and cooling", "electrical interconnection" in sequence, indicating that these elements have received extensive attention in 69 cases used as data sources. Similarly, the results obtained after the B-class grouping in Step 2 undergoes the topic extraction in Step 3 are as Figure 3 shown. The top three topic keywords ranked by weight are "energy", "system", and "comprehensive". The above results indicate that these keywords are the elements that a representative number of comprehensive energy service practice cases of a certain scale focus on.
[0073] Therefore, through the analysis results of the BERTopic model, it is possible to determine the physical boundaries and elements of the power source side and load side of the energy system for comprehensive energy services as wind power and photovoltaic, heating and cooling systems, gas-electricity interconnection, and the comprehensive energy service management and control platform, and the physical boundaries and elements of the energy storage side are electrochemical energy storage and electric vehicle V2G respectively.
[0074] A configuration method and system for comprehensive energy services disclosed by the present invention extracts topics and their keywords from a large amount of text data by inputting a representative number of comprehensive energy service practice cases of a certain scale, analyzes the distribution in each topic dataset, judges the topic importance, helps users intuitively understand the topic structure in the data, and further extracts the main physical boundaries, device content, and key technical elements of the system.
Claims
1. A method for extracting data resource themes for integrated energy services, characterized in that, Including: Loading data resources, setting custom stop words and custom multi-word terms, adjusting the word segmentation result function, and preprocessing the data resources to obtain preprocessed text; Loading a pre-trained BERT language model, converting the preprocessed text into a high-dimensional vector representation, using UMAP dimensionality reduction technology to reduce the dimensionality of the high-dimensional vector, and then using the HDBSCAN algorithm to cluster the reduced-dimensional vector; For each group obtained by clustering, by analyzing the center point and frequently occurring words of the group, using c-TF-IDF to extract the topic words that can represent the group, obtaining the weight values of different topic words for each group, and extracting the topic of the data resources based on the weight values.
2. The data resource theme extraction method according to claim 1, wherein The data resources are publicly available data related to integrated energy services.
3. The data resource topic extraction method according to claim 1, wherein The preprocessing is to perform word segmentation on the data resources using the word segmentation result function, and split the continuous text sequence to obtain words or phrases containing semantics.
4. The data resource theme extraction method according to claim 1, characterized in that, The BERT language model is used to capture the semantic information of word contexts and generate high-dimensional vectors. Its calculation formula is as follows: Among them, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d kk is the dimension of the key vector.
5. The data resource topic extraction method according to claim 4, wherein The UMAP dimensionality reduction technology reduces the dimensionality of the high-dimensional vector by minimizing the following cost function C: where w ii,jj is the weight of the high-dimensional vector in the high-dimensional space, and is the corresponding weight of the high-dimensional vector in the low-dimensional space.
6. The data resource theme extraction method according to claim 5, characterized in that The above-mentioned clustering process of the vectors after dimensionality reduction using the HDBSCAN algorithm is to calculate the core distance between data points, distinguish different density regions, and thus generate corresponding groups. The core distance d c between data points is calculated by the following formula: d cc d(x) = d(x, x kk ) Among them, d(x, x kk ) represents the distance from point x to its k-th nearest neighbor point x kk . k is a parameter in the HDBSCAN algorithm and is related to the density of clustering.
7. The data resource topic extraction method according to claim 6, wherein The clustering process analyzes to obtain the clustering result according to the set minimum clustering size, the number of samples of the core distance, and the stability affecting density estimation.
8. The data resource theme extraction method according to claim 1, wherein The weight calculation formula of the c-TF-IDF is: Among them, f(t, c) is the occurrence frequency of word t in topic c, and n cc is the total number of words in topic c, N is the total number of documents, and DF(t) is the number of documents containing word t.
9. A data resource theme extraction system for integrated energy services, comprising a memory, a processor, and a computer program stored on the memory, characterized in that The processor executes the computer program to implement the steps of the method described in claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in claims 1 to 8.