Building design short text-oriented topic extraction method and system, terminal and storage medium

By performing sentence-level semantic embedding and clustering optimization on architectural design short texts, the problems of poor adaptability and weak structure in existing technologies are solved, high-precision topic extraction and structured output are achieved, and information retrieval efficiency is improved.

CN120654705AActive Publication Date: 2025-09-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511149705.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-09-16
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing technologies have problems in topic extraction from architectural design short texts, such as poor adaptability, weak structure, and low accuracy. In particular, they are insufficient in processing sentence-level semantic expressions, high redundancy of clustering results, and difficulty in extracting clear strategy labels and structured output.

Method used

A sentence-level semantic embedding and clustering optimization method is adopted, including data preprocessing, sentence vector embedding, dimensionality reduction and clustering, and topic optimization processing. By obtaining a dataset of architectural design short texts for preprocessing, a sentence vector set is constructed and dimensionality reduction and clustering are performed. Combined with topic similarity calculation and keyword extraction, the topic clustering results are optimized to improve the degree of structure.

Benefits of technology

It enhances the adaptability and accuracy of topic extraction, improves the clustering accuracy and structured output of sentence-level semantic expression, reduces the redundancy of clustering results, and improves information retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654705A_ABST
    Figure CN120654705A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, and discloses a building design short text oriented topic extraction method and system, a terminal and a storage medium, and the method comprises the steps: obtaining a building design short text data set, and carrying out the preprocessing of the building design short text data set, and obtaining a structured text set; performing sentence vector embedding construction according to the structured text set to obtain a sentence vector set, performing dimension reduction processing on the sentence vector set to obtain a low-dimension sentence vector set, and performing clustering processing on the low-dimension sentence vector set to obtain an initial topic clustering result; and performing optimization processing on the initial topic clustering result to obtain a topic set, and performing structured processing on the topic set to obtain a target topic set. According to the method, sentence-level semantic embedding and clustering optimization processing are performed on the building design short text, so that adaptability of topic extraction is enhanced, and structuring and accuracy of topic extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular to a method, system, terminal and computer-readable storage medium for extracting topics from architectural design short texts. Background Art

[0002] Architectural design texts, especially in scenarios like design weekly reports and review minutes, often exhibit the following characteristics: short text units (i.e., each document typically contains only 200-500 words), semantic fragmentation (i.e., the text is mostly a record of key points or a transcription of oral expression, lacking contextual coherence), dense architectural terminology (i.e., the frequent presence of professional terms and complex spatial expressions, such as "modular teaching units"), strong subjectivity (i.e., the wording often carries intentional judgments or design strategy expressions), and sentence-level core information (i.e., most effective strategy information is concentrated in a single sentence, suitable for sentence-level clustering analysis). Due to these characteristics, traditional topic extraction models for long texts are significantly under-adapted to this scenario, resulting in overly coarse output topic granularity, severe semantic overlap, and poorly structured results.

[0003] Regarding automated clustering extraction of architectural design short texts (i.e., topic extraction of architectural design short texts, which automatically identifies and extracts text related to specific topics or fields from a large amount of text and is commonly used in fields such as data mining, information retrieval, and natural language processing), common technical solutions include methods based on document-level topic models, methods based on text clustering and bag-of-words models, and manual annotation and rule extraction. However, existing technologies for topic identification in architectural short text corpora generally suffer from granularity mismatch (i.e., most methods assume document-level modeling units and cannot handle sentence-level semantic expressions), insufficient purity (i.e., high redundancy in clustering results, making it difficult to extract clear strategic labels), weak structuredness (i.e., lack of implementable structured outputs such as keyword reconstruction, topic numbering, and sentence indexing), and poor adaptability (i.e., failure to incorporate architectural corpus characteristics (e.g., term density, semantic reusability, etc.).

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a topic extraction method, system, terminal and storage medium for architectural design short texts, aiming to solve the problems of poor adaptability, weak structure and low accuracy in the existing technology for topic extraction for architectural design short texts.

[0006] To achieve the above object, the present invention provides a topic extraction method for architectural design short texts, the topic extraction method for architectural design short texts comprising the following steps: Acquire an architectural design short text dataset, and preprocess the architectural design short text dataset to obtain a structured text set; Perform sentence embedding construction based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The initial topic clustering result is optimized to obtain a topic set, and the topic set is structured to obtain a target topic set.

[0007] Optionally, the topic extraction method for architectural design short texts, wherein the step of obtaining an architectural design short text dataset and preprocessing the architectural design short text dataset to obtain a structured text set, specifically includes: Acquire an architectural design short text dataset, perform sentence extraction on the architectural design short text dataset, and obtain a sentence text corpus set; Performing text denoising on basic characters of all sentence texts in the sentence text corpus to obtain a denoised sentence text set, and performing word segmentation processing on the denoised sentence text set to obtain a word segmentation result; Eliminate stop words from the segmentation result according to the stop word list to obtain an initial segmentation result, and filter high-frequency words from the initial segmentation result according to the high-frequency term filter table to obtain a target segmentation result; The target word segmentation results are subjected to synonym normalization processing to obtain a structured text set.

[0008] Optionally, the topic extraction method for architectural design short texts, wherein the sentence vector embedding is performed based on the structured text set to obtain a sentence vector set, specifically comprising: Obtaining a word vector model, establishing a mapping index between vocabulary and vectors based on the word vector model, and performing vector retrieval on the structured text set based on the mapping index to obtain multiple valid word vectors; All the valid word vectors are averaged element-wise to obtain a plurality of sentence embedding vectors, and a sentence vector set is obtained based on all the sentence embedding vectors.

[0009] Optionally, the topic extraction method for architectural design short texts, wherein the dimensionality reduction processing is performed on the sentence vector set to obtain a low-dimensional sentence vector set, and the low-dimensional sentence vector set is clustered to obtain an initial topic clustering result, specifically includes: Setting dimensionality reduction parameters, and performing dimensionality reduction processing on the sentence vector set according to the dimensionality reduction parameters to obtain a low-dimensional sentence vector set, wherein the dimensionality reduction parameters include the number of neighbors, the output dimension, the first distance metric, and the minimum distance; Clustering parameters are set, and the low-dimensional sentence vector set is clustered according to the clustering parameters to obtain an initial topic clustering result, wherein the clustering parameters include a minimum cluster size, a second distance measurement method, and a clustering mode.

[0010] Optionally, the topic extraction method for architectural design short texts, wherein the optimizing process of the initial topic clustering result to obtain a topic set, specifically includes: Performing topic similarity calculation on any two initial topics of the initial topic clustering result to obtain multiple first similarity calculation results; Performing keyword extraction on all the initial topics to obtain multiple semantic identifier sets, and performing word meaning similarity calculation on any two of the semantic identifier sets to obtain multiple second similarity calculation results; Comparing all the first similarity calculation results with a first preset threshold, and comparing all the second similarity calculation results with a second preset threshold; If there are two initial topics whose first similarity calculation results are greater than or equal to the first preset threshold, and whose second similarity calculation results are greater than or equal to the second preset threshold, then aggregate the two initial topics to obtain an aggregated result; The topic numbers are updated according to the aggregation results to obtain multiple updated topics, and a topic set is obtained based on all the updated topics.

[0011] Optionally, the topic extraction method for architectural design short texts, wherein the keyword extraction of all the initial topics to obtain multiple semantic identifier sets specifically includes: Obtaining the topic sentence texts belonging to each of the initial topics, and constructing subsets of all the topic sentence texts of each of the initial topics to obtain multiple topic corpus subsets; A word frequency-inverse document frequency calculation is performed on each word in each of the subject corpus subsets to obtain multiple calculation results, a keyword list for each of the initial topics is obtained based on all the calculation results, and a semantic identification set for each initial topic is obtained based on all the keyword lists.

[0012] Optionally, in the topic extraction method for architectural design short texts, the term frequency-inverse document frequency calculation is performed on each term in each of the topic corpus subsets, specifically: ; in, For words, For the initial theme, is the term frequency-inverse document frequency of the term in the initial topic, is the frequency of the word appearing in the initial topic, is the total number of initial topics, is the number of initial topics containing the word, is the inverse document frequency of the term.

[0013] Optionally, in the method for extracting topics from architectural design short texts, the topic extraction system for architectural design short texts includes: A data preprocessing module is used to obtain an architectural design short text dataset, preprocess the architectural design short text dataset, and obtain a structured text set; A text processing module is used to construct sentence vector embeddings based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The topic optimization module is used to optimize the initial topic clustering result to obtain a topic set, and perform structured processing on the topic set to obtain a target topic set.

[0014] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a topic extraction program for architectural design short texts stored on the memory and runnable on the processor, and when the topic extraction program for architectural design short texts is executed by the processor, the steps of the topic extraction method for architectural design short texts as described above are implemented.

[0015] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a topic extraction program for architectural design short texts, and when the topic extraction program for architectural design short texts is executed by a processor, the steps of the topic extraction method for architectural design short texts as described above are implemented.

[0016] In the present invention, a dataset of architectural design short texts is obtained, and the dataset is preprocessed to obtain a structured text set; sentence vector embedding is constructed based on the structured text set to obtain a sentence vector set, dimensionality reduction is performed on the sentence vector set to obtain a low-dimensional sentence vector set, and the low-dimensional sentence vector set is clustered to obtain an initial topic clustering result; the initial topic clustering result is optimized to obtain a topic set, and the topic set is structured to obtain a target topic set. By performing sentence-level semantic embedding and clustering optimization on architectural design short texts, the present invention not only enhances the adaptability of topic extraction, but also improves the structuring and accuracy of topic extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of a preferred embodiment of the method for extracting topics from architectural design short texts of the present invention; Figure 2 Schematic diagram of the terminal architecture of the subject extraction method for architectural design short texts of the present invention; Figure 3 This is a schematic diagram of the overall process of the topic extraction method for architectural design short texts of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the subject extraction system for architectural design short texts of the present invention; Figure 5 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] It should be noted that if the embodiments of the present invention involve directional indications (such as up, down, left, right, front, back, etc.), such directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0020] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features specified as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0021] The topic extraction method for architectural design short texts described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the topic extraction method for architectural design short texts includes the following steps: Step S10: Acquire an architectural design short text dataset, and preprocess the architectural design short text dataset to obtain a structured text set.

[0022] Specifically, in an embodiment of the present invention, in order to solve the problems in the prior art of topic extraction of architectural short text corpus such as granularity mismatch (i.e., most methods assume document-level modeling units and cannot process sentence-level semantic expressions), insufficient purity (i.e., high redundancy rate of clustering results, difficulty in extracting clear strategy labels), weak structure (i.e., lack of keyword reconstruction, topic numbering and sentence indexing and other feasible structured outputs), and poor adaptability (i.e., failure to combine architectural corpus characteristics (e.g., terminology density, semantic reusability, etc.), a topic extraction method for architectural design short texts is proposed, and the corresponding terminal architecture is as follows: Figure 2 As shown, it includes a data acquisition module, a preprocessing module, a sentence vector embedding module, a dimensionality reduction-clustering module, a topic aggregation module and a result output module, wherein the data acquisition module is used to collect short text data sets related to architectural design and initialize the auxiliary resources required by the system (for example, a stop word list, a high-frequency term filter table and a word vector model, etc.); the preprocessing module is used to perform text denoising, word segmentation, term filtering and synonym normalization operations on the input sentence text corpus set; the sentence vector embedding module is used to convert each cleaned sentence text into a low-dimensional dense semantic vector representation; the dimensionality reduction-clustering module is used to compress the sentence vector set into a clusterable low-dimensional semantic representation; the topic aggregation module is used to further optimize the initial topic clustering result; and the result output module is used to export the final topic clustering result in a structured form.

[0023] The specific processing process is as follows Figure 3As shown, the data acquisition module acquires data files sequentially imported by the user through a graphical user interface (GUI), including an architectural design short text dataset, a stop word vocabulary, a high-frequency term filter table, and a word vector model. The architectural design short text dataset contains text materials such as architect design descriptions, architectural reviews, and interview transcripts, and supports common formats such as .txt, .csv, and .xlsx. The stop words in the stop word vocabulary are common words that are ignored or deleted during text processing. These words are typically frequently occurring function words or meaningless words (e.g., prepositions, conjunctions, and pronouns). Stop words generally do not contribute significantly to text meaning analysis and consume a large amount of storage space and computing resources. Therefore, in text processing tasks (e.g., information retrieval), a set of stop words is often predefined and removed from the text during processing. The high-frequency term filter table is a list of high-frequency words statistically derived from an architectural design text corpus. It removes terms with excessive frequency but no semantic distinction (e.g., "project" and "building area") to improve subsequent clustering accuracy. The word vector model (e.g., a model trained on news corpus) can provide word-level semantic embedding capabilities for subsequent generation of sentence-level vector representations. After acquiring the architectural design short text dataset, the data acquisition module needs to extract sentences from the architectural design short text dataset. By performing sentence extraction on the architectural design short text dataset, a sentence text corpus set is obtained. Specifically, sentences are automatically extracted from each text in the architectural design short text dataset based on commonly used sentence segmentation symbols in Chinese (e.g., period, question mark, exclamation mark, etc.), and the names of the documents to which they belong are recorded. This is to facilitate the subsequent topic classification based on the sentences, indexing the documents containing the sentences, and organizing the different topic information involved in the documents. All sentences are then unified into a sentence text corpus set, represented by docs[], which serves as the basic unit for subsequent semantic modeling.

[0024] Afterwards, the data acquisition module inputs the sentence input corpus set into the preprocessing module, which performs text denoising, word segmentation, term filtering, and synonym normalization operations. Specifically, for text denoising, the basic characters of all sentence texts in the sentence text corpus set are subjected to text denoising to obtain a denoised sentence text set. That is, the basic characters of each sentence in the sentence text corpus set are cleaned, and English characters, Arabic numerals, and special symbols in the sentence are removed to eliminate the interference of non-Chinese information on word segmentation and vector modeling. This process follows a unified regularization rule to ensure processing consistency and controllability. For word segmentation processing, the denoised sentence text set is subjected to word segmentation processing to obtain a word segmentation result. That is, a word segmentation tool based on a hybrid mechanism of statistics and rules is used to segment each Chinese sentence in the denoised sentence text set into the smallest word unit. At this stage, the main parts of speech (for example, nouns, verbs, adjectives, directional words, etc.) are retained, and parts of speech with less impact on semantic recognition (for example, auxiliary words, conjunctions, etc.) are filtered out. Term filtering includes stop word removal and high-frequency word removal. Stop word removal refers to removing stop words from the word segmentation result according to a stop word list to obtain an initial word segmentation result, that is, matching and removing common function words appearing in the stop word list in the word segmentation result; high-frequency word removal refers to filtering high-frequency words from the initial word segmentation result according to a high-frequency term filter table to obtain a target word segmentation result, that is, combining the high-frequency term filter table to further remove terms that appear very frequently in architectural texts but have no distinguishing meaning (for example, projects, plans, etc.). Term filtering helps to compress the word vector space and improve the semantic purity of subsequent clustering. For synonym normalization processing, the target word segmentation results are subjected to synonym normalization processing, that is, the semantically similar words (for example, teaching building and school building, noise and noise, etc.) in the target word segmentation results are standardized and merged through the constructed synonym mapping table, and the synonymous expressions are automatically replaced with standardized terms, thereby improving the aggregation ability of similar strategy expressions in clustering; after the processing is completed, a structured cleaned text set (that is, a structured text set) is obtained, which is represented by cleaned_docs[] and serves as the input of the subsequent sentence vector embedding module. It has good semantic consistency and corpus cleanliness, which can effectively improve the clustering performance and interpretability of topic identification.

[0025] Step S20: construct sentence vector embedding based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and cluster the low-dimensional sentence vector set to obtain an initial topic clustering result.

[0026] Specifically, after receiving the structured text set, the sentence vector embedding module loads the word vector model trained based on Chinese news corpus by default. This word vector model covers more than 5 million commonly used terms, and each term corresponds to a 300-dimensional dense vector. After the model is loaded, a mapping index between vocabulary and vectors is established to support fast word vector retrieval and splicing operations. The process of establishing a mapping index between vocabulary and vectors is as follows: after the word vector model is loaded, all terms in the word vector model are pre-read, and a unique vector number index is assigned to each term. At the same time, the terms and their corresponding semantic vectors are organized into a key-value pair structure (i.e., "word→vector" form), and this key-value pair structure is stored in memory, allowing any subsequent term to directly find its corresponding vector representation in constant time, avoiding repeated calculations or disk read delays. Afterwards, for each sentence in the structured text set, the vector representation of each word in the sentence in the word vector model is first retrieved to obtain multiple valid word vectors, wherein the valid word vector is the vector corresponding to the valid word, and the valid word means that if a certain word has a corresponding vector representation in the word vector model, it is considered to be a valid word. Subsequently, only these valid word vectors are subjected to element-level averaging operations, the purpose of which is to avoid the interference of low-frequency words, special characters and other non-vector representation contents on the embedding results, thereby improving semantic stability and clustering accuracy, and finally obtaining multiple 300-dimensional sentence embedding vectors, and each 300-dimensional sentence embedding vector is used to reflect the semantic features of the corresponding sentence as a whole. The present invention retains the semantic differences of local words in the embedding calculation process, and reduces noise interference through vector averaging operations. It is suitable for processing short sentence corpus with simple structure and sparse vocabulary. For example, if the target platform has GPU capabilities, it also supports switching to a higher-performance deep semantic model to further improve the embedding accuracy and enhance the semantic separability between sentences. It is suitable for deployment scenarios with high requirements for topic boundary recognition accuracy.

[0027] Afterwards, we get the sentence vector set based on all the sentence embedding vectors and use Indicates that, is the sentence vector set, is the set of real numbers, is the number of sentences, is the embedding dimension, =300); the sentence vector set output by the sentence vector embedding module will serve as the direct input of the dimensionality reduction-clustering module to support the subsequent semantic structure recognition and topic label merging process.

[0028] After the dimensionality reduction-clustering module receives the sentence vector set, considering that the dimension of the sentence embedding vector in the sentence vector set is high (i.e., 300 dimensions), direct clustering may cause "dimensionality disaster" and degrade clustering performance. Therefore, in this embodiment of the present invention, the UMAP (Uniform Manifold Approximation and Projection) algorithm is used as a dimensionality reduction tool to map the sentence vector set to The low-dimensional manifold space is used to preserve local semantic adjacency. The dimensionality reduction parameters in the UMAP dimensionality reduction process are set as follows: number of neighbors n_neighbors=15, output dimension n_components=5, distance metric as cosine distance and minimum distance min_dist=0. This setting can effectively enhance the clustering of semantically similar sentences in the low-dimensional space and provide a good structural foundation for subsequent clustering. After the dimensionality reduction is completed, the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm is called to perform unsupervised topic clustering in the embedding space. Based on sample point density estimation and hierarchical tree pruning strategy, it can automatically identify cluster structures and filter out noisy sentences. The clustering parameters of the HDBSCAN aggregation process are set as follows: minimum cluster size , the distance metric is Euclidean distance and the clustering mode is set to leaf node priority pruning to improve clustering granularity; the low-dimensional sentence vector set is clustered according to the clustering parameters to obtain the initial topic clustering results, and the intermediate outputs returned during the clustering process include a topic label array (represented by topics[]) and a hierarchical relationship tree (represented by hierarchy_tree), wherein the topic label array is the topic number assigned to each sentence, and the hierarchical relationship tree is a display of the merging paths between topics for visualization and structural tracing. The dimensionality reduction-clustering module can effectively divide a large number of architectural strategy statements into multiple semantically well-aggregated topic clusters, providing basic support for subsequent topic optimization and semantic naming. Through dynamic clustering parameters and density adaptive mechanisms, the system can adapt to input corpora of different sizes and complexities, significantly reducing the probability of under- and over-aggregation.

[0029] Step S30: Optimize the initial topic clustering result to obtain a topic set, and structure the topic set to obtain a target topic set.

[0030] Specifically, after the dimension reduction-clustering module inputs the initial topic clustering result into the topic aggregation module, the topic aggregation module optimizes the initial topic clustering result to eliminate semantic redundancy and improve the purity and interpretability of the topic structure. The specific optimization process is to calculate the topic similarity of any two initial topics of the initial topic clustering result to obtain multiple first similarity calculation results, for example, for any two initial topics (for example, topics and themes ) is calculated using the cosine similarity (i.e. topic similarity) of the topic, and the corresponding calculation formula is: ; in, For the theme and themes The cosine similarity of For the theme The representative vector of For the theme The representative vector of is a vector is the Euclidean norm (i.e., modulus length), is a vector is the Euclidean norm of, and All are the initial number of topics; The value range of is [0, 1], and the closer it is to 1, the more similar the two topics are.

[0031] Afterwards, for each initial topic, it is necessary to extract keywords from all sentences within it. Specifically, the topic sentence texts of each initial topic are obtained, and all topic sentence texts of each initial topic are subset-constructed to obtain multiple topic corpus subsets; each word in each topic corpus subset is calculated by word frequency-inverse document frequency to obtain multiple calculation results. The corresponding calculation formula is: ; in, For words, For the initial theme, is the term frequency-inverse document frequency of the term in the initial topic, is the frequency of the word appearing in the initial topic, is the total number of initial topics, is the number of initial topics containing the word, is the inverse document frequency of the word; a keyword list of each of the initial topics is obtained based on all the calculation results, and a semantic identification set of each initial topic is obtained based on all the keyword lists, that is, a preset number (for example, the first 10) of keywords in each keyword list are selected as the semantic identification set of each initial topic for subsequent semantic overlap analysis. The purpose of setting the number of keywords to 10 is that if the number is too small (for example, 3 to 5), it may not cover the complex semantic features in the architectural strategy statement; if the number is too large (for example, 20), high-frequency and low-weight terms will be introduced, affecting the judgment of the semantic overlap between topics; and according to experiments, selecting 10 keywords can avoid redundancy while ensuring semantic coverage, and has a positive effect on the stability of the calculation results of the keyword overlap rate in the subsequent aggregation module. It can also be adjusted to any value in the range of 8 to 15 according to the characteristics of the industry corpus.

[0032] Before aggregating the initial topics, it is necessary to perform a two-factor merging rule determination, including semantic similarity determination and keyword overlap rate determination; for semantic similarity determination, all the first similarity calculation results are compared with the first preset threshold (for example, 0.85); for keyword overlap rate determination, all the second similarity calculation results are compared with the second preset threshold (for example, 0.4); if there are two initial topics whose first similarity calculation results are greater than or equal to the first preset threshold, and whose second similarity calculation results are greater than or equal to the second preset threshold, then the two initial topics are aggregated to obtain an aggregated result. For example, for any two initial topics (for example, topics and themes ), you need to judge the topic in turn and themes Whether the semantic similarity judgment and keyword overlap rate judgment are satisfied at the same time, that is, Is it greater than or equal to 0.85, and extract the theme separately Keyword list and themes Keyword list ,calculate and Jaccard similarity (i.e. and Intersection and and and the ratio of the union of the two sets), and determine whether the Jaccard similarity is greater than or equal to 0.4; if (considered as topic semantics close), and the Jaccard similarity is greater than or equal to 0.4 (there is significant semantic overlap), then the topic is determined to be and themes If there is a high degree of semantic redundancy, a topic merge operation is automatically triggered, generating a new topic number. After the topic merge operation is executed, the mapping relationship between the topic numbers before and after the merge is retained to form a traceable topic traceability table, that is, the mapping relationship between the original topic number and the merged topic is recorded for subsequent traceability verification and manual proofreading. If the first similarity calculation result of two initial topics is less than the first preset threshold, or the second similarity calculation result of two initial topics is less than the second preset threshold, the original topic numbers of the two initial topics are retained and the topic merge operation is not performed to ensure that the semantic boundary is preserved, thereby avoiding the loss of fine-grained policy information due to excessive merging.

[0033] Afterwards, the topic number is updated according to the aggregation result to obtain multiple updated topics. A topic set (represented by topics_final[]) is obtained based on all the updated topics, and the topic set is input into the result output module. The result output module performs structured processing on the topic set to obtain the target topic set, wherein the structured processing includes topic information table processing, topic hierarchy visualization processing and model persistence processing; for topic information table processing, the topic set is automatically generated into an Excel spreadsheet file, which includes the number of each topic, keyword list, number of sentences contained in the topic, representative sentences, etc.; for topic hierarchy visualization processing, the topic hierarchical relationship (for example, based on a tree structure) and topic similarity matrix constructed in the topic set are output as HTML files, which support interactive viewing on the browser side and can be used for review reports or strategy map construction; for model persistence processing, the clustering model, parameter settings and preprocessed corpus of the topic set are saved together as a binary model file, which supports subsequent reproduction, updating of clustering results or iterative training. The result output module ensures the transparency, operability and cross-platform compatibility of the full-link processing results.

[0034] In addition, although the present invention is based on short architectural design texts, the strategy of "sentence-level vector embedding + dynamic clustering + two-factor topic aggregation" is universal. After adapting the synonym vocabulary and corpus preprocessing rules, it can be applied to educational planning texts (for example, classroom observation records, teaching space feedback), urban governance documents (for example, block renovation opinions) and medical case key point clustering (for example, main symptom aggregation, rehabilitation suggestion classification).

[0035] Furthermore, if Figure 4 As shown, based on the above-mentioned topic extraction method for architectural design short texts, the present invention also provides a topic extraction system for architectural design short texts, wherein the topic extraction system for architectural design short texts includes: The data preprocessing module 51 is used to obtain an architectural design short text dataset and preprocess the architectural design short text dataset to obtain a structured text set; A text processing module 52 is configured to construct sentence vector embeddings based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The topic optimization module 53 is used to optimize the initial topic clustering result to obtain a topic set, and perform structured processing on the topic set to obtain a target topic set.

[0036] Furthermore, if Figure 5 As shown, based on the above-mentioned topic extraction method for architectural design short texts, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 5 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0037] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a topic extraction program 40 for architectural design short texts is stored on the memory 20. The topic extraction program 40 for architectural design short texts can be executed by the processor 10, thereby implementing the topic extraction method for architectural design short texts in the present application.

[0038] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program codes or process data stored in the memory 20, such as executing the topic extraction method for architectural design short texts.

[0039] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch screen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0040] In one embodiment, when the processor 10 executes the topic extraction program 40 for architectural design short texts in the memory 20, the following steps are implemented: Acquire an architectural design short text dataset, and preprocess the architectural design short text dataset to obtain a structured text set; Perform sentence embedding construction based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The initial topic clustering result is optimized to obtain a topic set, and the topic set is structured to obtain a target topic set.

[0041] The step of obtaining an architectural design short text dataset and preprocessing the architectural design short text dataset to obtain a structured text set specifically includes: Acquire an architectural design short text dataset, perform sentence extraction on the architectural design short text dataset, and obtain a sentence text corpus set; Performing text denoising on basic characters of all sentence texts in the sentence text corpus to obtain a denoised sentence text set, and performing word segmentation processing on the denoised sentence text set to obtain a word segmentation result; Eliminate stop words from the segmentation result according to the stop word list to obtain an initial segmentation result, and filter high-frequency words from the initial segmentation result according to the high-frequency term filter table to obtain a target segmentation result; The target word segmentation results are subjected to synonym normalization processing to obtain a structured text set.

[0042] The sentence vector embedding construction is performed based on the structured text set to obtain a sentence vector set, specifically including: Obtaining a word vector model, establishing a mapping index between vocabulary and vectors based on the word vector model, and performing vector retrieval on the structured text set based on the mapping index to obtain multiple valid word vectors; All the valid word vectors are averaged element-wise to obtain a plurality of sentence embedding vectors, and a sentence vector set is obtained based on all the sentence embedding vectors.

[0043] The step of performing dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and performing clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result specifically includes: Setting dimensionality reduction parameters, and performing dimensionality reduction processing on the sentence vector set according to the dimensionality reduction parameters to obtain a low-dimensional sentence vector set, wherein the dimensionality reduction parameters include the number of neighbors, the output dimension, the first distance metric, and the minimum distance; Clustering parameters are set, and the low-dimensional sentence vector set is clustered according to the clustering parameters to obtain an initial topic clustering result, wherein the clustering parameters include a minimum cluster size, a second distance measurement method, and a clustering mode.

[0044] The optimization process of the initial topic clustering result to obtain a topic set specifically includes: Performing topic similarity calculation on any two initial topics of the initial topic clustering result to obtain multiple first similarity calculation results; Performing keyword extraction on all the initial topics to obtain multiple semantic identifier sets, and performing word meaning similarity calculation on any two of the semantic identifier sets to obtain multiple second similarity calculation results; Comparing all the first similarity calculation results with a first preset threshold, and comparing all the second similarity calculation results with a second preset threshold; If there are two initial topics whose first similarity calculation results are greater than or equal to the first preset threshold, and whose second similarity calculation results are greater than or equal to the second preset threshold, then aggregate the two initial topics to obtain an aggregated result; The topic numbers are updated according to the aggregation results to obtain multiple updated topics, and a topic set is obtained based on all the updated topics.

[0045] The keyword extraction is performed on all the initial topics to obtain multiple semantic identification sets, specifically including: Obtaining the topic sentence texts belonging to each of the initial topics, and constructing subsets of all the topic sentence texts of each of the initial topics to obtain multiple topic corpus subsets; A word frequency-inverse document frequency calculation is performed on each word in each of the subject corpus subsets to obtain multiple calculation results, a keyword list for each of the initial topics is obtained based on all the calculation results, and a semantic identification set for each initial topic is obtained based on all the keyword lists.

[0046] The term frequency-inverse document frequency calculation is performed on each term in each of the subject corpus subsets, specifically: ; in, For words, For the initial theme, is the term frequency-inverse document frequency of the term in the initial topic, is the frequency of the word appearing in the initial topic, is the total number of initial topics, is the number of initial topics containing the word, is the inverse document frequency of the term.

[0047] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a topic extraction program for architectural design short texts, and when the topic extraction program for architectural design short texts is executed by a processor, the steps of the topic extraction method for architectural design short texts as described above are implemented.

[0048] In summary, the present invention provides a topic extraction method, system, terminal and storage medium for architectural design short texts, the method comprising: obtaining an architectural design short text dataset, preprocessing the architectural design short text dataset to obtain a structured text set; constructing sentence vector embedding based on the structured text set to obtain a sentence vector set, performing dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and clustering the low-dimensional sentence vector set to obtain an initial topic clustering result; optimizing the initial topic clustering result to obtain a topic set, and structuring the topic set to obtain a target topic set. The present invention not only enhances the adaptability of topic extraction but also improves the structuring and accuracy of topic extraction by performing sentence-level semantic embedding and clustering optimization processing on architectural design short texts.

[0049] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0050] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When executed, the program can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0051] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A topic extraction method for architectural design short texts, characterized by: The topic extraction method for architectural design short texts includes: Acquire an architectural design short text dataset, and preprocess the architectural design short text dataset to obtain a structured text set; Perform sentence embedding construction based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The initial topic clustering result is optimized to obtain a topic set, and the topic set is structured to obtain a target topic set.

2. The topic extraction method for architectural design short text according to claim 1 is characterized in that: The step of obtaining an architectural design short text dataset and preprocessing the architectural design short text dataset to obtain a structured text set specifically includes: Acquire an architectural design short text dataset, perform sentence extraction on the architectural design short text dataset, and obtain a sentence text corpus set; Performing text denoising on basic characters of all sentence texts in the sentence text corpus to obtain a denoised sentence text set, and performing word segmentation processing on the denoised sentence text set to obtain a word segmentation result; Eliminate stop words from the segmentation result according to the stop word list to obtain an initial segmentation result, and filter high-frequency words from the initial segmentation result according to the high-frequency term filter table to obtain a target segmentation result; The target word segmentation results are subjected to synonym normalization processing to obtain a structured text set.

3. The topic extraction method for architectural design short text according to claim 1 is characterized in that: The sentence vector embedding construction is performed based on the structured text set to obtain a sentence vector set, specifically including: Obtaining a word vector model, establishing a mapping index between vocabulary and vectors based on the word vector model, and performing vector retrieval on the structured text set based on the mapping index to obtain multiple valid word vectors; All the valid word vectors are averaged element-wise to obtain a plurality of sentence embedding vectors, and a sentence vector set is obtained based on all the sentence embedding vectors.

4. The topic extraction method for architectural design short text according to claim 1 is characterized in that: The step of performing dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and performing clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result, specifically includes: Setting dimensionality reduction parameters, and performing dimensionality reduction processing on the sentence vector set according to the dimensionality reduction parameters to obtain a low-dimensional sentence vector set, wherein the dimensionality reduction parameters include the number of neighbors, the output dimension, the first distance metric, and the minimum distance; Clustering parameters are set, and the low-dimensional sentence vector set is clustered according to the clustering parameters to obtain an initial topic clustering result, wherein the clustering parameters include a minimum cluster size, a second distance measurement method, and a clustering mode.

5. The topic extraction method for architectural design short text according to claim 1 is characterized in that: The optimizing process of the initial topic clustering result to obtain a topic set specifically includes: Performing topic similarity calculation on any two initial topics of the initial topic clustering result to obtain multiple first similarity calculation results; Performing keyword extraction on all the initial topics to obtain multiple semantic identifier sets, and performing word meaning similarity calculation on any two of the semantic identifier sets to obtain multiple second similarity calculation results; Comparing all the first similarity calculation results with a first preset threshold, and comparing all the second similarity calculation results with a second preset threshold; If there are two initial topics whose first similarity calculation results are greater than or equal to the first preset threshold, and whose second similarity calculation results are greater than or equal to the second preset threshold, then aggregate the two initial topics to obtain an aggregated result; The topic numbers are updated according to the aggregation results to obtain multiple updated topics, and a topic set is obtained based on all the updated topics.

6. The topic extraction method for architectural design short text according to claim 5 is characterized in that: The keyword extraction is performed on all the initial topics to obtain multiple semantic identification sets, specifically including: Obtaining the topic sentence texts belonging to each of the initial topics, and constructing subsets of all the topic sentence texts of each of the initial topics to obtain multiple topic corpus subsets; A word frequency-inverse document frequency calculation is performed on each word in each of the subject corpus subsets to obtain multiple calculation results, a keyword list for each of the initial topics is obtained based on all the calculation results, and a semantic identification set for each initial topic is obtained based on all the keyword lists.

7. The topic extraction method for architectural design short text according to claim 6 is characterized in that: The term frequency-inverse document frequency calculation is performed on each term in each of the subject corpus subsets, specifically: ; in, For words, For the initial theme, is the term frequency-inverse document frequency of the term in the initial topic, is the frequency of the word appearing in the initial topic, is the total number of initial topics, is the number of initial topics containing the word, is the inverse document frequency of the term.

8. A topic extraction system for architectural design short texts, characterized by: The topic extraction system for architectural design short texts includes: A data preprocessing module is used to obtain an architectural design short text dataset, preprocess the architectural design short text dataset, and obtain a structured text set; A text processing module is used to construct sentence vector embedding based on the structured text set to obtain a sentence vector set, perform dimensionality reduction processing on the sentence vector set to obtain a low-dimensional sentence vector set, and perform clustering processing on the low-dimensional sentence vector set to obtain an initial topic clustering result; The topic optimization module is used to optimize the initial topic clustering result to obtain a topic set, and perform structured processing on the topic set to obtain a target topic set.

9. A terminal, characterized in that: The terminal includes a memory, a processor, and a program stored in the memory and executable on the processor. When the program is executed by the processor, the steps of the topic extraction method for architectural design short texts as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer-readable storage medium stores a topic extraction program for architectural design short texts. When the topic extraction program for architectural design short texts is executed by a processor, the steps of the topic extraction method for architectural design short texts as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Theme information-based text segmentation method

    CN110110326A

  • Short text clustering method based on adaptive variational encoder

    CN114625879A

  • Text clustering method and device, electronic equipment and storage medium

    CN116992026A

  • Work order text processing method and device

    CN119128140A