Method for extracting spatial features of gravity dam design based on thought chain and creating dataset

By acquiring and integrating multi-source data, generating multidimensional datasets, and simulating engineers' design thinking, the problem of scattered gravity dam design data was solved, improving design efficiency and performance.

CN121234313BActive Publication Date: 2026-03-24POWER CHINA KUNMING ENG CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Gravity dam design data is scattered and diverse, lacking unified integration, resulting in low data utilization and insufficient design efficiency and performance.

Method used

By acquiring multi-source heterogeneous data, extracting multi-source feature parameters and design semantic data, using thought chain rules to associate and integrate data, generating a multidimensional dataset, and using artificial intelligence technology for data processing and analysis, the design thinking process of engineers is simulated, and a dataset is constructed to support design optimization.

Benefits of technology

It improves the efficiency and performance of gravity dam design. Through the construction of datasets and artificial intelligence analysis, it enables efficient use of data and supports design decisions, reduces redundant information, and improves the accuracy and efficiency of design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121234313B_ABST
    Figure CN121234313B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence for water conservancy and hydropower engineering, in particular to a gravity dam design space feature extraction and data set creation method based on a thinking chain, which collects and integrates safety and evaluation data of the gravity dam, integrates an efficient visual technology information platform to design a feature fusion algorithm design space, and generates the performance and efficiency of the gravity dam design. In order to effectively utilize historical design experience, reduce data dimension and reduce data training amount, design data sets for thinking chain guidance are added in the data set, the design semantics in the gravity dam design are converted into structured data representation through simulation of the thinking chain process of human designers, a data set capable of reflecting the design decision logic and function-scheme correlation is constructed, and the data set can support subsequent generative design training needs.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence technology for water conservancy and hydropower engineering, in particular to a method for extracting spatial features of gravity dam design based on thought chain and creating data set. BACKGROUND

[0002] The main problem of current gravity dam design data is that the information is complex, the data is scattered and insufficient. Gravity dam design involves a large amount of historical data, including hub layout, dam three-dimensional drawing, design scheme report and safety evaluation report, etc. These data are widely sourced and various in form, lack of unified integration and management, resulting in low data utilization rate.

[0003] Therefore, the gravity dam design data from multiple design units and high-level research institutions in the country are integrated, a standard design data set is constructed, and a design space exploration method based on feature learning and prior information is used to extract low-dimensional geometric features and key experience indicators from the design data, quickly estimate the performance of the gravity dam structure, filter the redundant information in the design space, and clearly define the optimization range, thereby significantly improving the design efficiency and performance. For the same project, the specification report drawing and model are associated to form a multi-dimensional feature fusion model of data fusion. Through such simultaneous data features, a basis is provided for subsequent machine learning and artificial intelligence learning. SUMMARY

[0004] To achieve the above purpose, the present application provides the following technical scheme:

[0005] According to the first aspect of the present application, a method for extracting spatial features of gravity dam design based on thought chain and creating data set is claimed, comprising:

[0006] Obtaining multi-source heterogeneous data of a gravity dam to be designed, and extracting and analyzing multi-source feature parameters of the gravity dam to be designed from the multi-source heterogeneous data;

[0007] Obtaining design requirement information of the gravity dam to be designed, and obtaining design semantic data of the design requirement by semantic analysis of the design requirement information;

[0008] Associating and integrating the multi-source feature parameters and design semantic data to obtain a multi-dimensional data set of the gravity dam to be designed;

[0009] Using thought chain rules to process the multi-dimensional data set of the gravity dam to be designed to generate a multi-dimensional data thought chain and record.

[0010] Further, the obtaining of multi-source heterogeneous data of a gravity dam to be designed, and the extracting and analyzing of multi-source feature parameters of the gravity dam to be designed from the multi-source heterogeneous data, further comprises:

[0011] collecting multi-source design data of a gravity dam, and extracting structured data in the multi-source design data;

[0012] extracting first design parameters from the multi-source design data, distinguishing the first design parameters based on key indicators, and unifying data formats and units;

[0013] extracting second design parameters from the multi-source design data, and constructing a hydrogeological generalization model;

[0014] extracting third design parameters from the multi-source design data, extracting graphic data features from the third design parameters, and obtaining graphic labeling and attribute information;

[0015] converting the graphic data features into JSON format, and storing the graphic data features and attribute information in a database;

[0016] extracting features of a BIM model, extracting geometric feature surfaces of a single dam section model, and decomposing into modeling feature parameters;

[0017] extracting feature information of the single dam section model based on a BIM model geometric API;

[0018] constructing an integrated data platform and a GIS information platform to manage data, and performing visual analysis and intelligent search of data.

[0019] Further, the design requirement information of the gravity dam to be designed is obtained, semantic analysis is performed on the design requirement information to obtain design semantic data of the design requirement, and the method further includes:

[0020] data alignment and entity analysis are performed using a data fusion algorithm, the multi-source feature parameters are associated with safety checking data, a high-dimensional design feature model is constructed, and the model is used for training a structure analysis proxy model;

[0021] a design function description semantic library is established, the design requirement information is expressed in a feature form, a multi-dimensional word vector is used to express the design requirement information, corresponding feature descriptions are mapped to design parameters and key design variables, multi-dimensional design information expression is formed, the design semantic recognition model is trained, and design semantic data is obtained;

[0022] the design requirement information is mapped to design parameters and key design variables through a long short-term memory network (LSTM).

[0023] Further, the multi-source feature parameters and design semantic data are associated and integrated to obtain a multi-dimensional data set of the gravity dam to be designed, and the method further includes:

[0024] In a low-dimensional design space, recursive feature elimination (RFE) is used to extract key variables in the low-dimensional design space;

[0025] The key variables are extracted, and principal component analysis (PCA) is used to support multi-objective optimization design training.

[0026] A multi-modal design dataset is constructed, multi-model objects are parameterized and reconstructed through code, and code ID is extracted as a feature to support data communication driven parameter mapping modification.

[0027] The design parameters, design semantic data, and multi-source feature parameters are associated and integrated using database technology or data fusion framework Apache Spark to generate a multi-modal data product that supports artificial intelligence training.

[0028] Further, the processing of the multi-dimensional dataset of the to-be-designed gravity dam using the thought chain rule generates a multi-dimensional data thought chain and records, further comprising:

[0029] Prepare a preset dataset case, extract labeled information and structure it;

[0030] The preset dataset case is preliminarily classified, and multi-dimensional classification is performed according to main function combination, presence or absence of key layout tendency;

[0031] Based on the basic principles and experience of gravity dam design, define multiple candidate thought chain templates or reasoning rules;

[0032] For each labeled preset dataset case, according to the functional characteristics, apply the candidate thought chain template or reasoning rule to simulate the thinking process of an engineer;

[0033] Use a rule engine for management and execution, automatically generate a target thought chain, and decompose the generation process of the target thought chain into several steps and record them;

[0034] The generation process of the target thought chain is associated and stored with the functional characteristics and scheme parameters of the corresponding preset dataset case;

[0035] The generation process of the target thought chain associated and stored with the functional characteristics and scheme parameters of the corresponding preset dataset case is integrated and structured with the dataset;

[0036] Similarity measurement algorithm is used for similar scheme recommendation.

[0037] Further, the dataset integration and structured storage of the generation process of the target thought chain associated and stored with the functional characteristics and scheme parameters of the corresponding preset dataset case further comprises:

[0038] According to the function feature vector or symbol of each preset data set case, the generated thought chain step sequence, the preliminary classification label, and the final dam layout scheme key parameter are integrated into a complete data sample;

[0039] The integrity and consistency of the data sample are checked, and the characteristic value is standardized or normalized;

[0040] The matching data structure is selected to store the data set, the table structure or document mode is designed, and the function features, thought chains, scheme parameters, and classification label fields are ensured to correspond.

[0041] Further, the similarity scheme recommendation according to the similarity measurement algorithm further includes:

[0042] Based on the constructed data set, a similarity measurement method is designed, and the similarity measurement method includes:

[0043] The function feature similarity is calculated by using the Jaccard similarity method to calculate the similarity of the query function feature and the sample function feature in the data set;

[0044] Or the thought chain similarity, for the thought chain represented by structured symbols or keywords, the similarity of the thought chain sequence is calculated;

[0045] Or the comprehensive similarity, the function feature similarity and the thought chain similarity are given different weights to obtain the comprehensive similarity score;

[0046] An approximate nearest neighbor search index is constructed based on the feature vector.

[0047] The present application relates to the field of water conservancy and hydropower engineering artificial intelligence technology, in particular to a gravity dam design space feature extraction and data set creation method based on thought chains, which collects and integrates safety and evaluation data of gravity dams, integrates an efficient visual technology information platform to design a feature fusion algorithm space, and generates a gravity dam design performance and efficiency. In order to effectively utilize historical design experience, reduce data dimension and data training amount, design data sets for thought chain guidance are added in the data set, the design semantics in the gravity dam design is converted into structured data representation by simulating the thought chain process of human designers, a data set reflecting design decision logic and function-scheme correlation is constructed, which can support subsequent generative design training needs. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 A gravity dam design space feature extraction and data set creation method based on thought chains is requested to protect the work flow chart of the embodiment of the present application;

[0049] Figure 2This is a schematic diagram illustrating the intelligent generation effect of multiple gravity dam schemes based on a method for extracting spatial features and creating datasets for gravity dam design using a thought chain-based approach, as claimed in an embodiment of this application. Detailed Implementation

[0050] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0051] The terms "first," "second," and "third" in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. All directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationships and movements between components in a specific orientation (as shown in the figures). If the specific orientation changes, the directional indications also change accordingly. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0052] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0053] According to a first embodiment of the present invention, the present invention claims protection for a method for extracting spatial features and creating a dataset for gravity dam design based on thought chains, referring to... Figure 1 ,include:

[0054] Obtain multi-source heterogeneous data of the gravity dam to be designed, and extract and analyze the multi-source characteristic parameters of the gravity dam to be designed from the multi-source heterogeneous data;

[0055] Obtain the design requirements information of the gravity dam to be designed, and perform semantic analysis on the design requirements information to obtain the design semantic data of the design requirements;

[0056] The multi-source feature parameters and design semantic data are correlated and integrated to obtain the multidimensional dataset of the gravity dam to be designed;

[0057] Using the thinking chain rule, the multidimensional dataset of the gravity dam to be designed is processed to generate a multidimensional data thinking chain and recorded.

[0058] Furthermore, the step of acquiring multi-source heterogeneous data of the gravity dam to be designed, and extracting and analyzing multi-source characteristic parameters of the gravity dam to be designed from the multi-source heterogeneous data, further includes:

[0059] Collect multi-source design data of gravity dams and extract structured data from the multi-source design data;

[0060] The first design parameter is extracted from the multi-source design data, and the first design parameter is distinguished based on key indicators and the data format and unit standard are unified.

[0061] A second design parameter is extracted from the multi-source design data to construct a generalized hydrogeological model.

[0062] A third design parameter is extracted from the multi-source design data, and graphic element data features are extracted from the third design parameter to obtain graphic element annotations and attribute information.

[0063] The graph data features are converted into JSON format and stored in a database.

[0064] Features are extracted from the BIM model, the geometric feature surfaces of a single dam section model are extracted, and decomposed into shape feature parameters;

[0065] Based on the BIM model geometry API, feature information of the individual dam section model is extracted;

[0066] Build an integrated data platform and GIS information platform to manage data and perform data visualization analysis and intelligent search.

[0067] In this embodiment, domestic gravity dam design data is collected, including the hub layout, dam 3D drawings, design scheme report, safety evaluation report and design specifications, and the structured data in the PDF documents is extracted using the PDF parsing library PyPDF2.

[0068] BIM models are generated through software APIs, which obtain the geometry of model components and iterate through the edges and faces of the geometry. Among them, the features that are coplanar, closed, and have the most edges are the geometric feature contours.

[0069] Design parameters were extracted from the design report, and the collected data were categorized based on key indicators including: dam axis length, maximum dam height, upstream dam slope ratio, downstream dam slope ratio, dam material strength parameters, net width of spillway structures, and spillway curves (foundation surface). A large language model was used to verify and normalize the extracted data format and unit standards, ensuring that data from different sources could be compared and integrated. The SVM method was then used to classify and rate the parameters.

[0070] By using regular expressions to extract data from the design report, such as maximum discharge, flood flow at various frequencies, foundation bearing capacity, foundation shear strength parameters, and permeability coefficients of various rock strata, a generalized hydrogeological model is constructed. This model allows a specific set of hydrogeological parameters to correspond to a gravity dam design scheme, providing design guidance under conditions of limited information.

[0071] Extracting engineering drawings: For CAD drawings, the Python ezdxf library is used to extract metadata, extracting design features in the form of graphic elements, and extracting element annotations and attribute information in text form. The extracted metadata is converted to JSON format for easier subsequent processing and analysis; a database is used to store the metadata and its attribute information. The X and Y values ​​of the element node data for the gravity dam design cross-section are extracted, and the element name, parameters, and attribute extensions are written into a JSON file.

[0072] Feature extraction is performed on the BIM model. Geometric feature surfaces of individual dam sections are extracted and decomposed into shape feature parameters. Based on the BIM model geometry API, the cross-sectional area, volume, and centroid information of individual dam sections are extracted. Information model attributes are extracted to obtain information such as model materials and engineering project details.

[0073] By building an integrated data platform and GIS information platform to manage data, we can ensure high data quality and ease of use, and support data visualization analysis and intelligent search functions.

[0074] Furthermore, the step of obtaining the design requirement information of the gravity dam to be designed, and performing semantic analysis on the design requirement information to obtain the design semantic data of the design requirements, further includes:

[0075] Data fusion algorithms are used for data alignment and entity parsing. The multi-source feature parameters are associated with security verification data to construct a high-dimensional design feature model for training the structural analysis proxy model.

[0076] Establish a semantic library for design function description, express the design requirement information using features, express the design requirement information using multi-dimensional word vectors, correspond to feature descriptions, map design parameters and key design variables to form a multi-dimensional design information expression, which is used to train the design semantic recognition model to obtain design semantic data;

[0077] The design requirements information is mapped to design parameters and key design variables using a Long Short-Term Memory (LSTM) network.

[0078] In this embodiment, the fused sequence data is transformed into a high-dimensional tensor representation. By calculating the spatial distance between different sequences, the similarity of the sequence data can be accurately evaluated, thereby mining the latent features in the sequence data to generate a design proxy model.

[0079] Design functional semantic descriptions are personalized descriptions of design intent and key parameter characteristics. Examples include design descriptions such as "complex regional geological conditions," "high ecological and environmental protection requirements," "high investment control requirements," and "conservative technical and economic indicators." By extracting highly subjective and specific functional descriptions from design reports and associating them with design scheme data, the system supports designers in inputting relatively subjective design strategies, and recommends and predicts corresponding technical solution data based on semantic similarity.

[0080] Furthermore, the step of associating and integrating the multi-source feature parameters and design semantic data to obtain the multidimensional dataset of the gravity dam to be designed also includes:

[0081] In the low-dimensional design space, recursive feature elimination (RFE) is used to extract key variables from the low-dimensional design space;

[0082] Extract key variables and use principal component analysis (PCA) to support multi-objective optimization design training.

[0083] Construct a multimodal design dataset, parametrically reconstruct multi-model objects through code and extract code IDs as features, and support data communication to drive parameter mapping modification;

[0084] By using database technology or the data fusion framework Apache Spark, the design parameters, design semantic data, and multi-source feature parameters are correlated and integrated to generate multimodal data products that support artificial intelligence training.

[0085] Furthermore, the step of using the thought chain rule to process the multidimensional dataset of the gravity dam to be designed, generate a multidimensional data thought chain, and record it, also includes:

[0086] Prepare a pre-defined dataset of cases, extract annotation information, and structure it.

[0087] The preset dataset cases are initially classified, and then classified in multiple dimensions based on the main functional combinations and whether there is a tendency to deploy key features;

[0088] Based on the fundamental principles and experience of gravity dam design, multiple candidate thought chain templates or reasoning rules are defined.

[0089] For each labeled preset dataset case, based on functional characteristics, the candidate thought chain template or reasoning rules are applied to simulate the engineer's thinking process.

[0090] The rule engine is used for management and execution, and the target thinking chain is automatically generated. The generation process of the target thinking chain is broken down into several steps and recorded.

[0091] The generation process of the target thinking chain is associated with and stored in relation to the functional features and solution parameters of the corresponding preset dataset cases;

[0092] The generation process of the target thinking chain, which is stored in association, is integrated and structured with the functional features and scheme parameters of the corresponding preset dataset cases.

[0093] Recommendations for similar solutions are made based on similarity measurement algorithms.

[0094] In this embodiment, from the perspective of dataset construction, the more design feature parameters there are, the longer and more dimensional the resulting sequence data becomes, and the larger the number of dataset samples required. This contradicts the current situation of sparse instance data in the engineering construction industry. By classifying and performing sensitivity analysis on the design feature data and filtering key data, the data dimensionality and dataset size can be reduced, while improving data matching accuracy.

[0095] The design characteristics of gravity dams reveal diverse layout options. For example, the combination of a dam and a power plant, a water diversion system, and the presence or absence of fish passages all present challenges to the applicability of the dataset. This paper extracts and categorizes pre-defined dataset cases, establishing a thought-chain template for different layout forms. This allows for the classification of design types based on broad categories, and the creation of targeted datasets for each category. This approach ensures data accuracy while reducing the difficulty of dataset construction.

[0096] Extract key information from the collected cases:

[0097] Basic project information: dam site, scale, main functions (power generation, flood control, navigation, water supply, etc.).

[0098] Functional features: Clearly indicate whether locks, fishways, and workshops are to be installed, as well as the designer's preferred layout (compact / open).

[0099] Key parameters of the dam layout scheme include: dam axis length, main dam section types and locations (overflow dam section, non-overflow dam section, powerhouse dam section, lock dam section, etc.), typical cross-sectional dimensions, and main structural relationships.

[0100] Design constraints and considerations: Record the geological conditions, environmental requirements, construction conditions, economic considerations, etc., explicitly mentioned in the design;

[0101] During the initial classification, each case is preliminarily categorized based on the extracted functional characteristics (locks, fishways, factory buildings, layout preferences). For example, classification can be carried out in multiple dimensions, such as by main functional combination (e.g., "power generation + shipping", "power generation + water supply"), by the presence or absence of key facilities (e.g., "with locks", "without locks"), and by layout preferences ("compact", "open").

[0102] Based on the fundamental principles and experience of gravity dam design, a series of possible "thought chain" templates or reasoning rules are defined. For example:

[0103] Rule 1 (Impact of Locks): If locks are to be installed, the size of the locks, the head of the water, and the layout of the approach channels must be considered. This usually requires reserving a specific location on the dam body (lock dam section), which may affect the overall length of the dam body and the local structural strength, and tends to be arranged in an open manner.

[0104] Rule 2 (Impact of Powerhouse): If a powerhouse is to be built (especially a downstream or riverbed type), the location of the powerhouse (inside the dam, downstream, or on the bank) must be determined. This will directly affect the downstream profile, foundation treatment, and relationship with the spillway section, and may require more complex structural connections.

[0105] Rule 3 (Impact of fish passages): The design of fish passages is usually related to ecological requirements. Their location and form need to take into account the migration routes of fish and the water flow conditions of the dam. This may increase the amount of engineering work, but the impact on the overall layout is relatively small compared to locks and powerhouses.

[0106] Rule 4 (Compact vs. Open): A compact layout may save land and reduce the amount of construction work, but it requires high structural stress and construction precision; an open layout facilitates construction, operation, and maintenance, but may increase land occupation and the amount of construction work. A balance must be struck between functional requirements, geological conditions, and economic efficiency.

[0107] Rule 5 (Functional Conflicts and Coordination): Analyze whether there are spatial or operational conflicts between different functional facilities (such as overlapping locations of locks and factory buildings, or water flow interference between fishways and overflow surfaces), and consider coordination solutions.

[0108] Break down the thought process into several steps and record them. For example, for a case that "includes a lock, a factory, and is oriented towards an open area," the thought process might be recorded as follows:

[0109] Step 1: Identify the lock and powerhouse to be installed, and preliminarily determine that the dam body needs to include a lock section and a powerhouse section.

[0110] Step 2: Consider the relative positions of the lock and the powerhouse to avoid mutual interference. It is preferable to place the lock at one end and the powerhouse on the riverbed or the other side.

[0111] Step 3: In line with the "open layout" tendency, determine that the length of the dam axis must meet the layout requirements of each dam section and leave sufficient spacing.

[0112] Step 4: Considering the overflow demand, arrange overflow surface or intermediate outlets in the remaining dam sections.

[0113] Step 5: Check whether the geological conditions support this layout, and make adjustments if necessary.

[0114] Discrete symbols (such as lock = yes, factory = no) or continuous / discrete hybrid vectors can be used. For layout tendency, layout tendency = compact or layout tendency = open can be used.

[0115] The thought chain can be stored in an ordered list, where each element is a step. The content of each step can be text, a list of keywords, or a structured dictionary (containing step type, involved objects, judgment result, etc.). For example: [{"step": 1,"type": "Function recognition", "content": ["lock", "factory"]}, {"step": 2, "type": "location coordination", "content": "lock end, factory riverbed"}].

[0116] The parameters of the scheme are stored using key-value pairs (dictionary / JSON), where the key is the parameter name (such as "dam axis length" or "powerhouse dam section location"), and the value is a specific numerical value or description.

[0117] Functional characteristics: Jaccard similarity is applicable to discrete symbols; cosine similarity is commonly used for vectors.

[0118] For thought chain sequences: If it's text, TF-IDF vectors plus cosine similarity can be used; if it's a structured representation, a specific matching algorithm can be designed, considering the step order and content overlap. For example, a weighted average of the content similarity of steps at the same position in two thought chain sequences can be calculated.

[0119] Overall similarity = w1 * functional feature similarity + w2 * thought chain similarity, where w1 + w2 = 1, and the weights can be adjusted according to importance.

[0120] Furthermore, the process of generating the target thought chain in association with the corresponding preset dataset cases, and integrating and structurally storing the data, also includes:

[0121] The functional feature vectors or symbols of each preset dataset case, the generated thought chain step sequence, the preliminary classification labels, and the key parameters of the final dam layout scheme are integrated into a complete data sample.

[0122] Check the integrity and consistency of the data samples, and standardize or normalize the feature values;

[0123] Choose a matching data structure to store the dataset, design table structures or document schemas, and ensure that functional features, thought processes, solution parameters, and classification label fields correspond.

[0124] In this embodiment, the feature value is the value of the extracted key design parameter. For example, the unit of design flood is flow rate, and the unit of characteristic water level of gravity dam is meter. The value range and unit of different design parameters are not the same. By using a scheme that integrates multi-source data to construct sequence feature data, the complex design features can be described as high-dimensional tensor features that can be calculated, analyzed and compared. However, in this process, the feature values ​​need to be normalized and standardized, while controlling the dimension of the data to a certain extent and highlighting the principal components.

[0125] By extracting data samples, the resulting sequence of feature values ​​constitutes a dataset for machine learning or deep learning.

[0126] For example, a relational database might contain a "case table" (ID, basic information, functional characteristics, preliminary classification), a "thinking list" (ID, case ID, thinking step number, step content), and a "solution parameter table" (ID, case ID, parameter name, parameter value).

[0127] Furthermore, the recommendation of similar schemes based on the similarity measurement algorithm also includes:

[0128] Based on the constructed dataset, a similarity measurement method is designed, which includes:

[0129] Functional feature similarity: The Jaccard similarity method is used to calculate the similarity between the query functional features and the functional features of the samples in the dataset.

[0130] Or, for thought chain similarity, calculate the similarity of thought chain sequences for thought chains represented by structured symbols or keywords;

[0131] Alternatively, a comprehensive similarity score can be obtained by assigning different weights to functional feature similarity and thought chain similarity.

[0132] An index structure is constructed based on an approximate nearest neighbor search index using feature vectors.

[0133] Reference Figure 2 Based on Tables 1 and 2, the sequence data characteristics of gravity dam design are presented from multiple dimensions. The resulting sequence data feature set can be used to support the matching of design feature data after the design intent has been identified.

[0134] Table 1. Schematic diagram of multi-dimensional feature fusion data structure for dam section layout

[0135]

[0136] Table 2. Schematic diagram of multi-dimensional feature fusion data structure for non-overflow dam sections.

[0137]

[0138] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0139] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units. The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

[0140] The specific embodiments of the invention have been described in detail above, but they are only examples, and this application is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of this application. Therefore, all equivalent changes, modifications, and improvements made without departing from the spirit and principles of this application should be covered within the scope of this application.

Claims

1. A method for extracting spatial features and creating a dataset for gravity dam design based on thought chains, characterized in that, include: Obtain multi-source heterogeneous data of the gravity dam to be designed, and extract and analyze the multi-source characteristic parameters of the gravity dam to be designed from the multi-source heterogeneous data; Obtain the design requirements information of the gravity dam to be designed, and perform semantic analysis on the design requirements information to obtain the design semantic data of the design requirements; The multi-source feature parameters and design semantic data are correlated and integrated to obtain the multidimensional dataset of the gravity dam to be designed; Using the thinking chain rule, the multidimensional data dataset of the gravity dam to be designed is processed to generate and record a multidimensional data thinking chain; The step of processing the multidimensional dataset of the gravity dam to be designed using the thinking chain rule to generate and record a multidimensional data thinking chain also includes: Prepare a pre-defined dataset of cases, extract annotation information, and structure it. The preset dataset cases are initially classified, and then classified in multiple dimensions based on the main functional combinations and whether there is a tendency to deploy key features; Based on the fundamental principles and experience of gravity dam design, multiple candidate thought chain templates or reasoning rules are defined. For each labeled preset dataset case, based on functional characteristics, the candidate thought chain template or reasoning rules are applied to simulate the engineer's thinking process. The rule engine is used for management and execution, and the target thinking chain is automatically generated. The generation process of the target thinking chain is broken down into several steps and recorded. The generation process of the target thinking chain is associated with and stored in relation to the functional features and solution parameters of the corresponding preset dataset cases; The generation process of the target thinking chain, which is stored in association, is integrated and structured with the functional features and scheme parameters of the corresponding preset dataset cases. Recommendations for similar solutions are made based on similarity measurement algorithms.

2. The method for extracting spatial features and creating a dataset for gravity dam design based on thought chain as described in claim 1, characterized in that, The process of acquiring multi-source heterogeneous data of the gravity dam to be designed, and extracting and analyzing multi-source characteristic parameters of the gravity dam to be designed from the multi-source heterogeneous data, further includes: Collect multi-source design data of gravity dams and extract structured data from the multi-source design data; The first design parameter is extracted from the multi-source design data, and the first design parameter is distinguished based on key indicators and the data format and unit standard are unified. A second design parameter is extracted from the multi-source design data to construct a generalized hydrogeological model. A third design parameter is extracted from the multi-source design data, and graphic element data features are extracted from the third design parameter to obtain graphic element annotations and attribute information. The graph data features are converted into JSON format and stored in a database. Features are extracted from the BIM model, the geometric feature surfaces of a single dam section model are extracted, and decomposed into shape feature parameters; Based on the BIM model geometry API, feature information of the individual dam section model is extracted; Build an integrated data platform and GIS information platform to manage data and perform data visualization analysis and intelligent search.

3. The method for extracting spatial features and creating a dataset for gravity dam design based on thought chain as described in claim 1, characterized in that, The step of obtaining the design requirement information of the gravity dam to be designed, and performing semantic analysis on the design requirement information to obtain the design semantic data of the design requirements, further includes: Data fusion algorithms are used for data alignment and entity parsing. The multi-source feature parameters are associated with security verification data to construct a high-dimensional design feature model for training the structural analysis proxy model. Establish a semantic library for design function description, express the design requirement information using features, express the design requirement information using multi-dimensional word vectors, correspond to feature descriptions, map design parameters and key design variables to form a multi-dimensional design information expression, which is used to train the design semantic recognition model to obtain design semantic data; The design requirements information is mapped to design parameters and key design variables using a Long Short-Term Memory (LSTM) network.

4. The method for extracting spatial features and creating a dataset for gravity dam design based on thought chain as described in claim 3, characterized in that, The step of associating and integrating the multi-source feature parameters and design semantic data to obtain the multidimensional dataset of the gravity dam to be designed further includes: In the low-dimensional design space, recursive feature elimination (RFE) is used to extract key variables from the low-dimensional design space; Extract key variables and use principal component analysis (PCA) to support multi-objective optimization design training; Construct a multimodal design dataset, parametrically reconstruct multi-model objects through code and extract code IDs as features, and support data communication to drive parameter mapping modification; By using database technology or the data fusion framework Apache Spark, the design parameters, design semantic data, and multi-source feature parameters are correlated and integrated to generate multimodal data products that support artificial intelligence-based training.

5. The method for extracting spatial features and creating a dataset for gravity dam design based on thought chain as described in claim 4, characterized in that, The process of generating the target thought chain in association with the corresponding preset dataset cases, and integrating and structurally storing the data, also includes: The functional feature vectors or symbols of each preset dataset case, the generated thought chain step sequence, the preliminary classification labels, and the key parameters of the final dam layout scheme are integrated into a complete data sample. Check the integrity and consistency of the data samples, and standardize or normalize the feature values; Choose a matching data structure to store the dataset, design table structures or document schemas, and ensure that functional features, thought processes, solution parameters, and classification label fields correspond.

6. The method for extracting spatial features and creating a dataset for gravity dam design based on thought chain as described in claim 5, characterized in that, The recommendation of similar schemes based on the similarity measurement algorithm also includes: Based on the constructed dataset, a similarity measurement method is designed, which includes: Functional feature similarity: The Jaccard similarity method is used to calculate the similarity between the query functional features and the functional features of the samples in the dataset. Or, for thought chain similarity, calculate the similarity of thought chain sequences for thought chains represented by structured symbols or keywords; Alternatively, a comprehensive similarity score can be obtained by assigning different weights to functional feature similarity and thought chain similarity. An index structure is constructed based on an approximate nearest neighbor search index using feature vectors.

Citation Information

Patent Citations

  • Product bionic design method based on knowledge graph and semantic fusion diffusion model

    CN119180134A

  • Reservoir dam safety evaluation report multi-dimensional examination method based on large language model

    CN120373316A