A method and device for optimizing the selection of variables of a deep learning model considering geographical causal relationships

By building a geographical expertise base and fine-tuning large language models, combined with a causal discovery algorithm, the shortcomings of deep learning models in revealing geographical causal logic are solved, and more efficient and accurate causal reasoning and variable optimization are achieved.

CN119166731BActive Publication Date: 2025-06-10NANJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411134763.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-06-10
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

When existing deep learning models deal with geographical phenomena, it is difficult to reveal the causal logic between variables, resulting in the model's in-depth exploration and application of geographical knowledge.

Method used

By building a geography expertise base, pre-training and fine-tuning large language models, generating geographic causality data sets, combining causal discovery algorithms and statistical causal inference algorithms, automating causal sequence inference and variable optimization selection.

Benefits of technology

It reduces the direct dependence on expert knowledge, improves the efficiency and accuracy of causal reasoning, and can more accurately model the causal effects between geographical variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166731B_ABST
    Figure CN119166731B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for optimizing the selection of variables in a deep learning model considering geographical causal relationships; the method pre-trains an open-source large language model to obtain a first geographical large model; fine-tunes the first geographical large model to obtain a second geographical large model; receives user input including a set of variables to be studied V and associated data, determines the processing variable X and the target variable Y, and calls the first geographical large model to generate additional metadata for each variable to be studied in V; based on the second geographical large model, determines the causal order among the variables to be studied in V; constructs a causal skeleton graph for V, and orients the edges in the causal skeleton graph through the causal order to obtain a causal directed acyclic graph DAG; screens the nodes in the DAG that meet the adjustment rules to obtain an additional variable set Z; uses X and Z as the input variables of the deep learning model and Y as the output variable to complete the optimization of the selection of variables in the deep learning model considering geographical causal relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of geographical modeling and causal science, and more specifically, to a method and device for optimizing the selection of variables in a deep learning model considering geographical causal relationships. Background Art

[0002] Geographical modeling, as a key tool for exploring complex geographical systems, plays a crucial role in deepening our understanding of geographical phenomena. Nevertheless, current deep learning models for geographical phenomena often focus on mining statistical associations between variables and fail to fully reveal the underlying causal logic, which to some extent limits the in-depth exploration and application of geographical knowledge by the models.

[0003] In the context of machine learning, accurately modeling and inferring causal variables and their mechanisms can lead to more robust feature representations. This is crucial for explaining the generation principle of observational data. To achieve this goal, we usually need to deeply understand geographical phenomena, construct a causal directed acyclic graph (DAG) between geographical variables, identify and exclude confounding variables, and select the remaining variables as additional inputs for modeling features. However, geographical causal relationship modeling highly depends on expert knowledge.

[0004] The geographical field has rich literature resources, and the knowledge contained in these literatures provides valuable information for geographical causal relationship modeling. Recently, large language models have shown great potential in simulating domain expert knowledge and performing causal reasoning. Large language models can act as agents for human domain knowledge, approximately providing causal orders like experts, thereby automating the causal reasoning process. In addition, when combined with existing causal discovery algorithms, large language models are expected to significantly improve the automation efficiency of causal inference, bringing a revolutionary change to the field of geographical modeling. In this way, we can more effectively utilize the knowledge in geographical literatures and promote geographical research towards a deeper causal understanding. Summary of the Invention

[0005] To solve the technical problems mentioned in the above background art, the present invention proposes a method and device for optimizing the selection of variables in a deep learning model considering geographical causal relationships.

[0006] The present invention provides a method for optimizing the selection of variables in a deep learning model considering geographical causal relationships, including:

[0007] Step S1: Construct a geographical professional knowledge base, pre-train an open-source large language model to obtain a first geographical large model;

[0008] Step S2: Construct a causal relationship dataset containing geographical context information, and fine-tune the first geographical large model to obtain a second geographical large model;

[0009] Step S3: Receive user input including the variable set V to be studied and associated data, identify the processing variable X and the target variable Y, and call the first geographical large model to generate additional metadata for each variable to be studied in the variable set V to be studied;

[0010] Step S4: Based on the second geographical large model, determine the causal order among the variables to be studied in the variable set V to be studied;

[0011] Step S5: Construct a causal skeleton graph for the variable set V to be studied, and orient the edges in the causal skeleton graph through the causal order to obtain a causal directed acyclic graph DAG;

[0012] Step S6: Screen the nodes in the causal directed acyclic graph DAG that meet the adjustment rules to obtain an additional variable set Z;

[0013] Step S7: Use the processing variable X and the additional variable set Z as the input variables of the deep learning model, and the target variable Y as the output variable of the deep learning model to complete the optimal selection of the variables of the deep learning model considering geographical causal relationships.

[0014] Further, in step S1, construct a geographical professional knowledge base, and perform pre-training on the open-source large language model to obtain the first geographical large model, including:

[0015] Widely extract geographical knowledge text data from various public resources. The public resources include but are not limited to professional books, academic papers, industry reports, encyclopedia entries, etc. The geographical knowledge text data includes but is not limited to geographical entity concepts, classic theories, geographical phenomena, etc.

[0016] Preprocess the collected data to construct a geographical knowledge base. The preprocessing methods include but are not limited to text cleaning, text deduplication, text segmentation, etc.

[0017] Continue to perform pre-training on the open-source language large model based on the geographical knowledge base to obtain the first geographical large model. After pre-training, the first geographical large model has the general knowledge ability of geographical domain knowledge.

[0018] Further, in step S2, construct a causal relationship dataset containing geographical context information, and fine-tune the first geographical large model to obtain the second geographical large model, including:

[0019] Screen out text paragraphs including clear causal relationship descriptions from the geographical text knowledge base, such as descriptions of geographical processes and geographical conclusions, extract the geographical context of events, manually identify causal relationship pairs, and construct a geographical causal relationship dataset D. Each piece of data d in the dataset D = {STA, GD, TAs, CRDs}, where:

[0020] STA = {Location, Time}, representing spatio - temporal attributes, including geographical location (Location) and time (Time);

[0021] GD = domain, an identifier representing the geographical discipline or research field (Geographical Domain) to which it belongs, such as urban planning, environmental science, etc.;

[0022] TAs = {Attribute 1 , Attribute 2 ,..., Attribute k}, representing thematic attributes, including a set of specific geographical features, such as climate, terrain, etc.;

[0023] CRDs = {CRD 1 , CRD 2 ,..., CRD n}, representing a set of causal relationship descriptions (Causal RelationshipDescription), where CRD = {Cause; Effect; Description}, Cause represents the cause variable, Effect represents the result variable, and Description represents the description of the causal relationship.

[0024] Use the STA, GD, and TAs in D to provide context information for geographical problems, use CRDs for reasoning and analysis of geographical problems, construct question - answer pairs, and fine - tune the first geographical large - model to obtain the second geographical large - model.

[0025] Furthermore, in step S3, receive user input including the set of variables to be studied V and associated data, identify the processing variable X and the target variable Y, and call the first geographical large - model to generate additional metadata for each variable to be studied in the set of variables to be studied V, including:

[0026] The user inputs the set of variables to be studied and associated data. In addition, the user also needs to input the current modeling research field and the geographical context description of the research area. The variables to be studied are various types of terms in the geographical research field, such as air temperature, precipitation, etc. The variable set should satisfy causal sufficiency, that is, the direct cause variables of any two variables in the variable set exist in the set;

[0027] The user needs to specify which pair of variables is the processing variable X and the target variable Y. The X - Y variable pair is the variable for which the user expects to model the causal effect.

[0028] Invoke the first geographical large model to generate additional metadata for each variable in the variable set, and output it in the form of <variable name, additional metadata>. The additional metadata includes at least the semantic description of the geographical phenomenon represented by the variable.

[0029] Further, in step S4, based on the second geographical large model, determine the causal order among the variables to be studied in the variable set V to be studied, including:

[0030] For the given set of geographical variables to be studied, construct a set of all possible triples, where each triple is composed of different elements in the set, expressed as T = {<vi, vj, vk>|i, j, k ∈ {1, 2, …, S} and i ≠ j ≠ k}, where vi, vj, and vk are the i-th, j-th, and k-th variables to be studied in the set respectively; S is the total number of elements in the set V;

[0031] Adopt the triple Prompt prompting word technique to determine the causal relationship among the three variables to be studied in each triple in T based on the second geographical large model, and form a directed acyclic graph subDAG;

[0032] Through the majority voting mechanism, merge the directed acyclic graphs subDAG of all triples to comprehensively obtain the overall causal structure; the majority voting mechanism is to count the number of times the causal relationship between every two variables to be studied is determined as a causal relationship in all triples, and select the direction with the most occurrences as the final causal direction;

[0033] Record the reasoning basis of the entire causal discovery process to ensure the transparency and traceability of the process, and provide detailed documentation for users. The reasoning basis includes the response output of the geographical large model during causal orientation.

[0034] Further, in step S5, construct a causal skeleton graph for the variable set V to be studied, and orient the edges in the causal skeleton graph through the causal order to obtain a causal directed acyclic graph DAG, including:

[0035] Create a causal skeleton graph for the set of geographical variables according to the constraint-based PC (Peter-Clark) causal discovery algorithm. The edges in the causal skeleton graph are undirected or partially directed.

[0036] Use the causal order obtained in S4 to orient the causal skeleton graph to obtain the final causal directed acyclic graph DAG.

[0037] Further, in step S6, screen the nodes in the causal directed acyclic graph DAG that meet the adjustment rules to obtain an additional variable set Z, including:

[0038] Select the nodes that satisfy the adjustment criterion in the causal directed acyclic graph to form a variable set, and use variable X and the variables in the set as input variables to train a deep learning model, and calculate the causal effect through the observed data:

[0039]

[0040] Among them, E[Y|X=x] is the observed expectation, and E[Y|do(X=x)] is the intervention expectation, that is, the causal effect.

[0041] For the path from X to Y, 1) if the nodes in the set block all non-causal paths from X to Y, and 2) the nodes in the set are not on the causal path from X to Y, then the adjustment criterion for (X, Y) is satisfied.

[0042] Furthermore, in step S7, use the processing variable X and the additional variable set Z as the input variables of the deep learning model, and the target variable Y as the output variable of the deep learning model to complete the optimal selection of the variables of the deep learning model considering geographical causal relationships.

[0043] The present invention also provides a device for optimizing the selection of variables of a deep learning model considering geographical causal relationships, and the device includes:

[0044] The large model pre-training module is used to construct a causal relationship data set containing geographical context information, and fine-tune the first geographical large model to obtain a second geographical large model;

[0045] The large model fine-tuning module is used to construct a causal relationship data set containing geographical context information, and fine-tune the first geographical large model to obtain a second geographical large model;

[0046] The user interaction module is used to receive user input including the variable set V to be studied and associated data, clarify the processing variable X and the target variable Y, and call the first geographical large model to generate additional metadata for each variable to be studied in the variable set V to be studied;

[0047] The causal order generation module is used to determine the causal order between the variables to be studied in the variable set V to be studied based on the second geographical large model by using the triple prompt word technique;

[0048] The causal graph generation module is used to construct a causal skeleton graph for the variable set V to be studied, and orient the edges in the causal skeleton graph through the causal order to obtain a causal directed acyclic graph DAG;

[0049] The additional variable screening module is used to screen the nodes that satisfy the adjustment rule in the causal directed acyclic graph DAG to obtain an additional variable set Z.

[0050] The present invention also provides a computer-readable storage medium storing one or more programs, the one or more programs including instructions which, when executed by a computing device, cause the computing device to execute the method as described above.

[0051] The present invention also provides an electronic device, including one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method as described above.

[0052] After adopting the above solution, the present invention has the following beneficial effects:

[0053] 1) The present invention performs automated causal order inference by simulating expert knowledge through a large language model, reducing the direct dependence on expert knowledge and improving the efficiency of causal reasoning.

[0054] 2) The present invention improves the interpretability and accuracy in the process of causal discovery of geographical variables by integrating knowledge-driven causal discovery and statistical causal inference algorithms.

[0055] 3) By considering the characteristic variables screened by causal relationships, the present invention can more accurately model the causal effects between geographical variables. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the technical process of the present invention.

[0057] Figure 2 It is a causal skeleton diagram and a causal directed acyclic graph schematically shown in an embodiment of the present invention, wherein (a) is the causal skeleton diagram and (b) is the causal directed acyclic graph.

[0058] Figure 3 It is a schematic diagram of a device for optimizing the selection of variables of a deep learning model considering geographical causal relationships provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0059] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and should not be used to limit the protection scope of the present invention.

[0060] The object of the present invention is to provide a method and device for optimizing the selection of variables in a deep learning model that takes into account geographical causal relationships. By simulating expert knowledge through a large language model for automated causal order inference, the direct dependence on expert knowledge is reduced, and the efficiency of causal reasoning is improved. By integrating knowledge-driven causal discovery and statistical causal inference algorithms, the interpretability and accuracy in the process of geographical variable causal discovery are improved. By considering feature variables screened by causal relationships, the causal effects between geographical variables can be modeled more accurately.

[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] As Figure 1 shown, the present invention provides a method for optimizing the selection of variables in a deep learning model that takes into account geographical causal relationships, including:

[0063] Step S1: Collect geographical knowledge texts, preprocess the data, and construct a geographical text knowledge base for continuing to pre-train an open-source large language model to obtain a first geographical large model.

[0064] Specifically, widely extract geographical knowledge text data from various public resources. The public resources include but are not limited to professional books, academic papers, industry reports, encyclopedia entries, etc. The geographical knowledge text data includes but is not limited to geographical entity concepts, classic theories, geographical phenomena, etc.

[0065] Preprocess the collected data to construct a geographical knowledge base. The preprocessing methods include but are not limited to text cleaning, text deduplication, text segmentation, etc.

[0066] Based on the geographical text knowledge base, continue to pre-train an open-source large language model to obtain a first geographical large model. After pre-training, the first geographical large model has the general knowledge ability of geographical domain knowledge.

[0067] Step S2: Construct a high-quality causal relationship data set containing geographical context information for fine-tuning the first geographical large model to obtain a second geographical large model.

[0068] Specifically, screen out text paragraphs including explicit causal relationship descriptions from the geographical text knowledge base, such as descriptions of geographical processes and geographical conclusions, extract the geographical context of events, and manually identify causal relationship pairs to construct a geographical causal relationship data set D. Each piece of data d in the data set D = {STA, GD, TAs, CRDs}, where:

[0069] STA = {Location, Time}, representing Spatial-Temporal Attributes, including Location (geographical position) and Time.

[0070] GD = domain, which is an identifier for the geographical discipline or research area (Geographical Domain) to which it belongs, such as urban planning, environmental science, etc.

[0071] TAs = {Attribute 1 , Attribute 2 ,..., Attribute k}, representing Thematic Attributes, including a set of specific geographical features, such as climate, terrain, etc.

[0072] CRDs = {CRD 1 , CRD 2 ,..., CRD n}, representing a set of Causal Relationship Descriptions, where CRD = {Cause; Effect; Description}, Cause represents the cause variable, Effect represents the result variable, and Description represents the description of the causal relationship.

[0073] Use the STA, GD, and TAs in D to provide context information for geographical problems, use CRDs for reasoning and analysis of geographical problems, construct question-answer pairs, and fine-tune the first geographical large model to obtain the second geographical large model.

[0074] In this example, assume we are studying the impact of urbanization on air quality. We can construct the following data set D entry:

[0075] STA: {Location: "Beijing", Time: "January 2023"}

[0076] GD: "Environmental Science"

[0077] TAs:

[0078] "PM2.5 Concentration": 40 micrograms per cubic meter

[0079] "Temperature": -5°C

[0080] "Relative Humidity": 40%

[0081] "Wind Speed": 2 m / s

[0082] CRDs:

[0083] CRD1: {Cause: "Increased industrial emissions"; Effect: "Elevated PM2.5 concentration";

[0084] Description: "The increase in industrial activities leads to an increase in the concentration of fine particulate matter in the air."}

[0085] CRD2: {Cause: "Elevated PM2.5 concentration"; Effect: "Deteriorated air quality";

[0086] Description: "The increase in the concentration of fine particulate matter will reduce the air quality index."}

[0087] Step S3: Receive user input including the set of variables to be studied V and associated data, identify the processing variable X and the target variable Y, and call the first geospatial large model to generate additional metadata for each variable to be studied in the set of variables to be studied V;

[0088] Specifically, the user inputs the set of variables to be studied and the associated data. In addition, the user also needs to input the current modeling research field and the geographical context description of the research area. The variables to be studied are terms of various types in the geographical research field, such as temperature, precipitation, etc. The variable set should satisfy causal sufficiency, that is, the direct cause variables of any two variables in the variable set exist in the set;

[0089] The user needs to specify which pair of variables is the processing variable X and the target variable Y. The X-Y variable pair is the variable for which the user expects to model the causal effect.

[0090] Call the first geospatial large model to generate additional metadata for each variable in the variable set and output it in the form of <variable name, additional metadata>. The additional metadata includes at least the semantic description of the geographical phenomenon represented by the variable.

[0091] In this example, if the user studies the urban heat island effect in the Shanghai area, a possible input is:

[0092] Variable set V: {Temperature (T), Green space coverage rate (GC), Building density (BD), Industrial emissions (IE), Population density (PD)}

[0093] Associated data: Include the observed values of the above variables in different cities and time periods.

[0094] Research field: Environmental science, especially urban climatology.

[0095] Research area geographical context description: For example, Shanghai, 2023.

[0096] Processing variable X: Industrial Emissions (IE)

[0097] Target variable Y: Temperature (T)

[0098] Call the first large-scale geographical model to generate additional metadata for each variable in the variable set V. For example:

[0099] <Temperature (T), "The average temperature in urban areas, affected by various factors, including green space coverage, building density, etc.">

[0100] <Green Space Coverage (GC), "The percentage of green area within urban areas, which has a positive effect on regulating temperature.">

[0101] <Building Density (BD), "The concentration of buildings per unit area, which may increase heat absorption and storage.">

[0102] <Industrial Emissions (IE), "Greenhouse gas emissions generated by industrial activities, which may increase the heat load in urban areas.">

[0103] <Population Density (PD), "The number of residents per unit area, which is related to energy consumption and heat generation.">

[0104] Step S4: Based on the second large-scale geographical model, determine the causal order among the variables to be studied in the variable set V to be studied;

[0105] Specifically, for the given set of geographical variables to be studied, construct all possible sets of triples, where each triple consists of different elements in the set, denoted as T = {<vi, vj, vk> | i, j, k ∈ {1, 2,..., S} and i ≠ j ≠ k}, where vi, vj, vk are the i-th, j-th, and k-th variables to be studied in it; S is the total number of elements in the set V;

[0106] Adopt the triple Prompt prompting technique to orient the nodes of each triple in the set based on the large-scale geographical model to form a directed acyclic graph subDAG;

[0107] In this example, given the variable set = {Temperature (T), Green Space Coverage (GC), Building Density (BD), Industrial Emissions (IE), Population Density (PD)}, construct all possible sets of triples T. Each triple t consists of different elements in it. For example, T = {<t1 = {T, GC, BD}>, t2 = {T, BD, IE}>,...}, and use the triple prompting words to orient the nodes in each triple in the set T through the second large-scale geographical model.

[0108] The triple prompt technique provided by the present invention uses the method of In-context Learning (ICL) during the prompt construction process. By injecting examples into the prompt, the large model is made to focus on the reasoning process. This method enables the model to have the ability of immediate learning without updating the model parameters. Table 1 below is an example of a prompt template:

[0109] Table 1

[0110]

[0111]

[0112] Specifically, through the majority voting mechanism, the sub-DAGs of all triples are merged to obtain the overall causal structure. For each pair of variables, count the number of times they are determined to have a causal relationship in all sub-DAGs, and select the direction with the most occurrences as the final causal direction. If there is a conflict (i.e., the voting counts for two or three possible edge orientations are the same), then the second large language model is used again through Chain-of-Thought (CoT) prompting to resolve the conflict and make a final decision.

[0113] In the example, assume that in the triple <T1 = {T, GC, BD}>, the model determines the causal relationships BD → T and GC → T. In the triple <T2 = {T, IE, PD}>, the model may determine the causal relationships IE → T and PD → T. Through majority voting, if the majority of sub-DAGs show that T is affected by BD and GC, the final causal directions will be determined as BD → T and GC → T.

[0114] Specifically, record the reasoning basis for the entire causal discovery process to ensure the transparency and traceability of the process, and provide detailed documentation for users. The reasoning basis includes the response output of the large language model during causal orientation. It includes the causal analysis results of all triples, the model response output, and the final causal structure diagram.

[0115] Step S5, construct a causal skeleton graph for the set of variables V to be studied, and orient the edges in the causal skeleton graph according to the causal order to obtain a causal directed acyclic graph DAG.

[0116] Specifically, according to the constraint-based PC (Peter-Clark) causal discovery algorithm, create a causal skeleton graph for the set of geographical variables, as shown in (a) below. The edges in the causal skeleton graph are undirected or partially directed. Use the causal order obtained in S4 to orient the causal skeleton graph to obtain the final causal directed acyclic graph DAG, as shown in (b) below. Figure 2 as shown in (a) below. The edges in the causal skeleton graph are undirected or partially directed. Orient the causal skeleton graph using the causal order obtained in S4 to obtain the final causal directed acyclic graph DAG, as shown in (b) below. Figure 2 as shown in (b) below.

[0117] In the example, an undirected edge is added for each variable pair in the variable set V. Using the conditional independence test of the PC algorithm, each pair of variables in the variable set V is tested to determine whether there is a direct causal relationship between them, and a causal skeleton graph is obtained. Using the causal order information obtained by the S4 algorithm, the edges in the causal skeleton graph are further oriented. The introduction of the PC algorithm increases the robustness of causal discovery.

[0118] Step S6: Screen the nodes in the causal directed acyclic graph DAG that satisfy the adjustment rule to obtain an additional variable set Z.

[0119] Specifically, screen the nodes in the causal directed acyclic graph that satisfy the adjustment criterion to form the variable set Z. Use the variable X and the variables in the set Z as input variables to train a deep learning model to calculate the causal effect through observational data:

[0120]

[0121] Among them, E[Y|X=x] is the observational expectation, and E[Y|do(X=x)] is the intervention expectation, that is, the causal effect.

[0122] For the path from X to Y, 1) if the nodes in the Z set block all non-causal paths from X to Y, and 2) the nodes in the Z set are not on the causal path from X to Y, then Z satisfies the adjustment criterion for (X, Y).

[0123] Step S7: Use the processed variable X and the additional variable set Z as the input variables of the deep learning model, and the target variable Y as the output variable of the deep learning model to complete the optimal selection of the variables of the deep learning model considering geographical causal relationships.

[0124] The present invention simulates expert knowledge through a large language model for automated causal order inference, reduces the direct dependence on expert knowledge, and improves the efficiency of causal reasoning; by integrating knowledge-driven causal discovery and statistical causal inference algorithms, it improves the interpretability and accuracy in the process of geographical variable causal discovery; by considering the characteristic variables screened by causal relationships, it can more accurately model the causal effect between geographical variables.

[0125] The present invention also provides a device for optimizing the selection of variables of a deep learning model considering geographical causal relationships. Specifically, as Figure 3 shown, the device includes:

[0126] A large model pre-training module for constructing a causal relationship data set containing geographical context information and fine-tuning the first geographical large model to obtain a second geographical large model;

[0127] A large model fine-tuning module, used to construct a causal relationship data set containing geographic context information, and fine-tune the first geographic large model to obtain a second geographic large model;

[0128] A user interaction module is used to receive user input including a set of variables to be studied V and associated data, specify a processing variable X and a target variable Y, and call the first geographic macromodel to generate additional metadata for each variable to be studied in the set of variables to be studied V;

[0129] A causal order generation module is used to determine the causal order between the variables to be studied in the variable set V to be studied based on the second geographic macromodel and using a triplet prompt word technique;

[0130] A causal graph generation module is used to construct a causal skeleton graph for the variable set V to be studied, and obtain a causal directed acyclic graph DAG by directing the edges in the causal skeleton graph through the causal order;

[0131] The additional variable screening module is used to screen the nodes that meet the adjustment rules in the causal directed acyclic graph DAG to obtain an additional variable set Z.

[0132] The technical solution of the above-mentioned deep learning model variable optimization selection device taking into account geographical causal relationships is similar to the aforementioned method and will not be repeated here.

[0133] Based on the same technical solution, the present invention also discloses a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, which, when executed by a computing device, enable the computing device to execute the above-mentioned deep learning model variable optimization selection method taking into account geographic causal relationships.

[0134] Based on the same technical solution, the present invention also discloses a computing device, including one or more processors, one or more memories and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the above-mentioned deep learning model variable optimization selection method taking into account geographic causal relationships.

[0135] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0136] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0139] Specific examples are used herein to illustrate the principles and embodiments of the present invention. The description of the above embodiments is only for helping to understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific embodiments and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for optimizing the selection of variables in a deep learning model taking into account geographical causal relationships, characterized in that: include: Step S1: Build a geographic professional knowledge base, pre-train the open source large language model, and obtain the first geographic large model; Step S2: construct a causal relationship dataset containing geographic context information, and fine-tune the first geographic big model to obtain a second geographic big model; Step S3: receiving user input including a variable set V to be studied and associated data, specifying a processing variable X and a target variable Y, and calling the first geographic macromodel to generate additional metadata for each variable to be studied in the variable set V to be studied; Step S4: determining the causal order between the variables to be studied in the variable set V to be studied based on the second geographic macro model; Step S5: construct a causal skeleton graph for the variable set V to be studied, and obtain a causal directed acyclic graph DAG by directing the edges in the causal skeleton graph through the causal order; Step S6: Filter the nodes in the causal directed acyclic graph DAG that satisfy the adjustment rules to obtain an additional variable set Z; Step S7: using the processing variable X and the additional variable set Z as input variables of the deep learning model, and the target variable Y as the output variable of the deep learning model, to complete the optimization selection of variables of the deep learning model taking into account the geographical causal relationship; The step S1 specifically includes: Geographic knowledge text data obtained from various public resources, including professional books, academic papers, industry reports and encyclopedia entries, and the geographic knowledge text data includes geographic entity concepts, classical theories and geographical phenomena; Preprocessing the acquired geographic knowledge text data to build a geographic knowledge base; wherein the preprocessing method includes text cleaning, text deduplication and text segmentation; Pre-training the open source language big model based on the geographic professional knowledge base to obtain a first geographic big model; The step S2 specifically includes: The text paragraphs containing clear causal relationship descriptions are selected from the geographic professional knowledge base, the geographic contexts are extracted, the causal relationship pairs are manually identified, and a causal relationship dataset D is constructed; each data d in the causal relationship dataset D = {STA, GD, TAs, CRDs}, where: Spatiotemporal attribute STA = {Locaion, Time}, Location is the geographical location, Time is the time; Geographical discipline or research field GD = domain, where domain is an identifier; Thematic attributes TAs = {Attribute1, Attribute2, ..., Attribute t }, Attribute k is the tth geographical feature; Causal relationship set CRDs = {CRD1,CRD2,...,CRD N }, where CRD n ={Cause n ;Effect n ;Description n }, Cause n The cause variable representing the nth causal relationship, Effect n Indicates the result variable of the nth causal relationship, Description n Represents the description of the nth causal relationship; STA, GD and TAs in D are used to provide contextual information for geographic questions, CRDs are used to reason and analyze geographic questions, question-answer pairs are constructed, the first geographic big model is fine-tuned, and the second geographic big model is obtained.

2. According to claim 1, a method for optimizing and selecting variables of a deep learning model taking into account geographical causal relationships is characterized in that: The step S3 specifically includes: Receive user input, the user input includes a set of variables V to be studied, associated data, and a description of the current modeling research field and the geographical context of the research area, wherein V should satisfy causal sufficiency, that is, the direct cause variables of any two variables to be studied in V exist in V; The user needs to specify which pair of variables is the processing variable X and the target variable Y to form an XY variable pair; the XY variable pair is the variable for which the user expects the causal effect to be modeled; The first geographic macromodel is called to generate additional metadata for each variable to be studied in the variable set V to be studied, and the output data is organized in the form of <variable name, additional metadata>; the additional metadata at least includes a semantic description of the geographic phenomenon represented by the variable.

3. The method for optimizing and selecting variables of a deep learning model taking into account geographical causal relationships according to claim 2, characterized in that: The step S4 comprises: For V, construct all possible triple sets T, T = {<vi,vj,vk> |i,j,k∈{1,2,…,S} and i≠j≠k}, where vi,vj,vk are the i,j,kth variables to be studied respectively; S is the total number of elements in the set V; Using the triple prompt word technique, based on the second geographic model, the causal relationship between the three variables to be studied in each triple in T is determined to form a directed acyclic graph subDAG; Through the majority voting mechanism, the directed acyclic graph subDAG of all triples is merged to comprehensively obtain the overall causal structure; the majority voting mechanism is to count the number of times the causal relationship between each two variables to be studied is determined as a causal relationship in all triples, and select the direction with the most occurrences as the final causal direction.

4. The method for optimizing and selecting variables of a deep learning model taking into account geographic causal relationships according to claim 1, characterized in that: The step S5 specifically includes: According to the constraint-based PC (Peter-Clark) causal discovery algorithm, a causal skeleton graph is created for the geographic variable set V; the edges in the causal skeleton graph are undirected or partially directed; The causal skeleton graph is oriented using the causal order obtained in S4 to obtain a final causal directed acyclic graph DAG.

5. The method for optimizing and selecting variables of a deep learning model taking into account geographic causal relationships according to claim 4, characterized in that: The adjustment rule in step S6 is: whether the node blocks all non-causal paths from X to Y and is not on the causal path from X to Y.

6. A device for implementing any of the methods of claims 1 to 5, characterized in that: The device comprises: The large model pre-training module is used to build a geographic professional knowledge base, pre-train the open source large language model, and obtain the first geographic large model; A large model fine-tuning module, used to construct a causal relationship data set containing geographic context information, and fine-tune the first geographic large model to obtain a second geographic large model; A user interaction module is used to receive user input including a set of variables to be studied V and associated data, specify a processing variable X and a target variable Y, and call the first geographic macromodel to generate additional metadata for each variable to be studied in the set of variables to be studied V; A causal order generation module is used to determine the causal order between the variables to be studied in the variable set V to be studied based on the second geographic macromodel and using a triplet prompt word technique; A causal graph generation module is used to construct a causal skeleton graph for the variable set V to be studied, and obtain a causal directed acyclic graph DAG by directing the edges in the causal skeleton graph through the causal order; The additional variable screening module is used to screen the nodes that meet the adjustment rules in the causal directed acyclic graph DAG to obtain an additional variable set Z.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that When the processor executes the computer program, the steps of the method described in any one of claims 1 to 5 are implemented.

8. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Causal medical knowledge graph construction method based on large-scale pre-training language model

    CN117668245A

  • Mechanistic causal reasoning for efficient analytics and natural language

    WO2021092099A1