Large model prompt project optimization system and method fusing domain knowledge graph
By introducing knowledge graph automatic extraction of field parameters and dynamic prompt template generation technology in the big model, combined with feedback-driven closed-loop optimization mechanism, the problem of low answer accuracy caused by the lack of professional constraints in the oil and gas exploration field is solved, and the continuous adaptation of professional constraints and language models is achieved, and the answer accuracy is improved.
Patent Information
- Application Number
- CN202510685760.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The large model has low answer accuracy due to lack of professional constraints in the field of oil and gas exploration. It is difficult for the existing technology to systematically integrate professional parameter constraints such as reservoir physical properties and geological structure, resulting in numerical deviations or logical errors in the model output.
The domain parameters are automatically extracted by the knowledge graph to generate a dynamic prompt template, and a feedback-driven closed-loop optimization mechanism is established. The analysis module extracts parameterized constraints through the entity relationship topology analysis algorithm. The template generation engine module uses the syntax tree dynamic recombination technology to process the constraint encoded signals to generate an enhanced prompt text flow. The large model interactive interface module receives the prompt text flow and loads the knowledge graph verification rule set. The feedback analysis module analyzes the model output results through the semantic error vector calculation algorithm. The optimization strategy module dynamically adjusts the reservoir physical parameter weight coefficient during the template generation process based on the reinforcement learning mechanism.
The continuous adaptation of the professional constraints of large models in the field of oil and gas exploration and language models is achieved, the accuracy of answers is improved, and the problems of poor interpretation, high iteration costs and weak field adaptability in the existing technology are solved.
Smart Images

Figure CN120196734A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image processing, and in particular to a large model prompting engineering optimization system and method integrating a domain knowledge graph. Background Art
[0002] In recent years, large language models represented by GPT have demonstrated powerful natural language processing capabilities in general fields, but have significant limitations in applications in professional fields such as oil and gas exploration. Traditional prompt engineering methods rely on manually designed fixed templates, which are difficult to systematically integrate into professional parameter constraints such as reservoir properties and geological structures, resulting in numerical deviations or logical errors in model output that are inconsistent with industry specifications.
[0003] In the existing technology, although knowledge graphs can store domain knowledge in a structured manner, their interaction with natural language generation models mostly remains at the static knowledge retrieval level and lacks a dynamic constraint injection mechanism. The field of oil and gas exploration is characterized by complex parameters and strict rules. For example, reservoir evaluation needs to consider the threshold relationship of multi-dimensional indicators such as porosity, permeability, and oil saturation at the same time, and general large models cannot spontaneously understand the inherent relationship between these professional constraints. Current solutions attempt to adapt to professional fields by fine-tuning model parameters, but face problems such as scarce training data and lagging knowledge updates, and it is difficult to cope with the evolving decision-making logic in exploration scenarios. On the other hand, although the rule-based post-processing verification method can correct some explicit errors, it destroys the end-to-end consistency of the generation process, resulting in a decrease in semantic coherence.
[0004] In the existing patent technology, it is proposed to link the knowledge graph entity to the question-answering system, but the problem of dynamic adaptation of constraints and prompt templates is not solved; reinforcement learning is used to optimize the dialogue strategy, but a closed-loop feedback mechanism for knowledge graph verification and template generation is not established. Industry practice shows that simply expanding the scale of training data or increasing manual verification links cannot fundamentally improve the reasoning reliability of large models in professional scenarios. These limitations lead to the existing systems having defects such as poor interpretability, high iteration cost, and weak domain adaptability in the actual application of oil and gas exploration, which seriously restricts the in-depth application of artificial intelligence technology in exploration decision support. Summary of the invention
[0005] In view of the above shortcomings of the prior art, the purpose of the present invention is to provide a large model prompt engineering optimization system and method that integrates domain knowledge graphs to solve the problem of low answer accuracy of large models in the field of oil and gas exploration due to the lack of professional constraints. The present invention automatically extracts domain parameters through knowledge graphs to generate dynamic prompt templates and establishes a feedback-driven closed-loop optimization mechanism.
[0006] The present invention provides a large model prompting engineering optimization system integrating domain knowledge graph, comprising: Parsing module, which extracts parametric constraint conditions through an entity relationship topology analysis algorithm to form a constraint encoding signal containing an entity attribute association matrix; Template generation engine module, which processes the constraint encoding signal using the syntax tree dynamic recombination technology, injects the constraint conditions into a preset template framework, and forms an enhanced prompt text stream with reservoir physical property parameter slots; Large model interaction interface module, which receives the enhanced prompt text stream and synchronously loads the knowledge graph verification rule set to generate a question and answer response data stream containing geological professional terms; Feedback analysis module, which receives the question and answer response data stream, and through the semantic error vector calculation algorithm, logically matches the model output result with the knowledge graph standard data to form a feedback signal containing semantic deviation metrics; Optimization strategy module, which receives the feedback signal and dynamically adjusts the weight coefficients of reservoir physical property parameters during the template generation process based on the reinforcement learning mechanism, forms a parameter optimization instruction signal, and transmits it to the parsing module to complete the iterative update of the constraint conditions.
[0007] In an embodiment of the present invention, the specific implementation method of the entity relationship topology analysis algorithm in the parsing module is as follows: a multi-dimensional topology network is constructed by combining the attribute similarity calculation of knowledge graph entity nodes and the relationship path weight assignment. Among them, the attribute similarity calculation uses a vector space model based on cosine similarity to perform embedding representation on the entity description text, and the relationship path weight assignment uses a decay function to dynamically adjust the multi-hop association path; the constraint encoding signal includes the sparse encoding format of the entity attribute association matrix and the normalized weight coefficient sequence. Among them, the sparse encoding format stores the combined relationship of entity types and attribute key-value pairs using a hash map, and the normalized weight coefficient sequence is calculated by the product of the node centrality index and the path decay value in the topology network. The parsing module filters out effective constraint conditions through a dynamic weight threshold filtering mechanism to form a hierarchical parametric constraint signal and transmits it to the template generation engine module.
[0008] In one embodiment of the present invention, when performing the dynamic restructuring technology of the syntax tree in the template generation engine module, dependency syntactic analysis is performed on the preset template framework to generate an initial syntax tree structure. The replaceable nodes in the syntax tree are located according to the parameter slot types in the constraint encoding signal. The bidirectional attention mechanism is used to calculate the semantic matching degree between the constraint conditions and the context. The parameter insertion position is dynamically selected and the syntax tree branches are reconstructed. The generation process of the enhanced prompt text stream includes the dynamic filling of parameter slots and the semantic coherence verification. Among them, the dynamic filling combines rule-based regular expression matching with neural network-based context prediction. The semantic coherence verification scores the fluency of the restructured prompt text through a pre-trained language model. When the score is lower than the set threshold, the syntax tree backtracking mechanism is triggered to reselect the insertion node. The template generation engine module encapsulates the verified prompt text stream into a structured instruction set and transmits it to the large model interaction interface module.
[0009] In one embodiment of the present invention, the method for constructing the knowledge graph verification rule set in the large model interaction interface module is to extract entity-relationship triples from the domain knowledge graph to form atomic verification units, construct composite verification rules by combining atomic units through logical operators, and use a rule engine to perform multi-level verification on the model output. The generation process of the question-and-answer response data stream includes real-time semantic constraint injection and multi-round dialogue state management. Among them, the semantic constraint injection is achieved by superimposing the knowledge graph subgraph embedding vector on the model input layer. The multi-round dialogue state management tracks the evolution path of the core parameters by maintaining the context entity relationship stack. The large model interaction interface module adopts a rule-triggered correction mechanism in the output stage. When a response that violates the knowledge graph constraints is detected, a predefined correction template is automatically called to overwrite and rewrite the key entity attributes.
[0010] In one embodiment of the present invention, when the feedback analysis module performs the semantic error vector calculation algorithm, the model output text is subjected to entity recognition and relationship extraction to obtain a structured semantic network, which is compared with the reference semantic network constructed from the knowledge graph standard data in terms of graph structure. The difference metric value is calculated through three dimensions: node matching degree, edge similarity, and path connectivity. The feedback signal of the semantic deviation metric includes fine-grained error classification information and correction priority marking. Among them, the error classification information distinguishes three types: entity missing, relationship contradiction, and numerical out-of-bounds. The correction priority marking is dynamically assigned according to the impact of the error on business decisions. The feedback analysis module traces the source of the deviation by constructing an error propagation graph and forms a feedback signal with causal chain annotation and transmits it to the optimization strategy module.
[0011] In one embodiment of the present invention, the specific implementation manner of the reinforcement learning mechanism in the optimization strategy module includes constructing a triple model of a state space, an action space, and a reward function. The state space is jointly defined by the semantic deviation type in the feedback signal and the historical optimization record. The action space corresponds to the adjustment direction and amplitude of the reservoir physical property parameter weight coefficient. The reward function performs multi-objective optimization by comprehensively considering the improvement amplitude of the question-answering accuracy and the parameter adjustment cost. The generation process of the parameter optimization instruction signal includes policy gradient update and exploration-exploitation balance control. The policy gradient update uses the proximal policy optimization algorithm to prevent training oscillation. The exploration-exploitation balance control adjusts the search range of the parameter space through a dynamic greedy policy. The optimization strategy module encodes the optimization instruction into a differential signal format and updates the parameters of the entity relationship topology analysis algorithm in the parsing module through the backpropagation link.
[0012] In one embodiment of the present invention, a constraint condition cache pool is set between the knowledge graph parsing module and the template generation engine module, which is used to store the effectively extracted historical constraint conditions and their application effect evaluation data. When the parsing module detects a new knowledge graph entity, it preferentially matches similar constraint patterns in the constraint condition cache pool. If the match is successful, it directly calls the cached constraint encoding signal. Otherwise, it starts the complete topology analysis process. The cache pool maintains the storage space using the least recently used (LRU) eviction policy and quickly retrieves the constraint conditions through the hash fingerprint technology. When the template generation engine module receives the cached constraint encoding signal, it synchronously loads the corresponding syntax tree recombination historical record to accelerate the template generation process.
[0013] In one embodiment of the present invention, the template generation engine module includes a multi-version hint template library, and each template version is associated with a combination of constraint conditions and syntax structure features in a specific business scenario. When receiving a new constraint encoding signal, it matches the optimal template version through a scenario classification model and performs domain adaptation fine-tuning on the template parameters based on the transfer learning mechanism. The multi-version template library uses a version control mechanism to manage the template evolution process. Each template version saves a complete snapshot of the syntax tree structure and performance evaluation metrics. When the optimization strategy module generates a parameter optimization instruction, it synchronously triggers the rollback test and optimal update mechanism of the template version.
[0014] In one embodiment of the present invention, the large model interaction interface module includes a knowledge graph real-time update interface. When there are changes in the entity attributes or relationship structures of the domain knowledge graph, it automatically triggers the incremental analysis process of the parsing module. The incremental analysis process identifies the changed area by comparing the knowledge graph version differences and only performs local topology analysis on the affected entity subgraph, generating a differential constraint encoding signal and merging it with the original constraint conditions. The real-time update interface is provided with a change impact evaluation unit. When detecting a modification of a key entity attribute, it immediately pauses the ongoing question-answering process and regenerates the hint template to ensure the timeliness of the output result.
[0015] The present invention also provides an optimization method for large model prompt engineering integrating domain knowledge graphs, including: S1: Extract parametric constraint conditions through an entity relationship topology analysis algorithm to form a constraint encoding signal containing an entity attribute association matrix; S2: Use a syntax tree dynamic recombination technique to process the constraint encoding signal, inject the constraint conditions into a preset template framework to form an enhanced prompt text stream with reservoir physical property parameter slots; S3: Call a GPT-like model according to the enhanced prompt text stream and synchronously load a knowledge graph verification rule set to generate a question-and-answer response data stream containing geological professional terms; S4: Process the question-and-answer response data stream through a semantic error vector calculation algorithm, perform logical matching between the model output result and the knowledge graph standard data to form a feedback signal containing a semantic deviation metric; S5: Process the feedback signal based on a reinforcement learning mechanism and dynamically adjust the weight coefficient of the reservoir physical property parameters in the template generation process to form a parameter optimization instruction signal and complete the update of the constraint conditions.
[0016] An optimization system and method for large model prompt engineering integrating domain knowledge graphs provided by the present invention, by constructing a closed-loop system for knowledge graph parsing and dynamic prompt generation, automatically extracts professional parameters in the oil and gas exploration field (such as reservoir physical property constraints), converts them into dynamic slots of a natural language prompt template, synchronously performs knowledge graph rule verification in the GPT model interaction, and optimizes the parameter weights through feedback analysis to achieve continuous adaptation of professional constraints and language models, thereby improving the answer accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is the system architecture diagram of the optimization system for large model prompt engineering integrating domain knowledge graphs; Figure 2 It is the method flow diagram of the optimization method for large model prompt engineering integrating domain knowledge graphs. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0020] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0021] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0022] Please refer to Figure 1 , which shows the large model prompt engineering optimization system integrating the domain knowledge graph of the present invention. The large model prompt engineering optimization system integrating the domain knowledge graph of the present invention includes a parsing module, a template generation engine module, a large model interaction interface module, a feedback analysis module, and an optimization strategy module. The parsing module extracts parameterized constraint conditions through an entity relationship topology analysis algorithm to form a constraint encoding signal including an entity attribute association matrix. The template generation engine module processes the constraint encoding signal using a syntax tree dynamic recombination technique, injects the constraint conditions into a preset template framework to form an enhanced prompt text stream with reservoir physical property parameter slots; the large model interaction interface module receives the enhanced prompt text stream and synchronously loads the knowledge graph verification rule set to generate a question-and-answer response data stream including geological professional terms; the feedback analysis module receives the question-and-answer response data stream, and through a semantic error vector calculation algorithm, logically matches the model output result with the knowledge graph standard data to form a feedback signal including a semantic deviation metric; the optimization strategy module receives the feedback signal and dynamically adjusts the weight coefficient of the reservoir physical property parameters in the template generation process based on a reinforcement learning mechanism to form a parameter optimization instruction signal and transmit it to the parsing module to complete the iterative update of the constraint conditions.
[0023] As Figure 1As shown in the figure, the present invention relates to a large model prompt engineering optimization system integrating domain knowledge graphs, which is characterized by including a parsing module, a template generation engine module, a large model interaction interface module, a feedback analysis module, and an optimization strategy module. The parsing module extracts parameterized constraint conditions from the oil and gas exploration knowledge graph through an entity relationship topology analysis algorithm. Specifically, it calculates the attribute similarity of entity nodes related to reservoir physical properties in the knowledge graph and analyzes the multi-hop relationship paths, constructs an associated topology network between entities. This module uses graph neural network-based embedding representation technology to convert professional parameters such as porosity and permeability into structured constraint conditions, forming a constraint encoding signal containing an entity attribute association matrix. The association matrix stores entity types, attribute key-value pairs, and their weight coefficients in the form of a sparse tensor, and filters out valid constraint subsets that match the current business scenario through a dynamic threshold filtering mechanism. The parsing module transmits the generated constraint encoding signal to the template generation engine module through the data bus to trigger the downstream processing flow. After receiving the constraint encoding signal, the template generation engine module adapts the preset general Prompt template for the domain using a syntax tree dynamic recombination technology. Its core lies in naturally integrating structured constraint conditions into the natural language prompt framework: first, it performs dependency syntactic analysis on the preset template to generate an initial syntax tree and identifies the key node positions where professional parameters can be inserted; then, according to the parameter types in the constraint encoding signal, it uses a bidirectional attention mechanism to calculate the semantic matching degree between the constraint conditions and the context, and dynamically selects the insertion positions of numerical parameters such as porosity thresholds and permeability ranges; finally, it generates an enhanced prompt text stream with dynamic slots through a syntax tree reconstruction algorithm. This text stream embeds reservoir physical property constraints while retaining the fluency of natural language. The template generation engine module encapsulates the processed enhanced prompt text stream into a structured instruction set and transmits it to the large model interaction interface module through the API interface. After receiving the enhanced prompt text stream, the large model interaction interface module, while calling a GPT-class large language model to generate a professional Q&A response, synchronously loads the knowledge graph verification rule set to implement double verification: on the one hand, it converts the reservoir physical property parameters in the prompt text into embedding vector constraints in the model input layer, and on the other hand, it performs a logical compliance check on the output result through the rule engine. When generating a Q&A response data stream containing geological professional terms, this module uses a multi-round dialogue state management technology to maintain the context entity relationship stack, ensuring that the parameter evolution path conforms to geological laws, and transmits the data stream to the feedback analysis module through the message queue.The feedback analysis module deeply analyzes the Q&A response data stream and quantifies the deviation between the model output and the standard data of the knowledge graph through the semantic error vector calculation algorithm: First, it uses entity recognition and relationship extraction techniques to convert the natural language response into a structured semantic network, and then compares the graph structure with the reference network constructed by the knowledge graph to generate error metric values from three dimensions: node matching degree, edge similarity, and path connectivity; The formed semantic deviation metric feedback signal not only contains classification information such as entity missing and relationship contradiction, but also attaches a correction priority mark. By constructing an error propagation graph to trace the source of the deviation, the feedback signal with causal chain annotation is finally transmitted to the optimization strategy module. The optimization strategy module establishes a dynamic optimization model based on the reinforcement learning mechanism. Its state space is jointly defined by the semantic deviation type in the feedback signal and the historical optimization record. The action space corresponds to the adjustment strategy of the reservoir physical property parameter weight coefficient, and the reward function conducts multi-objective optimization by comprehensively considering the improvement of Q&A accuracy and the cost of parameter adjustment; The parameter optimization instruction signal generated by this module adopts a differential coding format and updates the parameters of the entity relationship topology analysis algorithm in the parsing module through the backpropagation link, forming a closed-loop system from problem discovery to iterative update of constraint conditions. The five-level data transmission channels are constituted among the modules by constraint coding signals, enhanced prompt text streams, Q&A response data streams, feedback signals, and optimization instruction signals, realizing the full-link automation of knowledge graph parsing, prompt template generation, large model interaction, result analysis, and system optimization.
[0024] Furthermore, the specific implementation of the entity relationship topology analysis algorithm in the parsing module includes multi-dimensional feature extraction and dynamic weight assignment mechanism for knowledge graph entity nodes. The algorithm first partitions the oil and gas exploration knowledge graph into subgraphs, focusing on entity types related to reservoir evaluation (such as lithological units, logging curves, physical property parameters, etc.). It uses a graph embedding algorithm based on TransE to map entity nodes into a low-dimensional vector space, and establishes an entity attribute correlation matrix through cosine similarity calculation. Each matrix element represents the correlation strength between two entities in dimensions such as porosity distribution and permeability range. For multi-hop relationship path analysis, the algorithm adopts a path weight calculation model with a decay factor, setting that the path weight decays exponentially for each additional hop, and at the same time dynamically adjusts the propagation weights of key nodes in combination with node centrality indicators (such as betweenness centrality and closeness centrality). In the constraint condition generation stage, the algorithm implements a three-level filtering mechanism: the first level eliminates weakly related entity pairs based on the sparsity threshold of the attribute correlation matrix, the second level screens significant relationship chains through the cumulative value of path weights, and the third level applies a business rule library (such as reservoir classification criteria, physical property parameter threshold tables) for domain-specific filtering. The finally generated constraint encoding signal includes structured triples (subject entity, relationship type, constraint strength) and a dynamic weight coefficient sequence, where the weight coefficients are normalized to ensure the comparability of parameters with different dimensions. The algorithm is specifically designed with an incremental update mechanism. When new logging interpretation data or lithological analysis results are added to the knowledge graph, only local topological analysis is performed on the affected entity subgraph, and incremental constraint encoding signals are generated through difference comparison and weighted fusion with historical constraint conditions, thereby reducing the system calculation overhead.
[0025] In an embodiment of the present invention, the operation process and quality control mechanism of the syntactic tree dynamic recombination technology in the template generation engine module. This technology first performs in-depth semantic parsing on a preset general Prompt template, constructs an initial syntactic tree using a BERT-based dependency parser, and identifies key nodes (usually noun phrase modification positions or conditional clause insertion points) where domain parameters can be inserted. In the constraint injection phase, a bidirectional attention mechanism is used to achieve semantic alignment between structured parameters and natural language templates: convert the reservoir physical property parameters in the constraint encoding signal into feature vectors, perform similarity matching with the context embedding vectors of the syntactic tree nodes, and select the node position with the highest semantic compatibility to insert parameter slots. For numerical parameters (such as porosity > 15%), the system automatically generates conditional judgment sentences and embeds them into the syntactic tree; for enumerated parameters (such as lithology type), list sentences are used to expand the syntactic tree branches. A double verification mechanism is implemented during the syntactic tree reconstruction process: one is static verification, which checks the legality of the tree structure through a preset syntactic rule library to prevent the absence of predicates or misplacement of modifiers; the other is dynamic verification, which uses a pre-trained language model to score the fluency of the recombined prompt text. When the score is lower than the threshold, a backtracking mechanism is triggered to reselect the parameter insertion position. The generated enhanced prompt text stream adopts a hierarchical encapsulation structure: the outer layer is a natural language instruction, and the inner layer embeds machine-readable constraint parameter tags, enabling the large model to not only understand human-readable semantic requirements but also accurately parse structured constraint conditions. This module is also equipped with a multi-version template management function, stores optimized template variants according to different business scenarios (such as reservoir evaluation, well location deployment), and when receiving a new constraint encoding signal, matches the most suitable template version through a scenario classification model and performs parameter fine-tuning based on transfer learning to ensure a high degree of fit between prompt engineering and domain requirements.
[0026] Such as Figure 1As shown, the specific implementation method of the semantic error vector calculation algorithm in the feedback analysis module lies in establishing a multi-dimensional error quantification system and a deviation tracing mechanism. The algorithm first deeply analyzes the natural language responses output by the GPT model, uses an entity recognition model based on BiLSTM-CRF to extract key entities such as reservoir parameters and geological structures, and constructs a response semantic network through a graph attention network, where nodes represent entity attributes and edges represent the strength of the relationship between entities. Subsequently, the generated response semantic network is aligned with the reference semantic network constructed from the standard data of the knowledge graph, and a graph isomorphism detection algorithm is used to calculate the difference indicators in three dimensions: node matching degree, edge similarity, and path connectivity. The node matching degree is verified by both the cosine similarity of the attribute vectors and the type consistency. The edge similarity is calculated by combining the relationship type matching degree and the strength difference. The path connectivity evaluates the integrity of the logical chain between key entities. The calculated error metric values are normalized to generate a semantic error vector, which not only contains the overall deviation score but is also subdivided into three types of error components: entity omission (such as not mentioning the permeability parameter), relationship contradiction (such as a logical conflict between porosity and burial depth data), and numerical out-of-bounds (such as the oil saturation exceeding the physical limit of the formation). In the feedback signal generation stage, the system traces the root cause of the deviation by constructing an error propagation map: uses the backpropagation algorithm to analyze the propagation path of the error components in the semantic network, identifies the core nodes that trigger the chain error (usually misjudgment of key reservoir parameters), and attaches a correction priority label to each error component - dynamically adjusts the priority coefficient according to the importance of the parameter in the exploration decision (such as the influence weight of porosity on reserve calculation). The finally formed feedback signal adopts a hierarchical coding structure, with the bottom layer containing the original error data, the middle layer storing the results of the traceability analysis, and the top layer encapsulating the optimization strategy suggestions. When transmitted to the optimization strategy module through the message queue, it automatically triggers the weight adjustment process.
[0027] Specifically, the implementation architecture and dynamic adjustment strategy of the reinforcement learning mechanism in the optimization strategy module. This module constructs a Markov decision process model, whose state space is composed of multi-dimensional feature vectors, including the error type distribution in the current feedback signal, the execution effect record of historical optimization instructions, the version number of the knowledge graph constraint conditions, and the real-time load status of the template generation engine. The action space is defined as the adjustment operation of the reservoir physical property parameter weight coefficients in the parsing module, specifically including the amplitude gradient of weight increase / decrease (with 0.1 as the minimum unit) and the collaborative adjustment strategy of related parameter groups. The reward function is designed as a multi-objective optimization function, mainly including three calculation items: the core reward item is based on the improvement amplitude of the question-answering accuracy (calculated by comparing the change in the norm of the error vector before and after optimization), the auxiliary penalty item considers the incremental system calculation overhead caused by parameter adjustment (quantified according to CPU / GPU resource consumption indicators), and the stability constraint item evaluates the impact of weight fluctuations on the adaptability of historical templates. The policy update adopts the Proximal Policy Optimization (PPO) algorithm, which balances the exploration and exploitation relationship through importance sampling and trust region constraints, and at the same time introduces a curriculum learning mechanism to gradually expand the parameter search space - in the initial stage, only fine-tuning of core parameters such as porosity and permeability is allowed, and the adjustment permissions of secondary parameters such as sedimentary facies types and formation pressures are gradually opened as the number of training rounds increases. A safety verification mechanism is implemented during the optimization instruction generation process: when it is detected that the parameter weight adjustment may cause constraint conflicts (such as the porosity threshold exceeding the reasonable range of the associated burial depth), the Constraint Satisfaction Problem (CSP) solver is automatically triggered for feasibility verification, and only the optimization instructions that comply with all business rules are output. The finally generated parameter optimization instruction signal adopts a differential coding format, including the weight adjustment amount, the scope of action, and the effective time window information. When it is transmitted to the parsing module through the reverse data channel, the weight matrix in the entity relationship topology analysis algorithm is updated synchronously, and the version change log is recorded for tracing the optimization history.
[0028] Furthermore, the conditional cache pool module is used to optimize the data processing efficiency between the knowledge graph parsing module and the template generation engine module. The cache pool adopts a hierarchical storage architecture: the first layer is the hot constraint library, which stores the reservoir physical property constraint conditions frequently used recently and their associated syntax tree recombination schemes, and uses the LRU (Least Recently Used) algorithm for entry elimination management; the second layer is the pattern matching library, which classifies and stores the historical extraction constraint conditions according to business scenarios (such as reservoir classification and evaluation, well pattern deployment optimization) through clustering analysis, and each category is associated with a feature vector index; the third layer is the derivative rule library, which stores the constraint combination rules automatically generated by machine learning (such as "when porosity > 12% and permeability > 50 mD, oil-water interface analysis must be associated"). When the parsing module starts a new topology analysis task, first match the hash fingerprint of the current knowledge graph subgraph with the cache pool: calculate the Hamming distance of the subgraph feature fingerprint using the SimHash algorithm, and if it is within the threshold range, directly call the constraint encoding signal and the associated template generation record in the cache, otherwise execute the complete analysis process. When the cache is hit, the template generation engine module will load the historical syntax tree recombination scheme as the baseline template and then make incremental adjustments according to the current business requirements, which can reduce the template generation time by more than 40%. During the maintenance of the cache pool, a data integrity protection mechanism is implemented: each cache entry is associated with a confidence score (calculated based on the historical call times and the successful application ratio), and when the score is lower than the threshold, the re-verification process is automatically triggered to confirm the cache validity by comparing with the latest knowledge graph data. In addition, an anomaly detection unit is set up. When it is detected that the cache call causes an abnormal decrease in the Q&A accuracy rate, the relevant cache entries are immediately cleared and the constraint conditions are regenerated to ensure the reliability of the system output.
[0029] such as Figure 1As shown, the multi-version prompt template management function of the template generation engine module constructs an adaptive template system for different business scenarios. The system maintains a dynamically evolving template library, and each template version contains complete metadata: creation time, applicable scenario tags (such as lithology identification, production capacity prediction), associated constraint condition fingerprints, historical call statistics, and performance evaluation metrics (including generation response accuracy, semantic fluency score). When receiving a new constraint encoding signal, the scenario classification model (trained based on the XGBoost algorithm) will analyze the entity type distribution, parameter weight configuration, and business objective description contained in the signal, and match the most suitable existing template version as the basic framework. For scenarios with insufficient matching degree, start the template transfer learning process: select old templates with similar topological structures, retain the main structure of their syntax trees, replace the parameter slots in the leaf nodes to adapt to the new constraint conditions, and at the same time use a generative adversarial network (GAN) to polish the reorganized templates in natural language. The template version control adopts a Git-style management mechanism, records the differential patches of each syntax tree modification, and allows for a quick rollback to a historical stable version after the optimization strategy is adjusted. When the optimization strategy module issues a parameter optimization instruction, it synchronously triggers a backtracking test of the template version: run the new and old template versions in parallel in a sandbox environment, compare the question-and-answer response quality metrics, and only complete the version upgrade when the new template performs better in all test cases. In addition, establish template life cycle management rules, automatically archive stale versions that have not been called for three consecutive months, releasing storage resources while retaining key syntax tree features for future template reconstruction.
[0030] Furthermore, integrate the real-time update interface of the knowledge graph with the incremental analysis mechanism to construct a dynamically evolving constraint condition generation system. When entity attributes are updated or the relationship structure is adjusted in the oil and gas exploration knowledge graph, the real-time update interface automatically identifies the changed area through the version comparison engine: adopting a three-stage processing flow based on the graph difference algorithm (GraphDiff). First, locate the newly added / deleted entities through node fingerprint matching (using SHA-256 hash values to represent entity attributes). Second, discover topological structure changes through relationship path comparison. Finally, identify key parameter modifications through attribute value fluctuation detection (setting a threshold for the change rate of physical property parameters). For the detected changes, the system starts the incremental analysis process - only perform local topological analysis on the subgraph where the affected entities are located (determine the scope of influence through k-hop neighbor expansion), and use a lightweight graph neural network (GNN) to quickly calculate the entity association degree in the changed area, generating a differential constraint encoding signal. This signal is integrated with the original global constraint conditions through a weighted fusion algorithm: set a decay coefficient for historical constraints (linearly reduce the weight over time), assign an initial weight to newly added constraints, and resolve the logical contradiction between new and old constraints through a conflict detection mechanism (such as automatically discarding obsolete constraints when the porosity threshold is updated). The real-time update interface is specifically set with a change impact assessment unit, which uses a causal inference model to predict the impact of knowledge graph modifications on the existing template library: when a change in key reservoir parameters (such as the reference value of permeability) is detected, immediately send a template failure warning to the template generation engine module, triggering the re-generation process of relevant prompt templates; for non-critical parameter changes, use a background asynchronous update method to reduce system response latency. This mechanism ensures that in the high-frequency scenario where logging interpretation data is updated hourly, the system can complete constraint condition refreshing within an average of 300 ms, while maintaining the consistency of prompt templates with the latest industry standards.
[0031] Such as Figure 2As shown in the figure, the security guarantee system of the optimization strategy module introduces a sandbox test environment and an abnormal fusing mechanism. The sandbox environment clones the complete configuration of the current production system through containerization technology (including the weight matrix of the parsing module, the syntax tree library of the template generation engine, and the rule set of the large model interaction interface), and constructs a virtual test instance isolated from the production environment. When the optimization strategy module generates a new parameter optimization instruction, the instruction is first injected into the sandbox environment, and a batch test is carried out by loading the Q&A interaction data set within the past three months (including more than 5,000 labeled cases): the semantic error vector change, system resource consumption indicators, and the number of constraint conflicts of each test case are recorded during the execution process, and a multi-dimensional performance evaluation report is generated. The test result analysis uses a difference comparison algorithm to calculate the relative change rate of key indicators (such as average Q&A accuracy rate, maximum response delay) before and after optimization. When it is detected that the accuracy improvement is less than 5% or the resource consumption increases by more than 20%, it is marked as a sub-optimal strategy. In the security verification stage, three-level fusing thresholds are set: the primary threshold triggers strategy fine-tuning (retesting with a reduced parameter adjustment range), the intermediate threshold starts the manual review process (sending abnormal cases to the expert interface), and the high-level threshold directly discards the current optimization instruction and rolls back to the previous stable version. The sandbox environment is specially designed for adversarial test scenarios, and boundary condition cases (such as extreme reservoir parameter combinations) are generated through genetic algorithms to verify the robustness of the optimization strategy under abnormal conditions. The optimized instructions passing the test are encapsulated into a versioned update package, including an incremental deployment script (only modifying the configuration parameters of the affected modules) and a rollback snapshot, and are deployed to the production system through a secure channel after digital signature authentication. This mechanism enables the system to maintain 99.9% service availability during the parameter optimization process, and at the same time shortens the fault recovery time caused by misconfigurations to within 2 minutes.
[0032] As Figure 2 shown, it is a large model prompt engineering optimization method integrating a domain knowledge graph according to the present invention. It includes S1: extracting parameterized constraint conditions through an entity relationship topology analysis algorithm to form a constraint encoding signal including an entity attribute association matrix; S2: processing the constraint encoding signal by using a syntax tree dynamic recombination technique, injecting the constraint conditions into a preset template framework to form an enhanced prompt text stream with reservoir physical property parameter slots; S3: calling a GPT-like model according to the enhanced prompt text stream and synchronously loading a knowledge graph verification rule set to generate a Q&A response data stream including geological professional terms; S4: processing the Q&A response data stream by using a semantic error vector calculation algorithm, logically matching the model output result with the knowledge graph standard data to form a feedback signal including a semantic deviation metric; S5: processing the feedback signal based on a reinforcement learning mechanism and dynamically adjusting the weight coefficient of the reservoir physical property parameters in the template generation process to form a parameter optimization instruction signal, completing the update of the constraint conditions.
[0033] As Figure 2As shown in the figure, the multi-modal data fusion processing ability expands the input source and semantic understanding dimension of the knowledge graph parsing module. The system newly adds a logging curve parsing unit and a geological map recognition unit: The logging curve parsing unit uses wavelet transform to extract the characteristic segments of curves such as gamma and resistivity, and identifies reservoir interfaces and abnormal points of physical property parameters through a temporal convolutional network (TCN), and converts them into event entities in the knowledge graph (such as "The GR curve shows a box-shaped mutation at 3150m"); The geological map recognition unit integrates the pre-trained YOLOv7 model to detect elements such as fault lines and trap boundaries on the structural map, and uses a graph neural network (GNN) to construct a spatial topological relationship to generate geological structure triples with coordinate annotations (such as "Fault F1 cuts trap T2 at 108.7° east longitude"). The multi-modal data is fused with the text knowledge graph through a cross-modal alignment module: A contrastive learning framework is used to train a shared embedding space to make the logging curve feature vectors, geological map region vectors and text entity vectors comparable, and then establish cross-modal associations (such as automatically linking the resistivity curve peak with the "high permeability zone" entity in the knowledge graph). The parsing module is upgraded to a multi-modal constraint generation engine, which can not only process text-based knowledge graphs, but also extract spatial constraints from image data (such as "The distance between the well location deployment point and the fault > 200m") and derive dynamic constraints from temporal data (such as "The rising rate of water cut < 2% / month"). The generated enhanced constraint encoding signal adds a new modal marking dimension to guide the template generation engine module to create multi-modal prompt templates: Embed data reference marks in the text instructions (such as "Combined with the structural features of the attached Figure 2 ), and drive the large model interaction interface to synchronously load the corresponding logging curve segments or geological map regions for multi-modal reasoning. The feedback analysis module correspondingly upgrades its multi-modal error detection ability, which not only verifies the logic of the text response, but also checks the spatial consistency between the recommended plan and the geological map (such as whether the proposed well location crosses the fault), as well as the temporal matching degree with the logging curve (such as whether the predicted water cut change trend matches the curve shape), forming a multi-dimensional semantic error vector that combines text, space, and time, so that the optimization strategy module can implement precise parameter adjustment according to different modal characteristics.
[0034] The large model prompt engineering optimization system integrating domain knowledge graph of the present invention automatically extracts professional parameters in the oil and gas exploration field (such as reservoir physical property constraints) by constructing a closed-loop system of knowledge graph parsing and dynamic prompt generation, converts them into dynamic slots of natural language prompt templates, synchronously performs knowledge graph rule verification during the interaction with the GPT model, and optimizes the parameter weights through feedback analysis to achieve continuous adaptation of professional constraints and language models, thereby improving the answer accuracy.
[0035] Therefore, through the large model prompt engineering optimization system and method integrating domain knowledge graphs of the present invention, the problem of low answer accuracy of large models in the oil and gas exploration field due to the lack of professional constraints can be solved.
[0036] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A large model prompt engineering optimization system integrating domain knowledge graphs, characterized in that Including: A parsing module that extracts parameterized constraint conditions through an entity relationship topology analysis algorithm to form a constraint encoding signal containing an entity attribute association matrix; A template generation engine module that processes the constraint encoding signal using a syntax tree dynamic recombination technique, injects the constraint conditions into a preset template framework to form an enhanced prompt text stream with reservoir physical property parameter slots; A large model interaction interface module that receives the enhanced prompt text stream and synchronously loads a knowledge graph verification rule set to generate a question and answer response data stream containing geological professional terms; A feedback analysis module that receives the question and answer response data stream, performs logical matching between the model output result and the knowledge graph standard data through a semantic error vector calculation algorithm to form a feedback signal containing a semantic deviation metric; An optimization strategy module that receives the feedback signal and dynamically adjusts the weight coefficient of the reservoir physical property parameters in the template generation process based on a reinforcement learning mechanism, forms a parameter optimization instruction signal and transmits it to the parsing module to complete the iterative update of the constraint conditions.
2. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that The specific implementation method of the entity relationship topology analysis algorithm in the parsing module is: constructing a multi-dimensional topology network by combining the attribute similarity calculation of knowledge graph entity nodes and the relationship path weight assignment. Among them, the attribute similarity calculation uses a vector space model based on cosine similarity to perform embedding representation on the entity description text, and the relationship path weight assignment uses a decay function to dynamically adjust the multi-hop association path; the constraint encoding signal contains the sparse encoding format of the entity attribute association matrix and the normalized weight coefficient sequence. Among them, the sparse encoding format uses a hash mapping table to store the combined relationship of entity types and attribute key-value pairs, and the normalized weight coefficient sequence is calculated by the product of the node centrality index and the path decay value in the topology network. The parsing module filters out effective constraint conditions through a dynamic weight threshold filtering mechanism to form a hierarchical parameterized constraint signal and transmits it to the template generation engine module.
3. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, When performing the syntax tree dynamic recombination technique in the template generation engine module, perform dependency syntactic analysis on the preset template framework to generate an initial syntax tree structure, locate replaceable nodes in the syntax tree according to the parameter slot type in the constraint encoding signal, use a bidirectional attention mechanism to calculate the semantic matching degree between the constraint conditions and the context, dynamically select the parameter insertion position and reconstruct the syntax tree branch; the generation process of the enhanced prompt text stream includes the dynamic filling of parameter slots and semantic coherence verification. Among them, the dynamic filling uses a combination of rule-based regular expression matching and neural network-based context prediction, and the semantic coherence verification scores the fluency of the recombined prompt text through a pre-trained language model. When the score is lower than the set threshold, trigger the syntax tree backtracking mechanism to reselect the insertion node. The template generation engine module encapsulates the verified prompt text stream into a structured instruction set and transmits it to the large model interaction interface module.
4. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, The construction method of the knowledge graph verification rule set in the large model interaction interface module is as follows: extract entity-relationship triples from the domain knowledge graph to form atomic verification units, combine atomic units through logical operators to construct composite verification rules, and use a rule engine to perform multi-level verification on the model output. The generation process of the Q&A response data stream includes real-time semantic constraint injection and multi-round dialogue state management. Among them, semantic constraint injection is achieved by superimposing knowledge graph subgraph embedding vectors on the model input layer, and multi-round dialogue state management tracks the evolution path of core parameters by maintaining a context entity relationship stack. The large model interaction interface module adopts a rule-triggered correction mechanism in the output stage. When a response that violates the knowledge graph constraint is detected, a predefined correction template is automatically called to overwrite and rewrite the key entity attributes.
5. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, When the feedback analysis module performs the semantic error vector calculation algorithm, it performs entity recognition and relationship extraction on the model output text to obtain a structured semantic network, and compares the graph structure with the reference semantic network constructed from the knowledge graph standard data. The difference metric value is calculated through three dimensions: node matching degree, edge similarity, and path connectivity. The feedback signal of the semantic deviation metric includes fine-grained error classification information and a correction priority mark. Among them, the error classification information distinguishes three types: entity missing, relationship contradiction, and numerical out-of-bounds. The correction priority mark is dynamically assigned according to the impact of the error on business decisions. The feedback analysis module traces the source of the deviation by constructing an error propagation graph and forms a feedback signal with causal chain annotation to transmit to the optimization strategy module.
6. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that The specific implementation method of the reinforcement learning mechanism in the optimization strategy module includes constructing a triple model of the state space, action space, and reward function. Among them, the state space is jointly defined by the semantic deviation type in the feedback signal and the historical optimization record. The action space corresponds to the adjustment direction and amplitude of the reservoir physical property parameter weight coefficient. The reward function performs multi-objective optimization by comprehensively considering the improvement amplitude of the Q&A accuracy rate and the parameter adjustment cost. The generation process of the parameter optimization instruction signal includes policy gradient update and exploration-exploitation balance control. The policy gradient update uses the proximal policy optimization algorithm to prevent training oscillation. The exploration-exploitation balance control adjusts the search range of the parameter space through a dynamic greedy strategy. The optimization strategy module encodes the optimization instruction into a differential signal format and updates the parameters of the entity relationship topology analysis algorithm in the parsing module through the backpropagation link.
7. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, A constraint condition cache pool is set between the knowledge graph parsing module and the template generation engine module, which is used to store the effective constraint conditions extracted historically and the evaluation data of their application effects. When the parsing module detects a new knowledge graph entity, it preferentially matches similar constraint patterns in the constraint condition cache pool. If the match is successful, the cached constraint encoding signal is directly called; otherwise, a complete topology analysis process is started. The cache pool maintains the storage space using the least recently used (LRU) eviction policy and quickly retrieves the constraint conditions through the hash fingerprint technology. When the template generation engine module receives the cached constraint encoding signal, it synchronously loads the corresponding syntax tree recombination historical records to accelerate the template generation process.
8. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, The template generation engine module includes a multi-version prompt template library. Each template version is associated with a combination of constraint conditions and syntax structure features in a specific business scenario. When a new constraint encoding signal is received, the optimal template version is matched through a scenario classification model, and the template parameters are fine-tuned for domain adaptation based on the transfer learning mechanism. The multi-version prompt template library uses a version control mechanism to manage the template evolution process. Each template version saves a complete snapshot of the syntax tree structure and performance evaluation metrics. When the optimization strategy module generates a parameter optimization instruction, the rollback test and best-choice update mechanism of the template version are synchronously triggered.
9. The large model prompt engineering optimization system integrating domain knowledge graphs according to claim 1, characterized in that, The large model interaction interface module includes a real-time update interface for the knowledge graph. When there are changes in entity attributes or relationship structures in the domain knowledge graph, the incremental analysis process of the parsing module is automatically triggered. The incremental analysis process identifies the changed areas by comparing the knowledge graph version differences, and only performs local topology analysis on the affected entity subgraphs, generating difference constraint encoding signals and merging them with the original constraint conditions. The real-time update interface is provided with a change impact assessment unit. When a critical entity attribute modification is detected, the ongoing question-and-answer process is immediately paused and a new prompt template is regenerated to ensure the timeliness of the output results.
10. A large model prompt engineering optimization method for a large model prompt engineering optimization system using the fusion domain knowledge graph according to any one of claims 1-9, characterized in that, Including: S1: Extract parameterized constraint conditions through the entity relationship topology analysis algorithm to form a constraint encoding signal containing an entity attribute association matrix. S2: Use the syntax tree dynamic recombination technology to process the constraint encoding signal, inject the constraint conditions into a preset template framework to form an enhanced prompt text stream with reservoir physical property parameter slots. S3: Call a GPT-like model according to the enhanced prompt text stream and synchronously load the knowledge graph verification rule set to generate a question-and-answer response data stream containing geological professional terms. S4: Process the question-and-answer response data stream through the semantic error vector calculation algorithm, and logically match the model output result with the knowledge graph standard data to form a feedback signal containing semantic deviation metrics. S5: Process the feedback signal based on the reinforcement learning mechanism and dynamically adjust the weight coefficients of the reservoir physical property parameters in the template generation process to form a parameter optimization instruction signal, completing the update of the constraint conditions.
Citation Information
Patent Citations
Data display method, electronic device and storage medium
CN110209766A
Knowledge graph and rule constraint combined data intelligent analysis method and system
CN118606440A
DCS intelligent decision-making method and system fusing large language model and knowledge graph
CN118820778A
File analysis method for photovoltaic measurement and calculation and related device
CN119150843A
Operation and maintenance alarm processing method and system based on knowledge graph enhanced large model
CN119988154A
Cited By
Zero-sample named entity recognition method and device based on large model feedback optimization
CN120354855A
Zero-shot named entity recognition method and device based on large model feedback optimization
CN120354855B
Clinical information extraction method, system and equipment based on large language model and medium
CN120448551A
Visual setting calculation system and method thereof
CN120494072A
A visual tuning calculation system and method thereof
CN120494072B