Rag and preference alignment collaborative optimization method and system for power field
By constructing a collaborative optimization method for RAG and preference alignment in the power sector, acquiring and processing a power business decision knowledge base, and combining a large language model to generate power business decision problems and expert decision texts, and performing weighted fusion fine-tuning training, the problems of low update efficiency and insufficient professional knowledge coverage of power decision models are solved, and efficient and reliable power system decision support is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-14
AI Technical Summary
Decision-making models in the power sector rely on manually constructed structured rule bases and expert experience models, which make it difficult to quickly transform unstructured knowledge into machine-executable logic. They also suffer from low update efficiency. General-purpose large language models lack professional knowledge coverage and suffer from illusion problems, making it difficult to meet the dynamic optimization needs of power grid operation.
By constructing a collaborative optimization method for RAG and preference alignment in the power sector, a knowledge base for power business decisions is obtained, and text is segmented and vectorized. Combined with a large language model, power business decision-making questions and expert decision texts are generated. The target large language model is then fine-tuned by a weighted fusion of direct preference optimization and probability ratio preference optimization.
It achieves high-quality professional knowledge support and alignment with expert decision-making preferences, improves the professionalism and reliability of the model, solves the problems of low update efficiency, illusion phenomenon and poor adaptability of traditional models, and meets the real-time decision-making needs of power systems.
Smart Images

Figure CN121980039B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for collaborative optimization of RAG and preference alignment in the power sector. Background Technology
[0002] In the operation control, state management, and daily operation of power systems, the reliability and agility of decision-making models directly determine the safety, stability, and economy of power grid operation. Currently, decision support in the power sector still relies heavily on manually constructed structured rule bases and expert experience models. While these models possess a certain degree of static reliability, they struggle to quickly translate massive amounts of unstructured dispatching procedures, equipment manuals, fault reports, and other professional knowledge into machine-executable logic. Furthermore, the manual coding and updating methods typically consume considerable time, failing to adapt to the complex and ever-changing dynamic optimization needs of the power grid operating environment and hindering real-time decision-making and agile control.
[0003] With the development of generative artificial intelligence, Large Language Models (LLMs), with their powerful natural language understanding and instruction-following capabilities, have provided new possibilities for the efficient processing of unstructured knowledge and the automated construction of decision-making models in the power sector. However, the high professional barriers and safety-first core requirements of the power system pose multiple challenges to the direct application of general-purpose LLMs: First, the pre-training data for general-purpose LLMs lacks sufficient coverage of power-related professional knowledge, making it prone to biases in the semantic parsing of industry-specific terminology and complex rules; second, the inherent illusion problem of LLMs is particularly harmful in power scenarios, potentially fabricating rule exceptions or incorrectly associating business conditions, leading to serious decision-making errors; third, the probabilistic generative nature of LLMs fundamentally contradicts the requirement for strict determinism in power decision-making, making it difficult to guarantee the professional consistency and reliability of the generated content. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides a collaborative optimization method and system for RAG and preference alignment in the power sector. By leveraging technological collaboration, it overcomes the limitations of both RAG and preference alignment, constructing a dedicated decision-making model for the power sector that combines high-quality knowledge support with accurate preference alignment.
[0005] In a first aspect, embodiments of the present invention provide a collaborative optimization method for RAG and preference alignment in the power sector, comprising:
[0006] Acquire several documents related to power business decisions to obtain a power business decision knowledge base;
[0007] The power business decision knowledge base is segmented into text blocks, and each text block is generated through vectorization representation and retrieval enhancement to obtain power business-adapted text.
[0008] Based on the power business adaptation text, combined with the pre-called first major language model and retrieval enhancement, several power business decision questions are generated, as well as the expert decision text and ordinary decision text corresponding to each power business decision question.
[0009] Each power business decision problem is associated and combined with the corresponding expert decision text and ordinary decision text in a triplet structure to obtain a preference dataset;
[0010] Based on the aforementioned preference dataset, the pre-trained second-largest language model is fine-tuned using a weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the target large language model.
[0011] Preferably, the step of acquiring several documents related to power business decisions to obtain a power business decision knowledge base includes:
[0012] Several PDF documents related to power business decisions were collected from multiple publicly available data sources;
[0013] Each PDF document is converted to its corresponding Markdown document.
[0014] Each Markdown document is cleaned separately, and the cleaned Markdown documents are combined to form a power business decision knowledge base.
[0015] Preferably, the step of segmenting the power business decision knowledge base into text blocks and generating power business-adapted text by vectorizing and enhancing the representation of each text block includes:
[0016] The power business decision knowledge base is divided into text blocks according to document length to obtain several text blocks;
[0017] Each text block is vectorized using a preset text embedding model to obtain a vector representation of the corresponding text block, and the vector representations are combined to form a power business vector knowledge base;
[0018] Based on the aforementioned power business vector knowledge base, power business-adapted text is generated through enhanced retrieval.
[0019] Preferably, the step of generating several power business decision-making problems and corresponding expert decision-making texts and ordinary decision-making texts based on the power business adapted text, combined with the pre-invoked first large language model and retrieval enhancement, includes:
[0020] Based on the power business adaptation text, several power business decision questions are generated by combining the preset question generation template and the pre-called first major language model.
[0021] For each of the aforementioned power business decision problems, the expert decision text corresponding to the power business decision problem is generated by combining the first large language model with retrieval enhancement.
[0022] For each of the aforementioned power business decision problems, a general decision text corresponding to the power business decision problem is generated by combining the second major language model in the initial state.
[0023] Preferably, the step of fine-tuning the pre-trained second large language model based on the preference dataset using a weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the target large language model includes:
[0024] The preference dataset is divided into a training set, a test set, and a validation set according to a preset ratio;
[0025] Based on the training set, the pre-trained second large language model is fine-tuned by a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model.
[0026] The hyperparameters of the initial large language model are adjusted in real time based on the validation set to obtain an intermediate large language model;
[0027] The intermediate large language model is evaluated using the test set to obtain the target large language model.
[0028] Preferably, the step of fine-tuning the pre-trained second large language model based on the training set using a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model includes:
[0029] Based on the training set, the pre-trained second large language model is fine-tuned using a fusion loss function consisting of a dynamic weighted sum of direct preference optimization loss and probability ratio preference optimization loss, to obtain the initial large language model.
[0030] Preferably, the dynamic weighting factor of the direct preference optimization loss is determined by the basic weighting factor adjusted in segments based on the number of training steps, the model performance deviation during the training process, and the value range constraints.
[0031] Preferably, the adjustment strategy for the basic weighting factor includes:
[0032] The number of training steps is divided into a warm-up period, an ascent period, and a stabilization period according to a preset stage threshold.
[0033] When the training process is in the warm-up period, the basic weight factor gradually increases from the weight at the beginning of the warm-up period to the weight at the end of the warm-up period.
[0034] When the training process is in the rising phase, the basic weight factor gradually increases from the end of the warm-up period to the peak weight.
[0035] When the training process is in a stable period, the basic weight factor gradually decreases from the weight peak to the balanced weight at the end of the training.
[0036] Preferably, the step of evaluating the performance of the intermediate large language model using the test set to obtain the target large language model includes:
[0037] The intermediate large language model is evaluated using the test set from multiple dimensions, including semantic consistency, professional reliability, and preference alignment, to obtain the target large language model.
[0038] Secondly, embodiments of the present invention provide a RAG and preference alignment collaborative optimization system for the power sector, comprising:
[0039] The data acquisition and processing module is used to acquire several documents related to power business decisions and obtain a power business decision knowledge base.
[0040] The vectorization and RAG module is used to segment the power business decision knowledge base into text blocks, and generate power business-adapted text by vectorizing each text block and enhancing retrieval through vectorization representation and retrieval.
[0041] The problem text generation module is used to generate several power business decision questions and corresponding expert decision text and ordinary decision text based on the power business adaptation text, combined with the pre-called first major language model and retrieval enhancement;
[0042] The association and combination module is used to associate and combine each power business decision problem with the corresponding expert decision text and ordinary decision text according to the triplet structure to obtain a preference dataset;
[0043] The fine-tuning training module is used to fine-tune the pre-trained second large language model based on the preference dataset, using a weighted fusion of direct preference optimization and probability ratio preference optimization, to obtain the target large language model.
[0044] Compared with existing technologies, the RAG and preference alignment collaborative optimization method and system of this invention for the power sector has the following advantages at least one point:
[0045] (1) For the first time, a closed-loop link of "structured knowledge base accumulation - adaptive generation of professional data - weighted fusion and fine-tuning of dual algorithms" is constructed, which deeply couples the factual support capability of enhanced retrieval with the value orientation capability of preference alignment. Unlike the shortcomings of single RAG technology, which lacks expert preference adaptation and single preference alignment technology, which is insufficient in generalization, this invention provides a high-quality professional data foundation covering power business through the accurate construction of power business decision knowledge base and adaptive text segmentation. After the generation of triplet data pairs and dynamic weight fusion and fine-tuning, the model has the dual capabilities of accurate retrieval of professional knowledge and deep alignment of expert decision preferences. This solves the industry pain points of low efficiency of traditional manual modeling, prominent illusion of general LLM, and poor adaptability of single technical solutions.
[0046] (2) Breaking away from the conventional approach of independently applying the direct preference optimization algorithm and the probability ratio preference optimization algorithm, this paper designs a dynamic weight factor adaptive adjustment strategy based on the dual requirements of professionalism and efficiency in power modeling. During the training process, the knowledge constraint advantage of the direct preference optimization algorithm and the efficient training characteristics of the probability ratio preference optimization algorithm can be flexibly balanced according to the number of steps and performance feedback, thus avoiding the adaptability limitations of fixed weight fusion. At the same time, by pre-calling the first major language model to generate power business decision-making questions and expert decision-making texts that conform to power business specifications, the professionalism and relevance of the preference dataset are ensured, providing high-quality supervision signals for fine-tuning training and realizing the positive transmission from data quality to model performance. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating a collaborative optimization method for RAG and preference alignment in the power sector according to an embodiment of the present invention.
[0048] Figure 2 This is a schematic diagram of the process of obtaining the target large language model through fine-tuning training in an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the structure of a RAG and preference alignment collaborative optimization system for the power industry according to an embodiment of the present invention;
[0050] Figure label:
[0051] 01. Data Acquisition and Processing Module; 02. Vectorization and RAG Module; 03. Question Text Generation Module; 04. Association and Combination Module; 05. Fine-tuning Training Module. Detailed Implementation
[0052] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0053] In the description of this invention, it should be understood that the terms "first" and "second," etc., are used to distinguish different objects, rather than to describe a specific order.
[0054] In the description of this invention, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by those skilled in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0055] To address the issues of lag in updates caused by manual coding in traditional power business decision-making models and the lack of professional knowledge and illusion phenomena in general-purpose large language models in the power field, Retrieval Augmentation Generation (RAG) and preference alignment have emerged as two key technical approaches. RAG, by introducing a real-time retrieval mechanism from an external professional knowledge base, provides factual basis for LLM generation, effectively improving the accuracy of the output. It has been applied in scenarios such as power equipment condition assessment and question-answering systems. However, existing RAG solutions suffer from drawbacks such as a lack of expert decision preferences in the generated results, uncontrollable security, difficulty in outputting optimal solutions, and insufficient real-time performance due to complex processes. Preference alignment technology, by optimizing the fit between model output and expert decision preferences and security criteria, can suppress illusions and enhance decision interpretability. It has been attempted for sensitive tasks such as dispatch instruction generation. However, the effectiveness of this technology is highly dependent on data quality and the capabilities of the base model, cannot update the knowledge base independently, and faces risks of insufficient generalization and lagging rule iteration. Therefore, this invention proposes a deep collaborative optimization scheme.
[0056] like Figure 1 The diagram shown is a flowchart illustrating a collaborative optimization method for RAG and preference alignment in the power sector according to an embodiment of the present invention. (Refer to...) Figure 1 This invention provides a collaborative optimization method for RAG and preference alignment in the power sector, comprising the following steps:
[0057] S1. Obtain several documents related to power business decisions to create a power business decision knowledge base;
[0058] Specifically, step S1 includes:
[0059] 1) Collect several PDF documents related to power business decisions from multiple publicly available data sources;
[0060] This embodiment focuses on various core power business decision-making scenarios, such as photovoltaic capacity calculation, power system state estimation, load forecasting, grid economic dispatch, and unit combination. It traversed authoritative academic resource databases such as CNKI, Wanfang, and VIP, professional journal databases in the power field such as IEEE Xplore and Elsevier, and university dissertation databases. Through subject search and in-depth abstract reading, it screened literature and finally collected 465 PDF documents that focus on power business scenarios and modeling solutions, taking into account both theory and practice. These documents cover core directions such as system analysis, renewable energy, and smart grids, laying a data foundation for building a high-quality power business decision-making knowledge base.
[0061] 2) Convert the format of each PDF document to obtain the corresponding Markdown document;
[0062] The conversion process employs a collaborative approach combining MinerU software and Doubao AI. Specifically, MinerU software first batch-processes PDF documents, quickly converting them to Markdown format while preserving the original overall structure and core content. For documents containing numerous complex formulas and special formats, issues such as compilation errors and formula corruption after MinerU conversion are addressed through manual refinement and alignment correction using Doubao AI. This ensures the high fidelity and professionalism of core knowledge, key technical parameters, and formulas in the power industry, achieving lightweight document processing that facilitates subsequent structured parsing and vector embedding.
[0063] 3) Perform data cleaning on each Markdown document separately, and combine the cleaned Markdown documents to form a power business decision knowledge base.
[0064] Through multiple rounds of text parsing and intelligent filtering, noisy information in Markdown documents that does not substantially help in generating decision-making models is selectively filtered out, including author information, acknowledgments, and reference citations. The core text for power business decision-making is retained, such as problem descriptions, goal planning, constraints, and solution methods. At the same time, manual proofreading is used to correct logical deviations and formatting flaws in the documents and optimize the text structure to ensure that each cleaned document is logically coherent and formatted rigorously. Finally, all cleaned Markdown documents are integrated to form a comprehensive power business decision-making knowledge base covering core power business and a complete knowledge system.
[0065] S2. The power business decision knowledge base is segmented into text blocks, and each text block is generated through vectorization and retrieval enhancement to obtain power business-adapted text.
[0066] Specifically, step S2 includes:
[0067] 1) The power business decision knowledge base is divided into several text blocks according to document length;
[0068] An adaptive text segmentation strategy is adopted to fully balance the integrity of the document structure with the coherence of the modeling problem. For longer Markdown documents, logical segmentation is prioritized based on their chapter heading levels (#, ##, etc.) to ensure that each text block fully contains the problem definition, objective function, constraints, and other related content. For shorter documents such as journal articles and conference reports, a fixed maximum word count threshold (e.g., 5000) is set. If the document length does not exceed the threshold, it is retained as a whole; if it exceeds the threshold, it is segmented according to paragraph boundaries, while avoiding the truncation of formulas or key reasoning steps.
[0069] This embodiment uses an adaptive text segmentation strategy to divide 465 original documents in the power business decision knowledge base into 1266 semantically coherent text blocks, providing a reliable basic unit for subsequent vectorization and retrieval enhancement.
[0070] 2) Each text block is vectorized using a pre-defined text embedding model to obtain a vector representation of the corresponding text block, and each vector representation is combined to form a power business vector knowledge base;
[0071] This embodiment selects the BGE-M3 text embedding model as the preset text embedding model and performs standardized vectorization processing on each text block. Specifically, the text block is first tokenized and then mapped into a 1024-dimensional dense vector by the BGE-M3 encoder. This process can fully preserve the professional semantic information and the relevance of power industry terminology in the text block, ensuring the stability and comparability of the vector representation. After organizing the vector representations corresponding to all text blocks in a unified format, they are combined to construct a structurally standardized and highly efficient power business vector knowledge base.
[0072] 3) Based on the power business vector knowledge base, power business-adapted text is generated through retrieval enhancement.
[0073] The Retrieval Augmentation (RAG) technique is employed. First, the input power business-related topics or tasks are vectorized. Then, a similarity comparison is performed in a power business vector knowledge base to retrieve several text blocks highly relevant to the input as supplementary contextual information. Next, leveraging the text generation capabilities of the pre-invoked primary language model, the retrieved professional context is deeply integrated with the input requirements. The system systematically organizes and integrates core content such as power business problem definitions, goal planning, constraints, and solution methods to generate power business-adaptive text with professional accuracy, logical completeness, and instruction adaptability.
[0074] S3. Based on the power business adapted text, combined with the pre-called first major language model and retrieval enhancement, several power business decision questions are generated, as well as the expert decision text and ordinary decision text corresponding to each power business decision question.
[0075] Specifically, step S3 includes:
[0076] 1) Based on the power business-adapted text, several power business decision questions are generated by combining the preset question generation template with the pre-called first language model;
[0077] In this embodiment, the first major language model pre-called is DeepSeek V3. The preset question generation template must meet the requirements of the power business specifications. The core rules of the template include: the question must be highly consistent with the core content of the power business adaptation text and focus on multiple core power business scenarios; the description must be detailed and accurate and can stand alone as a question without relying on the original text context; avoid large blocks of formula symbols and diversify the expression methods. The specific template is shown in Table 1 below.
[0078] Table 1 Preset Question Generation Template
[0079]
[0080] The power business adaptation text is split into semantic units and used as the input context. It is then fed into the DeepSeek V3 model, which automatically generates questions based on preset question generation templates. This results in 25,230 power business decision-making questions covering scenarios such as photovoltaic capacity calculation, grid economic dispatch, and unit combination, ensuring that the question types are comprehensive and the professional attributes are prominent.
[0081] 2) For each power business decision problem, the expert decision text corresponding to the power business decision problem is generated by combining the first language model and retrieval enhancement;
[0082] Utilizing DeepSeek V3 as the primary language model, a collaborative generation mechanism of "retrieval enhancement + structured prompts" was employed. Specifically, each power business decision-making problem was first vectorized, and highly relevant text blocks were retrieved from the power business vector knowledge base as professional context support. Then, customized prompts (including format specifications and professional requirements) were loaded. These prompts explicitly required the output to include a problem description, a modeling objective function (LaTeX format + physical meaning annotations), constraints (including power-specific rules such as N-1 safety constraints), a variable list, and executable pseudocode. Based on the retrieved professional context and prompts, the primary language model generated text conforming to the decision-making logic of power industry experts, ultimately yielding 25,230 expert decision texts (Good Response). All texts were formatted uniformly, rigorously presented, and free from factual errors and misleading information, accurately matching power business modeling specifications.
[0083] 3) For each power business decision problem, generate the corresponding ordinary decision text by combining the second major language model in the initial state.
[0084] In the initial state of this embodiment, the second major language model is the Qwen2.5-1.5B-Instruct model without any fine-tuning. In this state, the model does not incorporate professional knowledge in the power field or expert decision-making preferences. Each power business decision question is directly used as the input prompt, without introducing professional contextual support for retrieval enhancement. The model autonomously generates responses based on native pre-trained knowledge, ultimately resulting in 25,230 ordinary decision texts (Bad Responses).
[0085] It should be noted that this type of text differs significantly from the corresponding expert decision-making text, mainly in that it lacks electricity-specific constraint logic, has insufficient formula standardization, and does not align with industry decision-making preferences (such as not reflecting the principle of "prioritizing power supply"). Its main purpose is to provide clear negative sample references for subsequent preference alignment training.
[0086] S4. Associate and combine each power business decision problem with the corresponding expert decision text and ordinary decision text according to the triplet structure to obtain the preference dataset;
[0087] Using a fixed triple structure of "power business decision-making problem - expert decision text - ordinary decision text", the triples are associated and combined according to a one-to-one correspondence principle. Specifically, each power business decision-making problem (Prompt) is taken as the core of the triple, and its exclusively generated expert decision text (Good Response) and ordinary decision text (Bad Response) are matched to ensure that the three types of texts within the same triple accurately correspond to the decision-making needs and response results of different quality levels under the same power business scenario.
[0088] Through this association method, a dataset containing 25,230 high-quality power sector preference-aligned triples was ultimately constructed, namely the preference dataset. This dataset not only ensures the consistency of scenarios and logical correlation of each training sample, but also provides clear supervision signals for subsequent model preference alignment fine-tuning, ensuring that the model can learn the decision-making preferences and professional norms of power sector experts.
[0089] To facilitate understanding, the following example illustrates the preference alignment triplet in the power sector, as shown in Table 2.
[0090] Table 2. Preference Alignment Triads in the Power Sector
[0091]
[0092] S5. Based on the preference dataset, the pre-trained second language model is fine-tuned by using a weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the target large language model.
[0093] While the pre-trained second language model possesses basic language generation capabilities, it has not deeply integrated the decision-making preferences and professional modeling standards of experts in the power sector. As a result, it suffers from issues such as a lack of industry adaptability in its output content, imprecise formula logic, and illusion phenomena, making it difficult to directly meet the high reliability requirements of power business decision-making.
[0094] Preference alignment refers to training a pre-trained large language model using manually designed preference-aligned data pairs. This allows the model to perform human-like alignment for specific tasks according to specific instructions, ensuring the professional accuracy and engineering practicality of the generated text. Through preference alignment training, the large model can better distinguish between high-quality and low-quality answers, thus avoiding poor output text during generation. The large model learns relative preferences rather than absolute answers; even if the reference answer has minor flaws, as long as it is better than the negative sample, the large model can still gradually optimize the output through preference relationships.
[0095] The preference dataset constructed in step S4, in the form of triplets of "power business decision-making problem - expert decision-making text - ordinary decision-making text" (preference-aligned data pairs), provides clear supervision signals for the model. The expert decision-making text carries the normative modeling logic, specific constraint rules and industry decision-making preferences in the power field, while the ordinary decision-making text clarifies the non-professional output forms that need to be avoided. The comparison between the two can guide the model to learn relative preferences rather than absolute answers.
[0096] To ensure the stability and evaluability of the fine-tuning effect, this step divides the preference dataset into a training set, a validation set, and a test set, which are used for model parameter training, hyperparameter dynamic adjustment, and final performance verification, respectively. Through dynamic weighted fusion of direct preference optimization and probability ratio preference optimization, the model achieves a deep alignment with power industry expertise and expert decision-making logic.
[0097] like Figure 2 As shown, this is a flowchart illustrating step S5. (Refer to...) Figure 2 Step S5 includes:
[0098] S501. Divide the preference dataset into training set, test set and validation set according to a preset ratio;
[0099] This embodiment adopts an "8-1-1" partitioning ratio to randomly hierarchically partition the preference dataset containing 25,230 power sector preference alignment triples, ensuring that triples of multiple core power business scenarios are evenly distributed across the three datasets and avoiding excessive concentration of data in a single scenario.
[0100] The training set contains 20,184 triples (80%), used for gradient updates and preference learning of the model's core parameters; the validation set contains 2,523 triples (10%), used for hyperparameter tuning and performance monitoring during training; and the test set contains 2,523 triples (10%), used as independent data for the final evaluation of the model's generalization ability.
[0101] S502. Based on the training set, the pre-trained second large language model is fine-tuned by using a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model.
[0102] Based on the training set, the pre-trained second large language model is fine-tuned using a fusion loss function consisting of a dynamic weighted sum of direct preference optimization loss and probability ratio preference optimization loss, to obtain the initial large language model.
[0103] To address the limitations of single reinforcement learning algorithms in power decision modeling scenarios, this step proposes a weighted fusion optimization strategy of Direct Preference Optimization (DPO) and Probability Ratio Preference Optimization (ORPO), forming an optimization framework that combines stability and efficiency.
[0104] Specifically, the fusion loss function is defined using the following formula:
[0105]
[0106] in, Represents the fusion loss function. This represents the dynamic weighting factor of the direct preference optimization loss. This represents the direct preference optimization loss. The dynamic weighting factor represents the probability-to-preference optimization loss. This represents the probability-to-preference optimization loss. It should be noted that... and The sum is 1, and it increases with the number of training steps. Adaptive adjustment allows the fusion loss function to meet the differentiated needs of the entire modeling process.
[0107] Furthermore, The loss is defined using the standard direct preference optimization formula:
[0108]
[0109] in, Represents the mathematical expectation. This refers to power business decision-making issues. This represents the expert decision-making text. This represents a typical decision-making text. This represents the Sigmoid function. Indicates hyperparameters, This represents the decision model that needs to be trained. A reference model representing the freeze parameters.
[0110] Furthermore, The loss is defined using the standard odds ratio preference optimization formula:
[0111]
[0112] in, Represents the standard language modeling loss, Indicates hyperparameters, It represents the probability ratio loss and is used to widen the probability gap between expert decision-making texts and ordinary decision-making texts.
[0113] The standard language modeling loss is calculated using the following formula:
[0114]
[0115] in, The effective token length of the expert decision text. This represents a conditional probability distribution.
[0116] The probability ratio loss is calculated using the following formula:
[0117]
[0118] in, Indicates the probability ratio. The probability is expressed, and the calculation formula is shown below:
[0119]
[0120] By using the aforementioned fusion loss function, the knowledge constraint advantage of the direct preference optimization algorithm and the efficient training characteristics of the probability ratio preference optimization algorithm are balanced, enabling the model to maintain the accuracy of professional knowledge while learning the preferences of power experts.
[0121] The dynamic weighting factor of the direct preference optimization loss is determined by the base weighting factor adjusted piecewise based on the number of training steps, the model performance bias during training, and the range constraints of the value.
[0122] Specifically, the dynamic weighting factor for the direct preference optimization loss is defined using the following formula:
[0123]
[0124] in, This represents a function that constrains the range of values. Indicates the basic weighting factor. This represents the feedback strength coefficient (reference value is 0.8), used to control the magnitude of weight adjustment based on validation set performance. This represents the target performance score of the validation set. This represents the overall performance score of the validation set. This represents the lower boundary of the value constraint (reference value is 0.2). This indicates the upper boundary of the value constraint (reference value is 0.7). It should be noted that... This indicates the model performance deviation during training (positive values indicate failure to meet the target, negative values indicate success). When performance fails to meet the target, Increase and strengthen the knowledge constraints of direct preference optimization; when performance meets the target, Reduce the size to avoid model overfitting.
[0125] Furthermore, the adjustment strategy for the basic weighting factor includes:
[0126] 1) Divide the number of training steps into a warm-up period, a growth period, and a stable period according to a preset stage threshold;
[0127] Preset stage thresholds include the threshold for the number of steps at the end of the warm-up period. (The first 20%-30% of the total steps) and the threshold for the end of the rising phase. (The last 60%-70% of the total steps). In other words, the warm-up period is... The rising period is The stable period is .in, This represents the total number of training steps.
[0128] 2) When the training process is in the warm-up period, the base weight factor gradually increases from the weight at the beginning of the warm-up period to the weight at the end of the warm-up period;
[0129] Specifically, when the training process is in the warm-up period, i.e. The basic weighting factor is calculated using the following formula:
[0130]
[0131] in, This indicates the initial weight value for the preheating period (a reference value of 0.1). This represents the base weight value at the end of the warm-up period (reference value is 0.3), used to bridge the weight transition between the warm-up and rising periods. The steepness coefficient of the Sigmoid function (reference value is 8) is used to control the rate of change of weight during the preheating process.
[0132] During the warm-up period, the high efficiency of single-stage training and no-reference model by leveraging the probability ratio preference optimization allows the model to quickly learn the preference ranking relationship of power business decisions. This avoids training oscillations caused by excessively high weights in the initial preference optimization and lays the foundation for subsequent professional knowledge constraints.
[0133] 3) When the training process is in the ascending phase, the base weight factor gradually increases from the end of the warm-up period to the peak weight.
[0134] Specifically, when the training process is in its ascending phase... The basic weighting factor is calculated using the following formula:
[0135]
[0136] in, This represents the base weight value at the end of the rising phase (reference value is 0.6), and also represents the peak value of the direct preference optimization weight (maximizing knowledge constraints). The steepness coefficient of the Sigmoid function (reference value is 6) is used to control the rate of change of weight on an upward trend.
[0137] As the training steps progress, the weight of the direct preference optimization loss is gradually increased, strengthening the KL divergence constraint advantage of direct preference optimization and ensuring that the model does not deviate from the professional knowledge and modeling standards of the power field while learning preferences.
[0138] 4) When the training process is in a stable period, the basic weight factor gradually decreases from the weight peak to the balanced weight at the end of the training.
[0139] Specifically, when the training process is in a stable period, i.e. The basic weighting factor is calculated using the following formula:
[0140]
[0141] in, This represents the base weight value at the end of training (reference value is 0.4), and also represents the final balanced weight between direct preference optimization and probability ratio preference optimization. The steepness coefficient of the Sigmoid function (reference value is 6) is used to control the rate of change of weight in stable options.
[0142] At this point, direct preference optimization loss dominates. Through its strong knowledge constraint characteristics, it deeply solidifies the model's memory of professional content such as objective function design and constraint definition in the power field, while retaining the efficient training advantage of probability ratio preference optimization, ensuring that the model takes into account both professional accuracy and preference alignment during the convergence process.
[0143] S503. Combine the validation set to adjust the hyperparameters of the initial large language model in real time to obtain the intermediate large language model;
[0144] The hyperparameters to be adjusted include the learning rate, batch size, weight decay coefficient, and weight boundaries of the fusion loss function. Specifically, this embodiment uses a Bayesian optimization algorithm. After every 100 training steps, the model performance is evaluated using a validation set. If the semantic consistency score does not improve or decreases for three consecutive rounds, hyperparameter adjustment is triggered: the learning rate is initially set to 5e-6, decaying by a factor of 0.9; the batch size is dynamically adjusted to 8 or 16 based on memory usage; the weight decay coefficient ranges from [0.005, 0.01] and only applies to the LoRA trainable matrix; simultaneously, the weight boundaries of the fusion loss function are adjusted to optimize weight fit. An early stopping strategy is adopted, stopping adjustment when the overall performance score on the validation set does not improve for five consecutive rounds, ultimately obtaining an intermediate large language model that balances training stability and professional performance.
[0145] S504. Use the test set to evaluate the performance of the intermediate large language model and obtain the target large language model.
[0146] The intermediate large language model is evaluated using a test set from multiple dimensions, including semantic consistency, professional reliability, and preference alignment, to obtain the target large language model.
[0147] Each dimension will be explained in detail below:
[0148] 1) Semantic consistency:
[0149] Semantic consistency refers to the degree of semantic matching between the model-generated text and the core requirements of power business decision-making issues and professional reference texts. It is quantified using BERT Score, and the formula is as follows:
[0150]
[0151] in, This indicates the accuracy of the token matching between the model-generated text and the professional reference text. Indicates the match recall rate. express and The harmonized average.
[0152] 2) Professional reliability:
[0153] Professional reliability refers to the technical accuracy and regulatory compliance of model-generated text in the power field, which includes the formula standardization compliance rate (number of standardized formulas / total number of formulas) and the illusion rate (number of factually incorrect statements / total number of statements).
[0154] 3) Preference Alignment:
[0155] Preference alignment refers to the degree to which the model-generated text closely resembles expert decision text and deviates significantly from ordinary decision text. It combines expert scoring with comparative similarity, where comparative similarity is defined as:
[0156]
[0157] in, Indicates the degree of similarity. Represents cosine similarity. This indicates that the model generates text.
[0158] If all dimensions of the intermediate large language model meet the corresponding threshold, and the overall score (weighted by factors such as semantic consistency 0.3, professional reliability 0.4, and preference alignment 0.3) ranks high (e.g., in the top 10%), then it is determined to be the target large language model; if it does not meet the threshold, return to step S503 to readjust the hyperparameters and evaluate again.
[0159] This invention presents a collaborative optimization method for RAG and preference alignment in the power sector. It is the first to construct a closed-loop link of "structured knowledge base accumulation - adaptive generation of professional data - weighted fusion and fine-tuning of dual algorithms", which deeply couples the factual support capability of enhanced retrieval with the value orientation capability of preference alignment. Unlike single RAG technologies that lack expert preference adaptation and single preference alignment technologies that suffer from insufficient generalization, this invention provides a high-quality professional data foundation covering the power business through the precise construction of a power business decision knowledge base and adaptive text segmentation. Further, through triplet data generation and dynamic weight fusion fine-tuning, the model simultaneously possesses the dual capabilities of accurate professional knowledge retrieval and deep alignment with expert decision preferences. This addresses industry pain points such as low efficiency of traditional manual modeling, prominent illusions of general LLM, and poor adaptability of single technical solutions. It breaks through the conventional mode of independently applying direct preference optimization algorithms and probability ratio preference optimization algorithms. Based on the dual requirements of professionalism and efficiency in power modeling, a dynamic weight factor adaptive adjustment strategy is designed. During training, the knowledge constraint advantages of direct preference optimization algorithms and the efficient training characteristics of probability ratio preference optimization algorithms can be flexibly balanced according to the number of steps and performance feedback, avoiding the adaptability limitations of fixed weight fusion. Meanwhile, by pre-calling the first major language model to generate power business decision-making questions and expert decision-making texts that conform to power business specifications, the professionalism and relevance of the preference dataset are ensured, providing high-quality supervision signals for fine-tuning training and realizing a positive transmission from data quality to model performance.
[0160] like Figure 3 The diagram shown is a structural schematic of a RAG and preference alignment collaborative optimization system for the power sector, according to an embodiment of the present invention. (Refer to...) Figure 3This invention provides a collaborative optimization system for RAG and preference alignment in the power sector, comprising:
[0161] Data acquisition and processing module 01 is used to acquire several documents related to power business decisions and obtain a power business decision knowledge base.
[0162] The Vectorization and RAG module 02 is used to segment the text of the power business decision knowledge base, and generate power business-adapted text by vectorizing each text block and enhancing retrieval through vectorization representation and retrieval enhancement.
[0163] The problem text generation module 03 is used to generate several power business decision questions and corresponding expert decision texts and ordinary decision texts based on power business adapted texts, combined with the pre-called first major language model and retrieval enhancement.
[0164] The association and combination module 04 is used to associate and combine each power business decision problem with the corresponding expert decision text and ordinary decision text according to the triplet structure to obtain the preference dataset;
[0165] The fine-tuning training module 05 is used to fine-tune the pre-trained second-largest language model based on the preference dataset, using a weighted fusion of direct preference optimization and probability ratio preference optimization, to obtain the target large language model.
[0166] It should be noted that the modules in the aforementioned RAG and preference alignment collaborative optimization system for the power sector can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module. For specific limitations regarding the RAG and preference alignment collaborative optimization system for the power sector, please refer to the limitations of the RAG and preference alignment collaborative optimization method for the power sector described above; both have the same function and role, and will not be repeated here.
[0167] In summary, the present invention provides a collaborative optimization method and system for RAG and preference alignment in the power sector. For the first time, it constructs a closed-loop link of "structured knowledge base accumulation - adaptive generation of professional data - weighted fusion and fine-tuning of dual algorithms", which deeply couples the factual support capability of enhanced retrieval with the value orientation capability of preference alignment. Unlike single RAG technologies that lack expert preference adaptation and single preference alignment technologies that suffer from insufficient generalization, this invention provides a high-quality professional data foundation covering the power business through the precise construction of a power business decision knowledge base and adaptive text segmentation. Further, through triplet data generation and dynamic weight fusion fine-tuning, the model simultaneously possesses the dual capabilities of accurate professional knowledge retrieval and deep alignment with expert decision preferences. This addresses industry pain points such as low efficiency of traditional manual modeling, prominent illusions of general LLM, and poor adaptability of single technical solutions. It breaks through the conventional mode of independently applying direct preference optimization algorithms and probability ratio preference optimization algorithms. Based on the dual requirements of professionalism and efficiency in power modeling, a dynamic weight factor adaptive adjustment strategy is designed. During training, the knowledge constraint advantages of direct preference optimization algorithms and the efficient training characteristics of probability ratio preference optimization algorithms can be flexibly balanced according to the number of steps and performance feedback, avoiding the adaptability limitations of fixed weight fusion. Meanwhile, by pre-calling the first major language model to generate power business decision-making questions and expert decision-making texts that conform to power business specifications, the professionalism and relevance of the preference dataset are ensured, providing high-quality supervision signals for fine-tuning training and realizing a positive transmission from data quality to model performance.
[0168] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0169] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention.
Claims
1. A collaborative optimization method for RAG and preference alignment in the power sector, characterized in that, include: Acquire several documents related to power business decisions to obtain a power business decision knowledge base; The power business decision knowledge base is segmented into text blocks, and each text block is generated through vectorization representation and retrieval enhancement to obtain power business-adapted text. Based on the power business adaptation text, combined with the pre-called first major language model and retrieval enhancement, several power business decision questions are generated, as well as the expert decision text and ordinary decision text corresponding to each power business decision question. Each power business decision problem is associated and combined with the corresponding expert decision text and ordinary decision text in a triplet structure to obtain a preference dataset; Based on the aforementioned preference dataset, a weighted fusion of direct preference optimization and probability ratio preference optimization is used to fine-tune the pre-trained second-largest language model, resulting in the target large language model, including: The pre-trained second large language model was fine-tuned by a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model. The method of fine-tuning the pre-trained second large language model by using a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model includes: The pre-trained second large language model is fine-tuned using a fusion loss function consisting of a dynamic weighted sum of direct preference optimization loss and probability ratio preference optimization loss to obtain the initial large language model; The dynamic weighting factor of the direct preference optimization loss is determined by the basic weighting factor adjusted in segments based on the number of training steps, the model performance deviation during the training process, and the value range constraints. The adjustment strategy for the basic weighting factor includes: The number of training steps is divided into a warm-up period, an ascent period, and a stabilization period according to a preset stage threshold. When the training process is in the warm-up period, the basic weight factor gradually increases from the weight at the beginning of the warm-up period to the weight at the end of the warm-up period. When the training process is in the rising phase, the basic weight factor gradually increases from the end of the warm-up period to the peak weight. When the training process is in a stable period, the basic weight factor gradually decreases from the weight peak to the balanced weight at the end of the training.
2. The RAG and preference alignment collaborative optimization method for the power sector according to claim 1, characterized in that, The process involves acquiring several documents related to power business decisions to obtain a power business decision knowledge base, including: Several PDF documents related to power business decisions were collected from multiple publicly available data sources; Each PDF document is converted to its corresponding Markdown document. Each Markdown document is cleaned separately, and the cleaned Markdown documents are combined to form a power business decision knowledge base.
3. The RAG and preference alignment collaborative optimization method for the power sector according to claim 1, characterized in that, The process of segmenting the power business decision knowledge base into text blocks and generating power business-adapted text by vectorizing and enhancing the retrieval of each text block includes: The power business decision knowledge base is divided into text blocks according to document length to obtain several text blocks; Each text block is vectorized using a preset text embedding model to obtain a vector representation of the corresponding text block, and the vector representations are combined to form a power business vector knowledge base; Based on the aforementioned power business vector knowledge base, power business-adapted text is generated through enhanced retrieval.
4. The RAG and preference alignment collaborative optimization method for the power sector according to claim 1, characterized in that, The process involves generating several power business decision-making questions based on the power business-adapted text, combined with a pre-invoked first major language model and retrieval enhancement, along with expert decision text and ordinary decision text corresponding to each power business decision-making question. These include: Based on the power business adaptation text, several power business decision questions are generated by combining the preset question generation template and the pre-called first major language model. For each of the aforementioned power business decision problems, the expert decision text corresponding to the power business decision problem is generated by combining the first large language model with retrieval enhancement. For each of the aforementioned power business decision problems, a general decision text corresponding to the power business decision problem is generated by combining the second major language model in the initial state.
5. The RAG and preference alignment collaborative optimization method for the power sector according to claim 1, characterized in that, The step of fine-tuning the pre-trained second large language model based on the preference dataset, using a weighted fusion of direct preference optimization and probability ratio preference optimization, to obtain the target large language model, also includes: The preference dataset is divided into a training set, a test set, and a validation set according to a preset ratio; Based on the training set, the pre-trained second large language model is fine-tuned by a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model. The hyperparameters of the initial large language model are adjusted in real time based on the validation set to obtain an intermediate large language model; The intermediate large language model is evaluated using the test set to obtain the target large language model.
6. The RAG and preference alignment collaborative optimization method for the power sector according to claim 5, characterized in that, The step of evaluating the performance of the intermediate large language model using the test set to obtain the target large language model includes: The intermediate large language model is evaluated using the test set from multiple dimensions, including semantic consistency, professional reliability, and preference alignment, to obtain the target large language model.
7. A RAG and preference alignment collaborative optimization system for the power sector, characterized in that, include: The data acquisition and processing module is used to acquire several documents related to power business decisions and obtain a power business decision knowledge base. The vectorization and RAG module is used to segment the power business decision knowledge base into text blocks, and generate power business-adapted text by vectorizing each text block and enhancing retrieval through vectorization representation and retrieval. The problem text generation module is used to generate several power business decision questions and corresponding expert decision text and ordinary decision text based on the power business adaptation text, combined with the pre-called first major language model and retrieval enhancement; The association and combination module is used to associate and combine each power business decision problem with the corresponding expert decision text and ordinary decision text according to the triplet structure to obtain a preference dataset; The fine-tuning training module is used to fine-tune the pre-trained second-largest language model based on the aforementioned preference dataset, employing a weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the target large-scale language model, including: The pre-trained second large language model was fine-tuned by a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model. The method of fine-tuning the pre-trained second large language model by using a dynamic weighted fusion of direct preference optimization and probability ratio preference optimization to obtain the initial large language model includes: The pre-trained second large language model is fine-tuned using a fusion loss function consisting of a dynamic weighted sum of direct preference optimization loss and probability ratio preference optimization loss to obtain the initial large language model; The dynamic weighting factor of the direct preference optimization loss is determined by the basic weighting factor adjusted in segments based on the number of training steps, the model performance deviation during the training process, and the value range constraints. The adjustment strategy for the basic weighting factor includes: The number of training steps is divided into a warm-up period, an ascent period, and a stabilization period according to a preset stage threshold. When the training process is in the warm-up period, the basic weight factor gradually increases from the weight at the beginning of the warm-up period to the weight at the end of the warm-up period. When the training process is in the rising phase, the basic weight factor gradually increases from the end of the warm-up period to the peak weight. When the training process is in a stable period, the basic weight factor gradually decreases from the weight peak to the balanced weight at the end of the training.
Citation Information
Patent Citations
Alignment model training method, information processing method and device
CN119106739A
Data insight method and system based on AI large model
CN120087372A
Modeling method and device of power business decision model, equipment and medium
CN120821832A
Knowledge fusion-based retrieval enhanced large language model system and generation method
CN120950676A