A method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning
By introducing knowledge enhancement and adaptive fine-tuning technologies into the power large language model, combining knowledge graphs and SeqGAN generators, a hybrid expert model is built, which solves the problems of data scarcity and insufficient generalization capabilities of traditional models when dealing with complex power tasks, and significantly improves the model's inference ability and adaptability.
Patent Information
- Application Number
- CN202510014091.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-06
AI Technical Summary
Traditional large language models face problems such as scarcity of data, insufficient domain knowledge, and insufficient model generalization capabilities when processing proprietary data in specific fields, resulting in poor performance in complex power tasks.
Using a power large language model construction method based on knowledge enhancement and adaptive fine-tuning, we use SeqGAN to generate diversified data samples, design knowledge graphs and enhance model reasoning capabilities in combination with their vector embedding, and build a hybrid expert model architecture to process diversified power data.
Effectively combining the structured information of the knowledge graph, strengthening the model's reasoning ability and the application ability of complex power tasks, and improving the model's performance and adaptability in diversified power tasks.
Smart Images

Figure CN119416874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electric power large language model construction, and in particular to an electric power large language model construction method based on knowledge enhancement and adaptive fine-tuning. Background Art
[0002] With the rapid increase in complexity and data volume in the power industry, the management and optimization of power systems face unprecedented challenges. The power industry covers a wide range of fields, including power equipment management, system status monitoring, process optimization, and compliance with standards and regulations. In order to improve efficiency and decision-making accuracy, more and more companies are beginning to seek artificial intelligence and natural language processing technologies to process and analyze large amounts of power data.
[0003] Large language models have shown great potential in processing natural language texts, and can perform prediction and generation tasks by learning patterns in data. However, traditional methods face problems such as data scarcity, insufficient domain knowledge, and insufficient model generalization when processing proprietary data in a specific field. Therefore, the power industry urgently needs a new method that can combine domain knowledge and adaptive capabilities to improve the performance of models in complex tasks. Summary of the invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide a method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning, so that the constructed large language model can effectively combine the structured information of the knowledge graph, enhance the model's reasoning ability and its application ability in complex power tasks.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning includes the following steps:
[0007] S1: Collect text data related to the power field, and pre-process to remove redundant, irrelevant or repeated information to obtain pre-processed text data;
[0008] S2: Based on the preprocessed text data, SeqGAN is used to generate diverse data samples, and an adversarial mechanism is added to the SeqGAN generation step;
[0009] S3: Use SpaCy for text annotation and basic NLP processing to identify and annotate power-related entities;
[0010] S4: Use Neo4j to design and build a knowledge graph in the power field, covering equipment, system, process concepts and their relationships;
[0011] S5: Based on the pre-trained large language basic model, a hybrid expert model architecture is designed to process different types of power data through different expert modules, improve the performance of the model on diversified tasks, and build a large power language model;
[0012] S6: Combine the knowledge graph with the power language model and enhance the reasoning ability of the model through prompts and attention mechanisms, as follows:
[0013] S61: Use the knowledge graph embedding method to convert entities and relations in the knowledge graph into vector representations;
[0014] S62: embed the knowledge graph vector into the input of the power large language model as supplementary information to enhance the model's context perception and reasoning capabilities, and inject the extracted graph features into the specific layer of the corresponding expert model to enhance domain knowledge;
[0015] S63: Construct specific natural language prompts, combined with graph embedding, as contextual activation parameters for the language model;
[0016] S64: Generate dynamic prompts based on current input content and related knowledge graph node information;
[0017] S65: Add a cross-attention layer inside the power large language model to focus on important information obtained through graph embedding and prompts.
[0018] Furthermore, S1 is specifically:
[0019] S11: Collect text data related to the power field, convert data in different formats into a unified text format, and perform text extraction;
[0020] S12: Remove redundant information from the text, including HTML tags, footnotes, headers and footers; use regular expressions to clean up special characters, blank lines and irrelevant tags;
[0021] S13: Use the text similarity algorithm Jaccard similarity to detect and remove duplicate documents and paragraphs, and filter out content related to the power field through keyword filtering and topic modeling LDA.
[0022] Furthermore, S13 is specifically:
[0023] Split each document or paragraph into a set of words and calculate the Jaccard similarity between each pair of documents or paragraphs;
[0024] Set a similarity threshold. Text pairs above this threshold are considered duplicates and one of them is removed.
[0025] Use the LDA model to train the text and set the number of topics;
[0026] Analyze the topic distribution of each document, filter out topics related to the power field, retain documents related to the target topic, and filter out other content;
[0027] Create a list of keywords related to the power sector, traverse each document or paragraph, check whether it contains words in the keyword list, retain the text containing the keywords, and filter out irrelevant content.
[0028] Furthermore, the LDA model is used to train the text and set the number of topics. The topic distribution of each document is analyzed, and topics related to the power field are screened out. Documents related to the target topic are retained, and other content is filtered out, as follows:
[0029] Split text data into words or phrases, form bags of words, and remove stop words;
[0030] Use stem extraction to normalize different forms of the same word, count the frequency of each word in the document, and form a word frequency matrix;
[0031] According to the prior knowledge of the domain, the number of topics is set to K, and the LDA model is used to train documents to extract topics.
[0032] The goal of the LDA model is to find the topic distribution by maximizing the likelihood:
[0033] ;
[0034] ;
[0035] in, is the topic distribution of document d; is the topic word distribution of topic k; Dirichlet prior for document topic distribution; is the Dirichlet prior of the keyword distribution; represents Dirichlet distribution;
[0036] Each document generation process:
[0037] ;
[0038] in, is the topic of the nth word in document d; is the nth word in document d; Multinomial represents multinomial distribution;
[0039] Analyze the LDA model training results and mark the specific content reflected by each topic. Identify topics directly related to the power sector;
[0040] For each document, calculate its topic distribution and identify the weight of electricity-related topics;
[0041] Documents with power-related topic weights greater than the threshold are saved, and other content that does not match the target topic is filtered out.
[0042] Furthermore, SeqGAN is used to generate diverse data samples based on the preprocessed text data, specifically:
[0043] Convert the preprocessed text data into a format suitable for SeqGAN input and train SeqGAN to generate diverse synthetic data;
[0044] The generator G is an RNN that generates text sequences. The goal of the generator is to generate realistic text sequences to deceive the discriminator. The generator loss is:
[0045] ;
[0046] in, is the loss function of the generator; represents the text sequence Y generated from the generator G; is the output of the discriminator for generating the text sequence Y; E represents the expected function;
[0047] The discriminator D is a classifier used to distinguish between real text sequences and generated text sequences. The goal of the discriminator is to accurately identify the generated text sequences;
[0048] ;
[0049] in, is the loss function of the discriminator; Represents the distribution of real data p data The real text sequence X sampled in is the logarithm of the discriminator's output for the true text sequence X.
[0050] Furthermore, an adversarial mechanism is added to the SeqGAN generation step, as follows:
[0051] Build a knowledge base in the power field, including key terms, standards, regulations and typical data patterns, and extract key features and patterns from the knowledge base as constraints in the generation process;
[0052] A conditional generation mechanism is introduced into the generator to ensure that the generated data meets the domain requirements. The generator formula is:
[0053] ;
[0054] in, is the previous word of the generated sequence, c is the conditional vector containing specific constraints in the power field; is a generated word;
[0055] The specific characteristics of the power field are input into the generator as conditions to guide the generation process. It is necessary to add domain constraints to the loss function to penalize the generated data that does not meet the domain standards:
[0056] ;
[0057] in, is the domain constraint loss, λ is the weight coefficient;
[0058] A domain feature detection module is added to the discriminator to identify features that do not meet the standards of the power field, and the pre-trained domain feature detection module is used to fine-tune the discriminator.
[0059] Furthermore, the power large language model includes a basic model, a routing layer and an expert module, wherein the basic model uses a pre-trained large language model for processing common language tasks, including semantic understanding, sentence generation, and general feature representation for further processing by the expert module;
[0060] The routing layer determines the best expert module based on the input data characteristics:
[0061] ;
[0062] in, is the probability of the input being assigned to the i-th expert; is the routing weight matrix; is the feature representation of the input;
[0063] The expert module includes several expert submodules, each of which focuses on a specific power data type or task, including an equipment expert submodule for processing equipment-related data; a system expert submodule for processing system-level data; and a process expert submodule for processing process-related data. Each expert submodule is a neural network trained with specific task data.
[0064] Overall loss function:
[0065] ;
[0066] Where N is the number of expert submodules; α i is the weight of the loss of the i-th expert module; D i is the dataset prepared for the i-th expert; is the loss function of the i-th expert on his specific task; is the global regularization term; is the regularization parameter, is the expected function.
[0067] Furthermore, the training process of the power language model is as follows:
[0068] Select the pre-trained large language model BERT as the base model;
[0069] Initialize the basic model and expert modules, and initialize the parameters through the pre-trained large language model; add specialized layers to each expert module to handle its task, and the initial parameters of these layers are selected from the modules in similar fields of the basic model;
[0070] Train the router to learn the mapping relationship between task characteristics and expert modules. Use a full dataset containing various tasks and data types to enable the router to learn to distinguish the characteristics of different task characteristics. Use the cross entropy loss function to optimize routing decisions so that it can correctly map task characteristics to appropriate expert modules.
[0071] Joint training is performed on the entire architecture so that routing decisions and expert modules are optimized together. In each training step, the parameters of the basic model, expert module and router are updated synchronously.
[0072] In another embodiment, a system for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning is provided, comprising a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning as described above.
[0073] In another embodiment, a computer storage medium is provided, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above method steps.
[0074] The present invention has the following beneficial effects:
[0075] 1. The large language model constructed by the present invention can effectively combine the structured information of the knowledge graph, enhance the reasoning ability of the model and its application ability in complex power tasks;
[0076] 2. The present invention combines Jaccard similarity and LDA topic modeling to effectively perform topic analysis and screening on large-scale texts, accurately locate documents closely related to the power field, and enable the model to focus more on key content in the field;
[0077] 3. The present invention uses SeqGAN to generate high-quality synthetic text data, enhances the richness and coverage of the power field data set, and introduces specific constraints in the power field into the generation process of SeqGAN to ensure the authenticity and compliance of the generated data;
[0078] 4. The present invention constructs a hybrid expert large language model for the power field, utilizes the general features of the pre-trained large language model, and processes tasks in the diversified power field through the specialized capabilities of the expert module, thereby improving the performance and adaptability of the overall model. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a flow chart of the method of the present invention;
[0080] Figure 2 A schematic diagram of a text processing method in one embodiment of the present invention;
[0081] Figure 3 This is a schematic diagram of the process of combining the knowledge graph with the power big language model in one embodiment of the present invention. DETAILED DESCRIPTION
[0082] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0083] refer to Figure 1 In this embodiment, a method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning is provided, comprising the following steps:
[0084] S1: Collect text data related to the power sector, including technical documents, standards, operating manuals, and academic papers, and preprocess to remove redundant, irrelevant, or repeated information to obtain preprocessed text data;
[0085] S2: Based on the preprocessed text data, SeqGAN is used to generate diverse data samples to enhance the richness of the data set, ensure that the data covers key changes in the power sector, and add an adversarial mechanism to the SeqGAN generation step to ensure the authenticity and compliance of the generated data;
[0086] S3: Use SpaCy for text annotation and basic NLP processing to identify and annotate power-related entities;
[0087] S4: Use Neo4j to design and build a knowledge graph in the power field, covering equipment, system, process concepts and their relationships:
[0088] S5: Based on the pre-trained large language basic model, a hybrid expert model architecture is designed to process different types of power data through different expert modules, improve the performance of the model on diversified tasks, and build a large power language model;
[0089] S6: Combine the knowledge graph with the power language model and enhance the model's reasoning ability through prompts or attention mechanisms, see Figure 3 , as follows:
[0090] S61: Use knowledge graph embedding methods (such as TransE, Node2Vec, GraphSAGE, etc.) to convert entities and relations in the knowledge graph into vector representations;
[0091] S62: embed the knowledge graph vector into the input of the power large language model as supplementary information to enhance the model's context perception and reasoning capabilities, and inject the extracted graph features into the specific layer of the corresponding expert model to enhance domain knowledge;
[0092] S63: Construct specific natural language prompts, combined with graph embedding, as contextual activation parameters for the language model;
[0093] S64: Generate dynamic prompts based on current input content and related knowledge graph node information;
[0094] S65: Add a cross-attention layer inside the power large language model to focus on important information obtained through graph embedding and prompts.
[0095] refer to Figure 2 In this embodiment, the text processing is specifically as follows:
[0096] S11: Collect text data related to the power sector, convert data in different formats (such as PDF, Word, HTML) into a unified text format, and use tools such as Apache Tika or PDFMiner to perform text extraction;
[0097] S12: Remove redundant information from the text, including HTML tags, footnotes, headers and footers; use regular expressions to clean up special characters, blank lines and irrelevant tags;
[0098] S13: Use the text similarity algorithm Jaccard similarity to detect and remove duplicate documents and paragraphs, and filter out content related to the power field through keyword filtering and topic modeling LDA.
[0099] In this embodiment, S13 is specifically:
[0100] Split each document or paragraph into a set of words and calculate the Jaccard similarity between each pair of documents or paragraphs;
[0101] Set a similarity threshold (such as 0.8). Text pairs above this threshold are considered duplicates and one of them is removed.
[0102] Use the LDA model to train the text and set the number of topics;
[0103] Analyze the topic distribution of each document, screen out the topics related to the power field, retain the documents related to the target topic, and filter out other content;
[0104] Create a keyword list related to the power field, traverse each document or paragraph, check whether it contains the words in the keyword list, retain the text containing the keywords, and filter out the irrelevant content.
[0105] In this embodiment, use the LDA model to train the text and set the number of topics; analyze the topic distribution of each document, screen out the topics related to the power field, retain the documents related to the target topic, and filter out other content, specifically as follows:
[0106] Split the text data into words or phrases to form a bag of words, and remove stop words (such as "de", "shi", etc.);
[0107] Use stemming to normalize the same vocabulary in different forms, count the frequency of each word in the document, and form a term frequency matrix;
[0108] According to the prior knowledge of the field, set the number of topics to K, and use the LDA model to train the documents to extract topics,
[0109] The goal of the LDA model is to find the topic distribution by maximizing the likelihood:
[0110] ;
[0111] ;
[0112] Among them, is the topic distribution of document d; is the topic word distribution of topic k; is the Dirichlet prior of the document topic distribution; is the Dirichlet prior of the topic word distribution; represents the Dirichlet distribution;
[0113] The generation process of each document:
[0114] ;
[0115] Among them, is the topic of the nth word in document d; is the nth word in document d; Multinomial represents the multinomial distribution;
[0116] Analyze the LDA model training results and mark the specific content reflected by each topic. Identify topics directly related to the power sector;
[0117] For each document, calculate its topic distribution and identify the weight of electricity-related topics;
[0118] Documents with power-related topic weights greater than the threshold are saved, and other content that does not match the target topic is filtered out.
[0119] In this embodiment, SeqGAN is used to generate diversified data samples based on the preprocessed text data, specifically:
[0120] Convert the preprocessed text data into a format suitable for SeqGAN input and train SeqGAN to generate diverse synthetic data;
[0121] The generator G is an RNN that generates text sequences. The goal of the generator is to generate realistic text sequences to deceive the discriminator. The generator loss is:
[0122] ;
[0123] in, is the loss function of the generator; represents the text sequence Y generated from the generator G; is the output of the discriminator for generating the text sequence Y; E represents the expected function;
[0124] The discriminator D is a classifier used to distinguish between real text sequences and generated text sequences. The goal of the discriminator is to accurately identify the generated text sequences;
[0125] ;
[0126] in, is the loss function of the discriminator; Represents the distribution of real data p data The real text sequence X sampled in is the logarithm of the discriminator's output for the true text sequence X.
[0127] In this embodiment, an adversarial mechanism is added to the SeqGAN generation step, as follows:
[0128] Build a knowledge base in the power sector, including key terms, standards, regulations, and typical data patterns, and extract key features and patterns from the knowledge base as constraints in the generation process; features include parameter ranges, typical operating modes, and compliance standards for power systems;
[0129] A conditional generation mechanism is introduced into the generator to ensure that the generated data meets the domain requirements. The generator formula is:
[0130] ;
[0131] in, is the previous word of the generated sequence, c is the condition vector, which contains specific constraints in the power field (such as terms, parameter ranges); is a generated word;
[0132] The specific characteristics of the power field are input into the generator as conditions to guide the generation process. It is necessary to add domain constraints to the loss function to penalize the generated data that does not meet the domain standards:
[0133] ;
[0134] in, is the domain constraint loss, λ is the weight coefficient;
[0135] A domain feature detection module is added to the discriminator to identify features that do not meet the standards of the power field, and the pre-trained domain feature detection module is used to fine-tune the discriminator.
[0136] In this embodiment, the power large language model includes a basic model, a routing layer and an expert module. The basic model uses a pre-trained large language model to process common language tasks, including semantic understanding, sentence generation, and general feature representation for further processing by the expert module;
[0137] The routing layer determines the best expert module based on the input data characteristics:
[0138] ;
[0139] in, is the probability of the input being assigned to the i-th expert; is the routing weight matrix; is the feature representation of the input;
[0140] The expert module includes several expert submodules, each of which focuses on a specific type of power data or task, including an equipment expert submodule for processing equipment-related data (such as sensor data, maintenance information); a system expert submodule for processing system-level data (such as load forecasting, system status monitoring); and a process expert submodule for processing process-related data (such as power generation process, transmission problems). Each expert submodule is a neural network, which is trained with specific task data.
[0141] Overall loss function:
[0142] ;
[0143] Where N is the number of expert submodules; α i is the weight of the loss of the i-th expert module; D i is the dataset prepared for the i-th expert; is the loss function of the i-th expert on his specific task; is the global regularization term; is the regularization parameter, is the expected function.
[0144] In this embodiment, the training process of the power language model is as follows:
[0145] Select the pre-trained large language model BERT as the base model;
[0146] Initialize the basic model and expert modules, and initialize the parameters through the pre-trained large language model; add specialized layers (for example, several neural network layers) to each expert module to handle its task, and the initial parameters of these layers are selected from the modules in the similar fields of the basic model;
[0147] Train the router to learn the mapping relationship between task characteristics and expert modules. Use a full dataset containing various tasks and data types to enable the router to learn to distinguish the characteristics of different task characteristics. Use the cross entropy loss function to optimize routing decisions so that it can correctly map task characteristics to appropriate expert modules.
[0148] Joint training is performed on the entire architecture so that routing decisions and expert modules are optimized together. In each training step, the parameters of the basic model, expert module and router are updated synchronously.
[0149] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0151] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0153] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning, characterized in that: The following steps are involved: S1: Collect text data related to the power field, and pre-process to remove redundant, irrelevant or repeated information to obtain pre-processed text data; S2: Based on the preprocessed text data, SeqGAN is used to generate diverse data samples, and an adversarial mechanism is added to the SeqGAN generation step; S3: Use SpaCy for text annotation and basic NLP processing to identify and annotate power-related entities; S4: Use Neo4j to design and build a knowledge graph in the power field, covering equipment, system, process concepts and their relationships; S5: Based on the pre-trained large language basic model, a hybrid expert model architecture is designed to process different types of power data through different expert modules, improve the performance of the model on diversified tasks, and build a large power language model; S6: Combine the knowledge graph with the power language model and enhance the reasoning ability of the model through prompts and attention mechanisms; The power large language model includes a basic model, a routing layer and an expert module. The basic model uses a pre-trained large language model to process common language tasks, including semantic understanding, sentence generation, and general feature representation for further processing by the expert module; The routing layer determines the best expert module based on the input data characteristics: ; in, is the probability of the input being assigned to the i-th expert; is the routing weight matrix; is the feature representation of the input; The expert module includes several expert submodules, each of which focuses on a specific power data type or task, including an equipment expert submodule for processing equipment-related data; a system expert submodule for processing system-level data; and a process expert submodule for processing process-related data. Each expert submodule is a neural network trained with specific task data. Overall loss function: ; Where N is the number of expert submodules; α i is the weight of the loss of the i-th expert module; D i is the dataset prepared for the i-th expert; is the loss function of the i-th expert on his specific task; is the global regularization term; is the regularization parameter, is the expected function; The adversarial mechanism is added to the SeqGAN generation step, as follows: Build a knowledge base in the power field, including key terms, standards, regulations and typical data patterns, and extract key features and patterns from the knowledge base as constraints in the generation process; A conditional generation mechanism is introduced into the generator to ensure that the generated data meets the domain requirements. The generator formula is: ; in, is the previous word of the generated sequence, c is the conditional vector containing specific constraints in the power field; is a generated word; The specific characteristics of the power field are input into the generator as conditions to guide the generation process. It is necessary to add domain constraints to the loss function to penalize the generated data that does not meet the domain standards: ; in, is the domain constraint loss, λ is the weight coefficient; Add a domain feature detection module to the discriminator to identify features that do not meet the standards of the power industry, and use the pre-trained domain feature detection module to fine-tune the discriminator; The S6 is specifically as follows: S61: Use the knowledge graph embedding method to convert entities and relations in the knowledge graph into vector representations; S62: embed the knowledge graph vector into the input of the power large language model as supplementary information to enhance the model's context perception and reasoning capabilities, and inject the extracted graph features into the specific layer of the corresponding expert model to enhance domain knowledge; S63: Construct specific natural language prompts, combined with graph embedding, as contextual activation parameters for the language model; S64: Generate dynamic prompts based on current input content and related knowledge graph node information; S65: Add a cross-attention layer inside the power large language model to focus on important information obtained through graph embedding and prompts.
2. The method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning according to claim 1 is characterized in that: The S1 is specifically: S11: Collect text data related to the power field, convert data in different formats into a unified text format, and perform text extraction; S12: Remove redundant information from the text, including HTML tags, footnotes, headers and footers; use regular expressions to clean up special characters, blank lines and irrelevant tags; S13: Use the text similarity algorithm Jaccard similarity to detect and remove duplicate documents and paragraphs, and filter out content related to the power field through keyword filtering and topic modeling LDA.
3. The method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning according to claim 2 is characterized in that: The S13 is specifically: Split each document or paragraph into a set of words and calculate the Jaccard similarity between each pair of documents or paragraphs; Set a similarity threshold. Text pairs above this threshold are considered duplicates and one of them is removed. Use the LDA model to train the text and set the number of topics; Analyze the topic distribution of each document, filter out topics related to the power field, retain documents related to the target topic, and filter out other content; Create a list of keywords related to the power sector, traverse each document or paragraph, check whether it contains words in the keyword list, retain the text containing the keywords, and filter out irrelevant content.
4. The method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning according to claim 3 is characterized in that: The LDA model is used to train the text and set the number of topics; the topic distribution of each document is analyzed, topics related to the power field are screened out, documents related to the target topic are retained, and other content is filtered out, as follows: Split text data into words or phrases, form bags of words, and remove stop words; Use stem extraction to normalize different forms of the same word, count the frequency of each word in the document, and form a word frequency matrix; According to the prior knowledge of the domain, the number of topics is set to K, and the LDA model is used to train documents to extract topics. The goal of the LDA model is to find the topic distribution by maximizing the likelihood: ; ; in, is the topic distribution of document d; is the topic word distribution of topic k; Dirichlet prior for document topic distribution; is the Dirichlet prior of the keyword distribution; Represents Dirichlet distribution; each document generation process: ; in, is the topic of the nth word in document d; is the nth word in document d; Multinomial represents multinomial distribution; Analyze the LDA model training results and mark the specific content reflected by each topic; identify topics directly related to the power field; For each document, calculate its topic distribution and identify the weight of electricity-related topics; Documents with power-related topic weights greater than the threshold are saved, and other content that does not match the target topic is filtered out.
5. The method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning according to claim 1 is characterized in that: Based on the preprocessed text data, SeqGAN is used to generate diversified data samples, specifically: Convert the preprocessed text data into a format suitable for SeqGAN input and train SeqGAN to generate diverse synthetic data; The generator G is an RNN that generates text sequences. The goal of the generator is to generate realistic text sequences to deceive the discriminator. The generator loss is: ; in, is the loss function of the generator; represents the text sequence Y generated from the generator G; is the output of the discriminator for generating the text sequence Y; E represents the expected function; The discriminator D is a classifier used to distinguish between real text sequences and generated text sequences. The goal of the discriminator is to accurately identify the generated text sequences; ; in, is the loss function of the discriminator; Represents the distribution of real data p data The real text sequence X sampled in is the logarithm of the discriminator's output for the true text sequence X.
6. The method for constructing a large electric power language model based on knowledge enhancement and adaptive fine-tuning according to claim 1 is characterized in that: The training process of the power language model is as follows: Select the pre-trained large language model BERT as the base model; Initialize the basic model and expert modules, and initialize the parameters through the pre-trained large language model; add specialized layers to each expert module to handle its task, and the initial parameters of these layers are selected from the modules in similar fields of the basic model; Train the router to learn the mapping relationship between task characteristics and expert modules. Use a full dataset containing various tasks and data types to enable the router to learn to distinguish the characteristics of different task characteristics. Use the cross entropy loss function to optimize routing decisions so that it can correctly map task characteristics to appropriate expert modules. Joint training is performed on the entire architecture so that routing decisions and expert modules are optimized together. In each training step, the parameters of the basic model, expert module and router are updated synchronously.
7. A system for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning, characterized in that: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the method for constructing a large power language model based on knowledge enhancement and adaptive fine-tuning as described in any one of claims 1 to 6.
8. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text auditing method, device and equipment, medium and program product
CN116484229A
Language processing question answering system and method based on AIGC large model
CN118093834A
Electric power vector knowledge base enhanced retrieval method and system based on artificial intelligence
CN118964648A