Electric power technology standard knowledge question and answer large model training data construction method and device

Through preprocessing of power technology standard documents, structured information extraction, dynamic optimization segmentation and multi-task multi-modal joint processing, the problem of low efficiency of power technology standard splitting in the existing technology is solved, efficient and accurate automatic splitting is achieved, and the performance of the search question-and-answer system is improved.

CN120162422APending Publication Date: 2025-06-17CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510342692.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The accurate splitting of power technology standards in the existing technology mainly relies on labor, which is low in efficiency and high in cost, making it difficult to meet practical application needs.

Method used

A large-scale training data construction method for power technology standard knowledge Q&A big model is adopted, including obtaining power technology standard documents, preprocessing documents, extracting structured information through GNN network, dynamically optimizing document segmentation, and performing multi-task and multi-modal joint processing to generate structured titles, abstracts and segmentation points.

Benefits of technology

It has achieved efficient and accurate automatic splitting of power technology standard documents, improved the performance of power technology standard retrieval Q&A system, reduced labor costs, improved the speed of information acquisition, and had significant socio-economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162422A_ABST
    Figure CN120162422A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of large model training, and particularly relates to a power technology standard knowledge question and answer large model training data construction method and device. The method comprises the following steps: acquiring a power technology standard document; preprocessing the power technology standard document to obtain a preprocessed power technology standard document; inputting the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; dynamically optimized document segmentation is carried out on the power technology standard document based on the structured information, and segmented document fragments are obtained; and performing multi-task and multi-mode combined processing on the segmented document segments to obtain training data of the power technology standard knowledge question-answer large model. According to the method, automatic splitting of the electric power technology standard document can be efficiently and accurately carried out, the performance of an electric power technology standard retrieval question-answering system is greatly improved, the labor cost is reduced, the information acquisition speed is increased, and remarkable social and economic benefits are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large model training, and particularly relates to a method and device for constructing training data of a large model for answering questions about electric power technical standards. Background Art

[0002] Electric power technical standards are unified regulations formulated for technical matters in the electric power field; they stipulate the requirements, test methods, inspection rules, etc. of electric power equipment, facilities, products, and related technologies. They have characteristics such as authority, normativity, scientificity, and practicality, and are formulated and issued by authoritative organizations such as governments, industry associations, or enterprises, and are binding on technical matters in the electric power field.

[0003] Electric power technical standards can be classified according to different dimensions, such as national standards, industry standards, enterprise standards, as well as mandatory standards and recommended standards. In addition, they can also be divided into basic standards, product standards, method standards, safety standards, etc.

[0004] The content of electric power technical standards involves the requirements for the design, manufacture, installation, operation, and maintenance of electric power equipment, as well as the specifications for the planning, construction, and operation of electric power systems. The formulation and implementation of electric power technical standards are of great significance and role in the development of the electric power industry: promoting technological innovation and industrial upgrading, standardizing the market order, improving economic benefits, ensuring safety, and promoting sustainable development.

[0005] With the intelligent and digital transformation of the electric power system, the formulation of electric power technical standards is developing towards a more scientific, rigorous, and advanced direction. In the fields of smart grid, photovoltaic power generation, and wind power grid connection, relevant research and formulation work on electric power technical standards are being actively carried out.

[0006] Electric power technical standards are an important support and guarantee for the development of the electric power industry. By formulating and implementing electric power technical standards, the healthy development of the electric power industry can be promoted, the safety and reliability of the electric power system can be improved, and energy efficiency and energy conservation and emission reduction can be promoted.

[0007] For the scenario of retrieving and answering questions about electric power technical standards, by processing the questions of users about technical standards, the corresponding answers and reference technical standard clauses are returned.

[0008] Power technology standard documents usually contain a large number of professional terms, charts, and flowcharts. Effective splitting and parsing of these documents are the key to building an efficient knowledge Q&A system. For the power technology standard retrieval Q&A model, it is necessary to split the power technology standards, and extract question-answer pairs based on the split documents as the model training data. Training the power technology standard retrieval Q&A model requires a large amount of training data, and the quantity of training data is also crucial for the performance of the model. Generally speaking, the more training data there is, the better the model performance. However, in the existing technology, most of the accurate splitting of power technology standards is still achieved manually, which has problems of high training cost and low efficiency, and it is difficult to meet the actual application requirements. Summary of the Invention

[0009] The purpose of the present invention is to provide a method and device for constructing training data for a large-scale power technology standard knowledge Q&A model to solve the technical problem of the low efficiency of the existing technology in accurately splitting power technology standards manually.

[0010] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for constructing training data for a large-scale power technology standard knowledge Q&A model, including: Obtain a power technology standard document; perform preprocessing on the power technology standard document to obtain a preprocessed power technology standard document; Input the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; Based on the structured information, perform dynamically optimized document segmentation on the power technology standard document to obtain segmented document fragments; Perform multi-task and multi-modal joint processing on the segmented document fragments to obtain training data for the large-scale power technology standard knowledge Q&A model.

[0011] A further improvement of the present invention lies in: in the step of performing preprocessing on the power technology standard document, the preprocessing specifically includes: Perform text tokenization, stop word removal, stemming, and image preprocessing on the power technology standard document; the image preprocessing includes one or more of scaling, cropping, and normalization; Among them, the power technology standard document includes text data and image data.

[0012] A further improvement of the present invention lies in: in the step of inputting the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information, the structured information includes: the hierarchical relationship and logical connection between different contents in the power technology standard document; The GNN network extracts the hierarchical relationships and logical connections between these different contents in the following ways: Node definition: Each paragraph, section, table, and chart in the power technology standard document is used as a node of the graph; for term definitions, test methods, and safety requirements, they are used as a single node separately; Edge definition: Edges are used to represent the relationships between nodes; connections are established between related content nodes to form edges; Feature representation: Based on the word embedding representation of the text content of each node, a feature representation is formed; Information passing: Through the information passing mechanism of the GNN network, a node receives information from its neighbor nodes, updates its own feature representation, and establishes the logical connections between various parts within the document; Hierarchical relationship modeling: For the hierarchical structure in the document, the relationship between the parent node and the child node is utilized to model the hierarchical relationship by designing the graph structure.

[0013] A further improvement of the present invention lies in: in the step where, through the information passing mechanism of the GNN network, a node receives information from its neighbor nodes, updates its own feature representation, and establishes the logical connections between various parts within the document, the node update formula of the GNN is expressed as:

[0014] Wherein, H (l) represents the output of the l layer, σ is the activation function, W (l) is the weight matrix connecting the l layer and the l +1 layer, N ( i ) is the neighbor set of node i , c i , m is the normalization constant, Hm (l) is the output of node m in the l layer. A further improvement of the present invention lies in: the step of performing dynamic optimization of the power technology standard document based on the structured information for document segmentation to obtain the segmented document fragments specifically includes: Initialize the spaces of state s and action a; wherein the state s includes the current processing position, the information of the segmented fragments, and the summary of the remaining unsegmented part; the action a includes moving the segmentation point forward, retreating backward, and ending the current segmentation operation; According to the current state s, calculate the keyword frequency and context information to determine the splitting granularity G; G = w1 * Fkw + w2 * Cct Among them, G is the splitting granularity. Fkw represents the influence of keyword frequency, which is obtained by calculating the average TF-IDF value of all keywords in the document. Cct represents the influence of context information, which is quantified by calculating the average cosine similarity of adjacent sentences. w1 and w2 are weighting coefficients. According to the current state s and the calculated splitting granularity G, select an action a; execute the action a to update the state s; calculate the reward R(s,a), and adjust the segmentation strategy according to the reward R(s,a) to obtain the segmented document fragments.

[0015] A further improvement of the present invention is that in the step of performing multi-task and multi-modal joint processing on the segmented document fragments to obtain the training data of the power technology standard knowledge Q&A large model, the training data of the power technology standard knowledge Q&A large model is structured titles, abstracts and segmentation points.

[0016] A further improvement of the present invention is that the step of performing multi-task and multi-modal joint processing on the segmented document fragments to obtain the training data of the power technology standard knowledge Q&A large model specifically includes: Use the VIT model to extract features from the image to generate image feature vectors; Use the BERT model to extract features from the text to generate text feature vectors; Through the alignment loss function L of multi-modal learning multi-modal , align the image feature vectors and text feature vectors to obtain the multi-modal fusion features; Perform multi-task learning based on the multi-modal fusion features; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; obtain structured titles, abstracts and segmentation points through the title recognition task, the text summarization task and the document segmentation task respectively.

[0017] A further improvement of the present invention is that in the step of performing multi-task and multi-modal joint processing on the segmented document fragments, the multi-task and multi-modal joint processing is performed through a pre-trained multi-task and multi-modal joint framework; The pre-trained multi-task and multi-modal joint framework is obtained through the following steps: Use the VIT model to extract features from the image to generate image feature vectors; use the BERT model to extract features from the text to generate text feature vectors; through the alignment loss function L of multi-modal learning multi-modal , align the image feature vectors and text feature vectors to obtain the multi-modal fusion features; Perform multi-task learning based on the features after multi-modal fusion; the multi-task learning includes a title recognition task, a text summarization task, and a document segmentation task; obtain structured titles, summaries, and segmentation points through the title recognition task, the text summarization task, and the document segmentation task respectively; Among them, the loss function of multi-task learning is expressed as:

[0018] Among them, L i is the loss function of the i-th task, w i is the corresponding weight, and N is the total number of tasks; The alignment loss function L multi-modal of multi-modal learning is expressed as:

[0019] Among them: I represents image data; T represents text data; L ViT (I) is the loss function of the vision transformer, and L BERT (T) is the loss function of the bidirectional encoder representation; L cross-modal (I,T) is the cross-modal loss function, which is used to align the feature representations of images and texts; Each time multi-task and multi-modal joint processing is performed, calculate the joint loss function once: L joint =L multi-task +L multi-modal ; Through the backpropagation algorithm, optimize the loss functions of multi-task and multi-modal at the same time, and adjust the parameters of the multi-task and multi-modal joint framework until the joint loss function converges, and obtain a pre-trained multi-task and multi-modal joint framework.

[0020] In a second aspect, the present invention provides a device for constructing training data for a large model of power technology standard knowledge Q&A, including: An acquisition module, configured to acquire a power technology standard document; preprocess the power technology standard document to obtain a preprocessed power technology standard document; A feature extraction module, configured to input the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; A segmentation module, configured to perform dynamically optimized document segmentation on the power technology standard document based on the structured information to obtain segmented document fragments; A joint processing module, configured to perform multi-task and multi-modal joint processing on the segmented document fragments to obtain training data for the large model of power technology standard knowledge Q&A.

[0021] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the processor is configured to execute a computer program stored in the memory to implement the method for constructing training data of the power technology standard knowledge Q&A large model.

[0022] In a fourth aspect, the present invention provides a computer-readable storage medium storing at least one instruction, and when the at least one instruction is executed by a processor, the method for constructing training data of the power technology standard knowledge Q&A large model is implemented.

[0023] In a fifth aspect, the present invention provides a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, the method for constructing training data of the power technology standard knowledge Q&A large model is implemented.

[0024] Compared with the prior art, the present invention has the following unexpected technical effects: The present invention provides a method for constructing training data of a power technology standard knowledge Q&A large model, including: obtaining a power technology standard document; preprocessing the power technology standard document to obtain a preprocessed power technology standard document; inputting the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; performing dynamically optimized document segmentation on the power technology standard document based on the structured information to obtain segmented document fragments; and performing multi-task and multi-modal joint processing on the segmented document fragments to obtain training data of the power technology standard knowledge Q&A large model. The present invention provides an efficient and accurate method for automatically splitting power technology standard documents, greatly improving the performance of the power technology standard retrieval Q&A system, reducing labor costs, accelerating the speed of information acquisition, and having significant social and economic benefits. Through the comprehensive application of dynamically optimized document segmentation, multi-task and multi-modal joint learning, and structured information extraction and adaptive context awareness technology, the present invention can better understand and process complex power technology standard documents, providing strong support for the informatization construction of the power industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 is a schematic flowchart of a method for constructing training data of a power technology standard knowledge Q&A large model according to an embodiment of the present invention; Figure 2 is a schematic diagram of a dynamically optimized document segmentation algorithm according to an embodiment of the present invention; Figure 3Schematic diagram of a multi-task and multi-modal joint learning framework according to an embodiment of the present invention; Figure 4 Schematic diagram of collaborative use among modules in a method for constructing training data of a large model for answering questions about power technology standards according to an embodiment of the present invention; Figure 5 Schematic structural diagram of an apparatus for constructing training data of a large model for answering questions about power technology standards according to an embodiment of the present invention; Figure 6 Schematic structural diagram of an electronic device according to the present invention. Detailed implementation manners

[0026] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0027] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise specified, all technical terms adopted by the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the present invention are only for describing specific implementation manners, and are not intended to limit the exemplary embodiments according to the present invention.

[0028] Please refer to Figure 1 As shown, an embodiment of the present invention provides a method for constructing training data of a large model for answering questions about power technology standards, and the specific steps are as follows: S1. Obtain a power technology standard document; perform preprocessing on the power technology standard document to obtain a preprocessed power technology standard document; S2. Input the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; S3. Perform dynamically optimized document segmentation on the power technology standard document based on the structured information to obtain segmented document fragments; S4. Perform multi-task and multi-modal joint processing on the segmented document fragments to obtain training data of a large model for answering questions about power technology standards.

[0029] In a specific implementation manner, in the step of performing preprocessing on the power technology standard document, the preprocessing specifically includes: Perform text tokenization, stop word removal, stemming, and image preprocessing on the power technology standard document; the image preprocessing includes one or more of scaling, cropping, and normalization; Among them, the power technology standard document includes text data and image data.

[0030] In a specific embodiment, in the step of inputting the preprocessed power technology standard document into the GNN network for feature extraction to obtain structured information, the structured information includes: the hierarchical relationship and logical connection between different contents in the power technology standard document; The GNN network extracts the hierarchical relationship and logical connection between these different contents in the following ways: Node definition: Each paragraph, section, table, and chart in the power technology standard document is used as a node of the graph; For term definitions, test methods, and safety requirements, they are used as a single node separately; Edge definition: Edges are used to represent the relationships between nodes; Connections are established between related content nodes to form edges; Feature representation: The feature representation is composed of the word embeddings of the text content of each node; Information transmission: Through the information transmission mechanism of the GNN network, a node receives information from its neighbor nodes and updates its own feature representation, establishing the logical connection between various parts of the document; Hierarchical relationship modeling: For the hierarchical structure in the document, the relationship between the parent node and the child node is used to model the hierarchical relationship by designing the graph structure.

[0031] In a specific embodiment, in the step of, through the information transmission mechanism of the GNN network, a node receives information from its neighbor nodes and updates its own feature representation, establishing the logical connection between various parts of the document, the node update formula of the GNN is expressed as:

[0032] Wherein, H (l) represents the output of the l th layer, σ is the activation function, W (l) is the weight matrix connecting the l th layer and the l th + 1 layer, N ( i ) is the neighbor set of node i , c i , m is the normalization constant, Hm (l) is the output of node m in the l th layer.

[0033] In a specific embodiment, the step of performing dynamic optimization document segmentation on the power technology standard document based on the structured information to obtain the segmented document fragments specifically includes: Initialize the spaces of state s and action a; where state s includes the current processing position, information of the segmented segments, and an overview of the remaining unsegmented part; action a includes moving the segmentation point forward, backing off, and ending the current segmentation operation. According to the current state s, calculate the keyword frequency and context information to determine the splitting granularity G. G = w1 * Fkw + w2 * Cct Where G is the splitting granularity, Fkw represents the influence of keyword frequency, obtained by calculating the average TF-IDF value of all keywords in the document; Cct represents the influence of context information, quantified by calculating the average cosine similarity of adjacent sentences; w1 and w2 are weighting coefficients. According to the current state s and the calculated splitting granularity G, select an action a; execute action a, update state s; calculate the reward R(s,a), and adjust the segmentation strategy according to the reward R(s,a) to obtain the segmented document segments.

[0034] In a specific embodiment, in the step of performing multi-task and multi-modal joint processing on the segmented document segments to obtain the training data of the power technology standard knowledge Q&A large model, the training data of the power technology standard knowledge Q&A large model is structured titles, abstracts, and segmentation points.

[0035] In a specific embodiment, the step of performing multi-task and multi-modal joint processing on the segmented document segments to obtain the training data of the power technology standard knowledge Q&A large model specifically includes: Use the VIT model to extract features from the image and generate image feature vectors. Use the BERT model to extract features from the text and generate text feature vectors. Through the alignment loss function L of multi-modal learning multi-modal , align the image feature vectors and text feature vectors to obtain the multi-modal fusion features. Perform multi-task learning based on the multi-modal fusion features; the multi-task learning includes title recognition task, text summarization task, and document segmentation task; obtain structured titles, abstracts, and segmentation points through the title recognition task, text summarization task, and document segmentation task respectively.

[0036] In a specific embodiment, in the step of performing multi-task and multi-modal joint processing on the segmented document segments, the multi-task and multi-modal joint processing is performed through a pre-trained multi-task and multi-modal joint framework. The pre-trained multi-task and multi-modal joint framework is obtained through the following steps: Use the VIT model to extract features from images and generate image feature vectors; use the BERT model to extract features from text and generate text feature vectors; and use the alignment loss function L of multimodal learning to extract features from text. multi-modal , align the image feature vector and the text feature vector to obtain the multimodal fusion features; Multi-task learning is performed based on the multimodal fusion features; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; structured titles, summaries and segmentation points are obtained through the title recognition task, the text summarization task and the document segmentation task respectively; Among them, the loss function of multi-task learning is expressed as:

[0037] Among them, L i is the loss function of the i-th task, w i is the corresponding weight, N is the total number of tasks; Alignment loss function L for multimodal learning multi-modal It is expressed as:

[0038] Where: I represents image data; T represents text data; L ViT (I) is the loss function of the visual transformer, L BERT (T) is the loss function of the bidirectional encoder representation; L cross-modal (I,T) is a cross-modal loss function used to align the feature representations of images and texts; Each time multi-task and multi-modal joint processing is performed, the joint loss function is calculated once: L joint =L multi-task +L multi-modal ; Through the back-propagation algorithm, the multi-task and multi-modal loss functions are optimized simultaneously, and the parameters of the multi-task and multi-modal joint framework are adjusted until the joint loss function converges to obtain the pre-trained multi-task and multi-modal joint framework.

[0039] The embodiment of the present invention provides a method for constructing training data of a large model of electric power technical standard knowledge question and answer, and the specific steps are as follows: First, through the dynamically optimized document segmentation algorithm, the appropriate splitting unit is automatically selected according to the content of the electric power technical standard document; for example, the electric power technical standard document is a national standard document on relay protection of power systems, with "Basic Principles of Relay Protection" as one section.

[0040] Next, using a multi-task and multi-modal joint learning framework, during the process of processing documents, the model not only learns how to accurately identify titles and generate summaries, but also can understand circuit diagrams and flowcharts in the documents through multi-modal learning, improving the accuracy of question answering.

[0041] Finally, through the graph neural network, the model successfully extracts the connection method and working principle of relay protection from the document, providing users with a more intuitive understanding.

[0042] In a specific embodiment, the core algorithm used in the method of the present invention is described: 1. Dynamic optimization of document segmentation algorithm; Please refer to Figure 2 As shown, the embodiment of the present invention provides a document segmentation algorithm combining dynamic adaptability and reinforcement learning. By analyzing the keyword frequency and context information features in the document, it automatically determines the most appropriate splitting unit size. On this basis, reinforcement learning technology is introduced, and by defining a reasonable reward function, the segmentation strategy is continuously optimized.

[0043] 1.1. Dynamic adaptive splitting First, the document is preprocessed, including steps of removing stop words and stemming. Then, by analyzing the keyword frequency and context information features in the document, the most appropriate splitting unit size is automatically determined. The formula is as follows: G = w1 * Fkw + w2 * Cct Where G is the splitting granularity, Fkw represents the influence of keyword frequency, which can be obtained by calculating the average TF-IDF value of all keywords in the document; Cct represents the influence of context information, which can be quantified by calculating the average cosine similarity of adjacent sentences. The higher the similarity, the smaller the Cct value, indicating that splitting is not appropriate here; w1 and w2 are weighting coefficients, which can be adjusted according to the actual application scenario to balance the influence of different factors on G.

[0044] 1.2. Reinforcement learning optimization By setting a reasonable reward function R(s,a), where s represents the current state and a represents the action taken, the model can continuously learn and optimize the segmentation strategy during the process of processing a large number of documents. The design of the reward function needs to consider factors such as the accuracy, efficiency, and complexity of segmentation. For example:

[0045] Where α, β, γ are weight parameters, representing the importance of accuracy, efficiency, and complexity respectively.

[0046] Reinforcement learning optimizes the segmentation strategy by defining a reasonable reward function R(s,a). Specifically, before each segmentation operation, the algorithm predicts possible results based on the current state s and the available actions a, and gives corresponding rewards according to the actual effects.

[0047] The state s can include the following information: The current processing position; Information of the segmented segments (such as keyword frequency, sentence length, etc.); Summary of the remaining unsegmented part.

[0048] The action a can include: Moving the segmentation point forward; Moving backward; Ending the current segmentation operation.

[0049] The reward function R(s,a) needs to consider factors such as segmentation accuracy, efficiency, and complexity. The specific formula is as follows:

[0050] Where: Accuracy: Whether the segmentation result can accurately reflect the logical structure of the document, without missing important information or wrongly interrupting a certain topic.

[0051] Eficiency: The time consumption of each segmentation operation and the overall time required to complete the segmentation.

[0052] Complexity: The computational resource consumption involved in the segmentation process, including memory usage and CPU utilization.

[0053] α, β, γ: Weight parameters, representing the importance of accuracy, efficiency, and complexity respectively.

[0054] 1.3 Specific implementation of the combination of dynamic adaptive splitting and reinforcement learning 1) Initialization: Preprocess the document, including removing stop words and stemming.

[0055] Initialize the spaces of the state s and the action a.

[0056] 2) Definition of state and action: State s: Includes the current processing position, information of the segmented segments (such as keyword frequency, sentence length, etc.), and summary of the remaining unsegmented part.

[0057] Action a: Includes moving the segmentation point forward, moving backward, and ending the current segmentation operation.

[0058] 3), Dynamic Adaptive Splitting: Based on the current state s, calculate the keyword frequency, sentence length distribution, and context information.

[0059] Calculate the current splitting granularity G.

[0060] 4), Reinforcement Learning Optimization: Based on the current state s and the calculated splitting granularity G, select an action a.

[0061] Execute the action a and update the state s.

[0062] Calculate the reward R(s,a) and adjust the policy according to the reward.

[0063] 5), Iterative Optimization: Through multiple iterations, continuously adjust the weight parameters α, β, γ, optimize the reward function R(s,a) until convergence, and obtain a trained model. Finally, the model can continuously learn and optimize the segmentation strategy during the process of processing a large number of documents.

[0064] 2. Multi-Task and Multi-Modal Joint Learning Framework; Please refer to Figure 3 As shown, the present invention adopts a multi-task learning framework to jointly train multiple tasks such as title recognition, text summarization, and document segmentation. By sharing the underlying feature representation, the model can fully utilize the correlation between different tasks and improve the overall performance. In addition, combined with multi-modal learning technology, text data is combined with non-text information such as images and charts, and a multi-modal deep learning framework (such as the ViT + BERT joint model) is used to achieve in-depth analysis of complex charts and flowcharts in power technology standards.

[0065] 2.1. Multi-Task Learning The loss function of multi-task learning is expressed as:

[0066] where L i is the loss function of the i-th task, w i is the corresponding weight, and N is the total number of tasks.

[0067] Specific tasks:

[0068] where is the true title label, is the title label predicted by the model. CE is the cross-entropy loss, which is used to measure the difference between the model prediction value and the true label.

[0069]

[0070] Among them, is the real text summary, is the text summary generated by the model. NLL is the Negative Log Likelihood Loss, which is used to measure the difference between the sequence generated by the model and the real sequence.

[0071]

[0072] Among them, is the real document segmentation label, is the document segmentation label predicted by the model. BCE is the Binary Cross-Entropy Loss, which is used to measure the difference between the predicted value of the model and the real label.

[0073] 2.2. Multimodal Learning In the processing of power technology standard documents, multimodal learning combines text data and image data (such as charts, flowcharts, etc.) to achieve in-depth analysis of complex content. The loss function of multimodal learning is expressed as:

[0074] Among them: I represents image data; T represents text data; L ViT (I) is the loss function of the Vision Transformer (ViT), which is used to process image data; L BERT (T) is the loss function of Bidirectional Encoder Representations from Transformers (BERT), which is used to process text data; L cross-modal (I,T) is the cross-modal loss function, which is used to align the feature representations of images and texts to ensure the consistency and complementarity of multimodal information.

[0075] Specific implementation: Image processing: Use the VIT model to extract features from images and generate image feature vectors.

[0076] Text processing: Use the BERT model to extract features from texts and generate text feature vectors.

[0077] Cross-modal alignment: Align image features and text features through the attention mechanism or alignment loss function to ensure the consistency and complementarity of multimodal information.

[0078] 2.3. Joint Optimization of Multimodal and Multitask Joint loss function: L joint =L multi-task +L multi-modal Among them, L multi-taskis the loss function for multi-task learning, L multi-modal is the loss function for multi-modal learning.

[0079] Training process: Input multi-modal data (text and images) into a shared Transformer encoder to extract multi-modal features; separately pass the multi-modal features to the specific output layers of each task to calculate the losses of each task; through the backpropagation algorithm, simultaneously optimize the loss functions of multi-task and multi-modal, and adjust the model parameters.

[0080] 3. Structured Information Extraction Algorithm The present invention uses graph neural network (GNN) technology to process the structured information in the document, and effectively extracts the hierarchical relationships and logical connections between different parts in the document.

[0081] In power technology standard documents, there are usually a large number of normative reference documents, term definitions, test methods, safety requirements, etc. There are complex hierarchical relationships and logical connections between these contents. For example, the relationship between term definitions and test methods, and the association between safety requirements and specific operation steps. The graph neural network (GNN) can help extract these relationships in the following ways: Node definition: First, regard each paragraph, section, table, and chart in the document as nodes of the graph. For the key concepts of term definitions, test methods, and safety requirements, they can also be regarded as a single node separately.

[0082] Edge definition: Edges are used to represent the relationships between nodes. For example, if a term definition appears in the description of a certain test method, then an edge can be established between the node representing the term definition and the node representing the test method. Similarly, if a certain safety requirement applies to multiple operation steps, then a connection can be established between the safety requirement node and these operation step nodes.

[0083] Feature representation: The initial feature of each node can be based on the word embedding representation of its text content. For term definitions, domain knowledge can also be added, such as the professional vocabulary extracted from power technology standards, to enhance the feature representation of the node.

[0084] Information transmission: Through the information transmission mechanism of GNN, a node can receive information from its neighbor nodes and update its own feature representation accordingly. This process can be iterated, so that each node not only contains its own information, but also integrates the information of other parts associated with it, thereby establishing the logical connection between various parts in the document.

[0085] Hierarchical relationship modeling: For the hierarchical structure in a document, such as chapter headings and subheadings, the relationship between parent nodes and child nodes can be utilized to model this hierarchical relationship by designing a graph structure. For example, when constructing the graph, it can be ensured that all child nodes are directly connected to their parent nodes to reflect the organizational structure of the document.

[0086] The node update formula of GNN is expressed as:

[0087] Where, H (l) represents the output of the l th layer, σ is the activation function, W (l) is the weight matrix connecting the l th layer and the l +1 th layer, N ( i ) is the neighbor set of node i , c i , m is the normalization constant, Hm (l) is the output of node m in the l th layer.

[0088] 4. Synergy between modules; Please refer to Figure 4 as shown, the dynamically optimized document segmentation algorithm provides structured input data for the multi-task and multi-modal joint learning framework, which helps to improve the processing efficiency and accuracy of the latter.

[0089] The multi-task and multi-modal joint learning framework enhances the ability of structured information extraction by processing various types of data such as text and images, enabling it to understand the document content more accurately.

[0090] The structured information extraction algorithm, on the other hand, provides a structured information output and feeds it back to the previous two modules to help them better understand the logical structure and professional terms of the document, forming a closed-loop optimization.

[0091] The structured information extraction algorithm can identify key parts in the document, such as titles, subsections, list items, etc., and provide the boundary information of these parts. This helps the document segmentation algorithm to divide different parts of the document more accurately, especially when the document contains complex formats or non-standard layouts. In addition, by understanding the logical structure of the document, the segmentation algorithm can handle parts with fuzzy boundaries more intelligently, such as merging closely related paragraphs or splitting long paragraphs reasonably.

[0092] The information provided by the structured information extraction algorithm can help the framework better understand the logical structure and professional terms of the document. For example, by providing the location of the term definition and its relationship with other parts, it can assist tasks in the multi-task learning framework (such as term explanation, test method parsing, etc.) to be executed more precisely. At the same time, these structured information can also be used as additional input features to improve the performance of the model when performing tasks such as text summarization and title generation.

[0093] Please refer to Figure 5 As shown, an embodiment of the present invention provides a device for constructing training data of a large model for answering questions about power technology standards, including: An acquisition module, configured to acquire a power technology standard document; perform preprocessing on the power technology standard document to obtain a preprocessed power technology standard document; A feature extraction module, configured to input the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; A segmentation module, configured to perform dynamically optimized document segmentation on the power technology standard document based on the structured information to obtain segmented document fragments; A joint processing module, configured to perform multi-task and multi-modal joint processing on the segmented document fragments to obtain training data of a large model for answering questions about power technology standards.

[0094] Please refer to Figure 6 As shown, an embodiment of the present invention provides an electronic device 100 for implementing a method for constructing training data of a large model for answering questions about power technology standards; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0095] The memory 101 can be used to store the computer program 103. By running or executing the computer program stored in the memory 101 and invoking the data stored in the memory 101, the processor 102 implements the steps of the method for constructing training data of the power technology standard knowledge Q&A large model. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0096] The at least one processor 102 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.

[0097] The memory 101 in the electronic device 100 stores multiple instructions to implement a method for constructing training data of the power technology standard knowledge Q&A large model. The processor 102 can execute the multiple instructions to implement: Obtain a power technology standard document; preprocess the power technology standard document to obtain a preprocessed power technology standard document; Input the preprocessed power technology standard document into a GNN network for feature extraction to obtain structured information; Perform dynamically optimized document segmentation on the power technology standard document based on the structured information to obtain segmented document fragments; The segmented document fragments are subjected to multi-task and multi-modal joint processing to obtain the training data for the large model of power technology standard knowledge Q&A.

[0098] If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).

[0099] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0100] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0101] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the processes in Figure 1one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.

[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or boxes Figure 1 one box or multiple boxes.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for constructing training data for a large model of electric power technical standards knowledge question and answer, characterized in that: include: Obtain power technical standard documents; Preprocessing the electric power technical standard document to obtain the preprocessed electric power technical standard document; The pre-processed power technical standard documents are input into the GNN network for feature extraction to obtain structured information; Based on structured information, the document of electric power technical standard is segmented by dynamic optimization to obtain segmented document fragments; The segmented document fragments are processed jointly in a multi-task and multi-modal manner to obtain the training data of the large model of electric power technical standards knowledge question and answer.

2. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1 is characterized in that: In the step of preprocessing the electric power technical standard document, the preprocessing specifically includes: Performing text segmentation, stop word removal, stem extraction, and image preprocessing on the electric power technical standard document; the image preprocessing includes one or more of scaling, cropping, and normalization; Wherein, the electric power technical standard document includes text data and image data.

3. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1 is characterized in that: Inputting the preprocessed electric power technical standard document into the GNN network for feature extraction to obtain structured information, wherein the structured information includes: hierarchical relationships and logical connections between different contents in the electric power technical standard document; The GNN network extracts the hierarchical relationships and logical connections between these different contents in the following ways: Node definition: each paragraph, chapter, table, and chart in the power technical standard document is used as a node of the diagram; term definitions, test methods, and safety requirements are treated as separate nodes; Definition of edge: Edge is used to represent the relationship between nodes; the connection between related content nodes constitutes an edge; Feature representation: The feature representation is constructed based on the word embedding representation of the text content of each node; Information transmission: Through the information transmission mechanism of the GNN network, a node receives information from its neighboring nodes and updates its own feature representation to establish a logical connection between the various parts of the document; Hierarchical relationship modeling: For the hierarchical structure in the document, the relationship between parent nodes and child nodes is used to model the hierarchical relationship by designing a graph structure.

4. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 3 is characterized in that: In the step of establishing a logical connection between the parts of a document through the information transmission mechanism of the GNN network, a node receives information from its neighboring nodes and updates its own feature representation, the node update formula of the GNN is expressed as: in, H (l) Indicates l The output of the layer, σ is the activation function, W (l) Is the connection l Layer and l +1 layer weight matrix, N ( i ) is a node i The neighbor set of c i , m is the normalization constant, H (l) It is l The output of node m in the layer.

5. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1 is characterized in that: The step of dynamically optimizing the document segmentation of the electric power technical standard document based on the structured information to obtain the segmented document fragments specifically includes: Initialize the space of state s and action a; state s includes the current processing position, information of the segmented segments, and a summary of the remaining unsegmented parts; action a includes moving the segmentation point forward, rolling back, and ending the current segmentation operation; According to the current state s, the keyword frequency and context information are calculated to determine the split granularity G; G = w1 * Fkw + w2 * Cct Where G is the split granularity, Fkw represents the influence of keyword frequency, which is obtained by calculating the average TF-IDF value of all keywords in the document; Cct represents the influence of context information, which is quantified by calculating the average cosine similarity of adjacent sentences; w1, w2 are weighting coefficients; According to the current state s and the calculated split granularity G, select an action a; execute action a and update the state s; calculate the reward R(s,a), and adjust the segmentation strategy according to the reward R(s,a) to obtain the segmented document fragments.

6. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1, characterized in that: The segmented document fragments are subjected to multi-task and multi-modal joint processing to obtain the large model training data of electric power technical standards knowledge question and answer, wherein the large model training data of electric power technical standards knowledge question and answer is structured title, abstract and segmentation point.

7. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1, characterized in that: The steps of performing multi-task and multi-modal joint processing on the segmented document segments to obtain training data for the large model of electric power technical standards knowledge question and answer include: Use the VIT model to extract features from the image and generate image feature vectors; Use the BERT model to extract features from text and generate text feature vectors; Alignment loss function L through multimodal learning multi-modal , align the image feature vector and the text feature vector to obtain the multimodal fusion features; Multi-task learning is performed based on the features after multimodal fusion; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; structured titles, summaries and segmentation points are obtained respectively through the title recognition task, the text summarization task and the document segmentation task.

8. The method for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 7, characterized in that: In the step of performing multi-task and multi-modal joint processing on the segmented document segments, the multi-task and multi-modal joint processing is performed by a pre-trained multi-task and multi-modal joint framework; The pre-trained multi-task and multi-modal joint framework is trained by the following steps: Use the VIT model to extract features from images and generate image feature vectors; use the BERT model to extract features from text and generate text feature vectors; and use the alignment loss function L of multimodal learning to extract features from text. multi-modal , align the image feature vector and the text feature vector to obtain the multimodal fusion features; Multi-task learning is performed based on the multimodal fusion features; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; structured titles, summaries and segmentation points are obtained through the title recognition task, the text summarization task and the document segmentation task respectively; Among them, the loss function of multi-task learning is expressed as: Among them, L i is the loss function of the i-th task, w i is the corresponding weight, N is the total number of tasks; Alignment loss function L for multimodal learning multi-modal It is expressed as: Where: I represents image data; T represents text data; L ViT (I) is the loss function of the visual transformer, L BERT (T) is the loss function of the bidirectional encoder representation; L cross-modal (I,T) is a cross-modal loss function used to align the feature representations of images and texts; Each time multi-task and multi-modal joint processing is performed, the joint loss function is calculated once: L joint =L multi-task +L multi-modal ; Through the back-propagation algorithm, the multi-task and multi-modal loss functions are optimized simultaneously, and the parameters of the multi-task and multi-modal joint framework are adjusted until the joint loss function converges to obtain the pre-trained multi-task and multi-modal joint framework.

9. A device for constructing training data for a large model of electric power technical standard knowledge question and answer, characterized in that: include: An acquisition module is used to acquire power technical standard documents; Preprocessing the electric power technical standard document to obtain the preprocessed electric power technical standard document; The feature extraction module is used to input the pre-processed power technical standard document into the GNN network for feature extraction to obtain structured information; A segmentation module is used to dynamically optimize the document segmentation of the electric power technical standard document based on structured information to obtain segmented document fragments; The joint processing module is used to perform multi-task and multi-modal joint processing on the segmented document fragments to obtain the training data of the large model of electric power technical standards knowledge question and answer.

10. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 9, characterized in that: In the step of preprocessing the electric power technical standard document, the preprocessing specifically includes: Performing text segmentation, stop word removal, stem extraction, and image preprocessing on the electric power technical standard document; the image preprocessing includes one or more of scaling, cropping, and normalization; Wherein, the electric power technical standard document includes text data and image data.

11. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 9, characterized in that: Inputting the preprocessed electric power technical standard document into the GNN network for feature extraction to obtain structured information, wherein the structured information includes: hierarchical relationships and logical connections between different contents in the electric power technical standard document; The GNN network extracts the hierarchical relationships and logical connections between these different contents in the following ways: Node definition: each paragraph, chapter, table, and chart in the power technical standard document is used as a node of the diagram; term definitions, test methods, and safety requirements are treated as separate nodes; Definition of edge: Edge is used to represent the relationship between nodes; the connection between related content nodes constitutes an edge; Feature representation: The feature representation is constructed based on the word embedding representation of the text content of each node; Information transmission: Through the information transmission mechanism of the GNN network, a node receives information from its neighboring nodes and updates its own feature representation to establish a logical connection between the various parts of the document; Hierarchical relationship modeling: For the hierarchical structure in the document, the relationship between parent nodes and child nodes is used to model the hierarchical relationship by designing a graph structure.

12. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 11, characterized in that: In the step of establishing a logical connection between the parts of a document through the information transmission mechanism of the GNN network, a node receives information from its neighboring nodes and updates its own feature representation, the node update formula of the GNN is expressed as: in, H (l) Indicates l The output of the layer, σ is the activation function, W (l) Is the connection l Layer and l +1 layer weight matrix, N ( i ) is a node i The neighbor set of c i , m is the normalization constant, H (l) It is l The output of node m in the layer.

13. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 9, characterized in that: The step of dynamically optimizing the document segmentation of the electric power technical standard document based on the structured information to obtain the segmented document fragments specifically includes: Initialize the space of state s and action a; state s includes the current processing position, information of the segmented segments, and a summary of the remaining unsegmented parts; action a includes moving the segmentation point forward, rolling back, and ending the current segmentation operation; According to the current state s, the keyword frequency and context information are calculated to determine the split granularity G; G = w1 * Fkw + w2 * Cct Where G is the split granularity, Fkw represents the influence of keyword frequency, which is obtained by calculating the average TF-IDF value of all keywords in the document; Cct represents the influence of context information, which is quantified by calculating the average cosine similarity of adjacent sentences; w1, w2 are weighting coefficients; According to the current state s and the calculated split granularity G, select an action a; execute action a and update the state s; calculate the reward R(s,a), and adjust the segmentation strategy according to the reward R(s,a) to obtain the segmented document fragments.

14. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 9, characterized in that: The segmented document fragments are subjected to multi-task and multi-modal joint processing to obtain the large model training data of electric power technical standards knowledge question and answer, wherein the large model training data of electric power technical standards knowledge question and answer is structured title, abstract and segmentation point.

15. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 1, characterized in that: The steps of performing multi-task and multi-modal joint processing on the segmented document segments to obtain training data for the large model of electric power technical standards knowledge question and answer include: Use the VIT model to extract features from the image and generate image feature vectors; Use the BERT model to extract features from text and generate text feature vectors; Alignment loss function L through multimodal learning multi-modal , align the image feature vector and the text feature vector to obtain the multimodal fusion features; Multi-task learning is performed based on the features after multimodal fusion; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; structured titles, summaries and segmentation points are obtained respectively through the title recognition task, the text summarization task and the document segmentation task.

16. The device for constructing training data of a large model of electric power technical standard knowledge question and answer according to claim 15, characterized in that: In the step of performing multi-task and multi-modal joint processing on the segmented document segments, the multi-task and multi-modal joint processing is performed by a pre-trained multi-task and multi-modal joint framework; The pre-trained multi-task and multi-modal joint framework is trained by the following steps: Use the VIT model to extract features from images and generate image feature vectors; use the BERT model to extract features from text and generate text feature vectors; and use the alignment loss function L of multimodal learning to extract features from text. multi-modal , align the image feature vector and the text feature vector to obtain the multimodal fusion features; Multi-task learning is performed based on the multimodal fusion features; the multi-task learning includes a title recognition task, a text summarization task and a document segmentation task; structured titles, summaries and segmentation points are obtained through the title recognition task, the text summarization task and the document segmentation task respectively; Among them, the loss function of multi-task learning is expressed as: Among them, L i is the loss function of the i-th task, w i is the corresponding weight, N is the total number of tasks; Alignment loss function L for multimodal learning multi-modal It is expressed as: Where: I represents image data; T represents text data; L ViT (I) is the loss function of the visual transformer, L BERT (T) is the loss function of the bidirectional encoder representation; L cross-modal (I,T) is a cross-modal loss function used to align the feature representations of images and texts; Each time multi-task and multi-modal joint processing is performed, the joint loss function is calculated once: L joint =L multi-task +L multi-modal ; Through the back-propagation algorithm, the multi-task and multi-modal loss functions are optimized simultaneously, and the parameters of the multi-task and multi-modal joint framework are adjusted until the joint loss function converges to obtain the pre-trained multi-task and multi-modal joint framework.

17. An electronic device, characterized in that: It includes a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the method for constructing training data for a large model of knowledge questions and answers of electric power technical standards as described in any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the method for constructing training data for the large model of knowledge questions and answers of electric power technical standards as described in any one of claims 1 to 8.

19. A computer program product, characterized in that The computer program product includes a computer program / instruction, which, when executed by a processor, implements the method for constructing training data for a large model of knowledge questions and answers of electric power technical standards as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Electric power knowledge retrieval system based on large language model

    CN120910228A