Power switch equipment control large model training method and device, equipment, storage medium and program product

By employing a large-scale model training method for power switchgear control, utilizing parallel training and gradient balancing mechanisms, and combining electrical rules and visual language models, the multi-task adaptability problem of traditional power switchgear control in complex power grid scenarios is solved, achieving efficient, accurate, and safe operation and maintenance control of power equipment.

CN121765385APending Publication Date: 2026-03-31SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing power switchgear control technologies struggle to handle complex tasks in highly dynamic, multi-variable coupled power grid scenarios. Traditional methods rely on static rules and low-dimensional features, which cannot adapt to the multiple task requirements of power equipment operation and maintenance.

Method used

A large-scale model training method for power switchgear control is adopted. The pre-trained model is trained in parallel by acquiring electrical images and problem texts. The gradient is calculated using a gradient equalization mechanism and the model parameters are updated synchronously. The data is labeled and augmented by combining electrical rules and visual language models to achieve multi-task parallel training and safety constraint loss and adversarial training.

Benefits of technology

It improves the accuracy and efficiency of the control model for power switching equipment, enabling it to efficiently handle complex tasks in the operation and maintenance of power equipment, adapt to multi-variable coupled scenarios, and ensure safety and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765385A_ABST
    Figure CN121765385A_ABST
Patent Text Reader

Abstract

The invention relates to a power switch equipment control large model training method and device, equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a power switch equipment control training data set; each data sample group in the power switch equipment control training data set comprises an electrical image and a problem text; using each electrical image and the corresponding problem text to parallelly train a pre-training model deployed on each group of computing nodes to obtain at least two prediction sequences of each pre-training model; wherein each prediction sequence corresponds to one prediction task; for each pre-training model, calculating the gradient of each group of calculation nodes based on a gradient equalization mechanism and at least two prediction sequences; and synchronizing the gradient of each group of calculation nodes, and updating the parameters of the pre-training model by using the synchronized gradient until the pre-training model meets a termination condition, thereby obtaining the large power switch equipment control model. By adopting the method, complex tasks under the operation and maintenance condition of the power equipment can be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for training a large-scale control model of power switching equipment. Background Technology

[0002] With the rapid evolution of smart grids, the demand for intelligent control of power system switching equipment is becoming increasingly urgent.

[0003] Control technologies for power switching equipment primarily rely on control strategies based on human experience and rules, classic PID algorithms, and shallow machine learning models (such as support vector machines and random forests). These methods can achieve basic control functions under steady-state operation or simple fault scenarios in the power grid. Their core logic involves generating and responding to switching commands through a pre-set rule base or static parameter configuration. For example, rule-based strategies rely on expert knowledge to define action thresholds, PID algorithms achieve steady-state tracking through linear feedback regulation, and shallow machine learning models are trained on limited labeled data for classification or regression tasks. These technologies generally employ offline training, with model parameters fixed and deployed in a fixed hardware environment. Data processing is concentrated on single electrical quantities (such as voltage and current) or simplified environmental variables, forming a traditional control framework centered on static rules and low-dimensional features.

[0004] Although these methods have basic control capabilities under steady-state conditions, they are still limited to single-task models in complex power grid scenarios with high dynamics and multivariate coupling, making it difficult to handle complex tasks in power equipment operation and maintenance. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for training a large-scale control model of power switching equipment that can handle complex tasks in the operation and maintenance of power equipment, in order to address the above-mentioned technical problems.

[0006] Firstly, this application provides a method for training a large-scale control model of power switching equipment, the method comprising:

[0007] Obtain a training dataset for power switchgear control; wherein each data sample group in the power switchgear control training dataset includes an electrical image and a question text;

[0008] The pre-trained model deployed on each set of computing nodes is trained in parallel using each electrical image and corresponding question text in the power switchgear control training dataset to obtain at least two prediction sequences for each pre-trained model; wherein each prediction sequence corresponds to a prediction task.

[0009] For each pre-trained model, the gradient of each set of computation nodes is calculated based on the gradient equalization mechanism and the at least two prediction sequences.

[0010] The gradients of each group of computing nodes are synchronized, and the parameters of the pre-trained model are updated using the synchronized gradients until the pre-trained model meets the termination condition, thus obtaining the large control model for power switchgear.

[0011] In one embodiment, obtaining the power switchgear control training dataset includes:

[0012] Acquire electrical signals, environmental parameters, and internal state diagrams of power switchgear; convert the electrical signals into waveform diagrams; and convert the environmental parameters into heat maps.

[0013] Based on the waveform overlaid with the thermogram, a composite electrical feature image is determined;

[0014] Based on each of the composite electrical feature images and the corresponding internal state diagram, a first image text group is determined;

[0015] The first image text group is labeled using electrical rules and a visual language model to determine the second image text group;

[0016] The second image-text group is enhanced to obtain a data sample group; wherein the data sample group includes electrical images and question text.

[0017] In one embodiment, the step of annotating the first image text group using electrical rules and a visual language model to determine the second image text group includes:

[0018] The images in the first image-text group are analyzed using a visual language model to obtain descriptive text;

[0019] The description text is modified based on at least one of the physical rules, safety constraints, and action timing interlocking logic in the electrical rules, and the modified text is determined.

[0020] The second image text group is determined based on the corrected text and the images in the first image text group.

[0021] In one embodiment, enhancing the second image text group to obtain a data sample group includes:

[0022] The electrical images in the second image text group are subjected to geometric transformation, noise simulation, or illumination perturbation enhancement to obtain enhanced electrical images;

[0023] The text in the second image text group is enhanced based on the power knowledge graph and reference electrical parameters to obtain the enhanced text.

[0024] Based on the enhanced electrical image and the enhanced text, a data sample group is determined.

[0025] In one embodiment, the step of using each electrical image and corresponding question text in the power switching equipment control training dataset to train a pre-trained model deployed on each set of computing nodes in parallel includes:

[0026] The power switchgear control training dataset is divided into power switchgear control training subsets; wherein the number of power switchgear control training subsets is the same as the number of computing node groups;

[0027] In each set of computing nodes, the pre-trained model is divided into model stages according to the network layer order through pipeline parallelism; the power switch equipment control training sub-dataset is divided into multiple power switch equipment control training micro-datasets according to the number of model stages.

[0028] Based on the number of computing nodes and the number of model stages, the number of parallel tensors for each model stage is determined; the network layer of the pre-trained model is trained using each electrical image and corresponding question text in the power switch equipment control training sub-dataset according to the number of parallel tensors, the calculation results are obtained, and the calculation results are passed in the order of the network layers.

[0029] In one embodiment, each data sample group further includes the answer text corresponding to the question text; the step of calculating the gradient of each set of computation nodes based on the gradient equalization mechanism and the at least two prediction sequences includes:

[0030] Calculate the task loss between each of the predicted sequences and the answer text;

[0031] The total loss is determined based on the task loss, security constraint loss, and adversarial training loss; wherein the security constraint loss represents the difference between the features of the data sample group and the predefined security feature library; and the adversarial training loss represents the difference between the answer text output based on the data sample group and the answer text output based on the second image text group.

[0032] Calculate the gradient norm of each of the task loss, security constraint loss, and adversarial loss, and adjust the weights of the task loss, security constraint loss, and adversarial loss in the total loss according to the gradient norm.

[0033] Based on the task loss, security constraint loss, adversarial loss, and their corresponding weights, the weighted total loss is determined.

[0034] The gradient of each set of computation nodes is calculated based on the weighted total loss.

[0035] Secondly, this application also provides a large-scale model training device for power switchgear control, the device comprising:

[0036] The acquisition module is used to acquire a power switchgear control training dataset; wherein, each data sample group in the power switchgear control training dataset includes an electrical image and a question text;

[0037] The training module is used to train the pre-trained model deployed on each set of computing nodes in parallel using each electrical image and corresponding question text in the power switchgear control training dataset, to obtain at least two prediction sequences for each pre-trained model; wherein each prediction sequence corresponds to a prediction task.

[0038] The gradient calculation module is used to calculate the gradient of each set of computing nodes for each pre-trained model, based on the gradient equalization mechanism and the at least two prediction sequences.

[0039] The update module is used to synchronize the gradients of each group of computing nodes and update the parameters of the pre-trained model using the synchronized gradients until the pre-trained model meets the termination condition, thus obtaining the large control model for power switchgear.

[0040] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0041] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0042] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0043] The aforementioned training method, apparatus, computer equipment, computer-readable storage medium, and computer program product for a large-scale power switchgear control model firstly acquires a power switchgear control training dataset. Each data sample group in the training dataset includes an electrical image and a question text. Using each electrical image and corresponding question text from the training dataset, a pre-trained model deployed on each set of computing nodes is trained in parallel, resulting in at least two prediction sequences for each pre-trained model. Each prediction sequence corresponds to a prediction task. By using a specially constructed power switchgear dataset as the training set, a highly accurate dedicated large-scale model can be trained in parallel, improving training efficiency. Secondly, for each pre-trained model, based on a gradient balancing mechanism and at least two prediction sequences, the gradient of each set of computing nodes is calculated. This trains a multi-task model capable of adapting to the complex tasks of power equipment operation and maintenance. Finally, the gradients of each set of computing nodes are synchronized, and the parameters of the pre-trained model are updated using the synchronized gradients until the pre-trained model meets the termination condition, resulting in a large-scale power switchgear control model. This model can efficiently and accurately train a batch of dedicated large-scale models to handle complex tasks in power equipment operation and maintenance. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating a method for training a large-scale control model of power switching equipment in one embodiment.

[0046] Figure 2 This is a schematic diagram of the process for obtaining a power switchgear control training dataset in one embodiment;

[0047] Figure 3 This is a schematic diagram of a process for annotating a first image text group using electrical rules and a visual language model in one embodiment;

[0048] Figure 4 This is a schematic diagram of the process for enhancing a second image text group in one embodiment;

[0049] Figure 5 This is a flowchart illustrating the process of training a pre-trained model deployed on each set of computing nodes in parallel using each electrical image and the corresponding question text in one embodiment.

[0050] Figure 6This is a flowchart illustrating the process of calculating the gradient of each set of computing nodes in one embodiment.

[0051] Figure 7 This is a structural block diagram of a large-scale training device for controlling power switching equipment in one embodiment;

[0052] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] In one embodiment, such as Figure 1 As shown, a method for training a large-scale control model of power switching equipment is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps S102 to S108. Wherein:

[0055] Step S102: Obtain the power switchgear control training dataset.

[0056] Each data sample group in the power switchgear control training dataset includes an electrical image and a question text. Each data sample group may also include answer texts corresponding to the question texts; there can be multiple question texts, and similarly, there can be multiple corresponding answer texts.

[0057] Optionally, the terminal constructs a power switchgear control training dataset and obtains the completed power switchgear control training dataset.

[0058] Step S104: Use each electrical image and corresponding question text in the power switch equipment control training dataset to train the pre-trained model deployed on each set of computing nodes in parallel, and obtain at least two prediction sequences for each pre-trained model.

[0059] Each prediction sequence corresponds to a prediction task. The pre-trained model is the initial model for the large-scale control model of power switching equipment.

[0060] Optionally, the terminal uses the electrical images and corresponding question texts of each sample data group in the constructed power switchgear control training dataset to train the pre-trained model deployed on each set of computing nodes in parallel, obtaining prediction sequences corresponding to at least two prediction tasks for each pre-trained model. For example, task a is state description generation, task b is fault diagnosis, and task c is control command generation. Here, a set of computing nodes includes multiple devices, and there are also multiple sets of computing nodes.

[0061] Pre-trained models can be fine-tuned in a fully supervised manner. For a set of data samples... and the corresponding real sequence The model predicts the output sequence as follows: Each y is the word probability distribution of the model output at time step t. The expression for the cross-entropy loss function is shown in formula (1).

[0062] Formula (1)

[0063] In the formula: Represents a given previous input sequence and model parameters In the case of correctly predicting the t-th target word y t The probability of.

[0064] Each training sample consists of three parts: an image, a question, and its corresponding answer. The system first performs multimodal encoding on the input data: image data is processed by a visual encoder to extract feature representations, and the question and answer texts are transformed into discrete token sequences through word segmentation and embedding operations. Subsequently, the image features and text tokens are concatenated in a preset order to form a unified input sequence, enabling the model to perceive both visual and textual information simultaneously within the same semantic space.

[0065] During model training, the terminal calculates the prediction loss only for the tokens in the answer portion. To achieve this, the model adds an autoregressive mask to the input sequence, allowing it to generate the current token based solely on its preceding tokens and image features, thus enabling conditional generation. The calculated loss value reflects the model's prediction error on the current sample.

[0066] The cross-entropy loss function effectively measures the difference between the model's predicted distribution and the true target sequence distribution. By minimizing this loss function, the model gradually learns how to make accurate word predictions based on context, thereby improving its understanding and response to instructions.

[0067] Step S106: For each pre-trained model, calculate the gradient of each set of computation nodes based on the gradient equalization mechanism and at least two prediction sequences.

[0068] Optionally, for each pre-trained model deployed on each set of computing nodes, the terminal calculates the gradient of each set of computing nodes based on a gradient equalization mechanism and at least two prediction sequences.

[0069] Step S108: Synchronize the gradients of each group of computing nodes, and use the synchronized gradients to update the parameters of the pre-trained model until the pre-trained model meets the termination condition, thus obtaining the large control model for power switching equipment.

[0070] Optionally, the terminal acquires the gradient of each set of computing nodes and synchronizes each gradient. It then uses the synchronized gradient to update the parameters of the pre-trained model until the pre-trained model meets the termination condition, thus obtaining a large control model for power switching equipment. The termination condition can be that the prediction result output by the updated pre-trained model meets the accuracy condition or that the number of update iterations meets the maximum number of iterations.

[0071] Multiple computing nodes perform the aforementioned forward and backward computations on their respective data shards to obtain local gradients. Subsequently, the gradients from each node are aggregated and averaged through the AllReduce communication mechanism, and the synchronized global gradient is used to update the model parameters.

[0072] In this process, the data parallel layer ensures computational independence and global synchronization between different samples, the model parallel layer is responsible for allocating computational tasks in the parameter dimension, and the pipeline parallel layer realizes overlapping execution between batches in the time dimension, thereby significantly improving fine-tuning efficiency and computational throughput while maintaining model consistency.

[0073] In the aforementioned training method for a large-scale power switchgear control model, firstly, a power switchgear control training dataset is acquired. Each data sample group in the training dataset includes an electrical image and a question text. Using each electrical image and corresponding question text from the training dataset, pre-trained models deployed on each set of computing nodes are trained in parallel, resulting in at least two prediction sequences for each pre-trained model. Each prediction sequence corresponds to a prediction task. By using a specially constructed power switchgear dataset as the training set, a highly accurate dedicated large-scale model can be trained in parallel, improving training efficiency. Secondly, for each pre-trained model, the gradient of each set of computing nodes is calculated based on a gradient balancing mechanism and at least two prediction sequences. This trains a multi-task model capable of adapting to the complex tasks of power equipment operation and maintenance. Finally, the gradients of each set of computing nodes are synchronized, and the parameters of the pre-trained model are updated using the synchronized gradients until the pre-trained model meets the termination condition, resulting in a large-scale power switchgear control model. This method can efficiently and accurately train a batch of dedicated large-scale models to handle complex tasks in power equipment operation and maintenance.

[0074] In one exemplary embodiment, such as Figure 2As shown, obtaining the power switchgear control training dataset includes the following steps S202 to S208. Wherein:

[0075] Step S202: Obtain the electrical signals, environmental parameters, and internal state diagram of the power switchgear; convert the electrical signals into waveform diagrams; and convert the environmental parameters into heat maps.

[0076] Optionally, the electrical signal acquisition equipment collects the electrical signals of the power switchgear and sends the electrical signals to the terminal, while the environmental acquisition equipment collects the environmental parameters of the power switchgear and sends the environmental parameters to the terminal. High-precision industrial cameras and infrared imaging technology collect multi-angle real-world images of the switchgear and internal state diagrams such as contact status.

[0077] The terminal acquires electrical signals from power switchgear, such as three-phase current, three-phase voltage, power factor, harmonic content, and sampling time, and converts these signals into visual images, such as waveforms. The terminal also acquires environmental parameters, such as temperature and humidity, and converts these parameters into heat maps.

[0078] Step S204: Based on the waveform diagram superimposed with the heat map, determine the composite electrical feature image; based on each composite electrical feature image and the corresponding internal state diagram, determine the first image text group.

[0079] Optionally, the terminal overlays the waveform and heat map to obtain a composite electrical feature image and its corresponding text description, such as "Top left: Three-phase current waveform (showing A-phase current 1250A, with slight imbalance)". The terminal matches the corresponding mechanical parameter text label with the internal state diagram corresponding to each composite electrical feature image, and determines the first image text group based on each composite electrical feature image and the mechanical parameter text label.

[0080] Step S206: The first image text group is labeled using electrical rules and a visual language model to determine the second image text group.

[0081] Optionally, the terminal uses a visual language model such as GPT-4o to annotate the first image text group and generate an initial description, such as "Input image: The temperature of the C-phase contact of the circuit breaker is too high; Model output: The C-phase of the circuit breaker is overheating, and it is recommended to check the contact resistance. The current temperature is 82°C, slightly higher than the normal range." The terminal corrects the initial description through pre-configured electrical rules, such as: if a specific threshold is missing, "slightly higher than the normal range" is changed to "exceeds the warning threshold of 80°C"; if a standard basis is missing, "According to DL / T 596 standard, the operating temperature should not exceed 80°C" is added; if a safety constraint is missing, "Operation under load is prohibited, and switching operation tickets must be followed," to obtain the second image text group, such as the input image showing the high temperature of the C-phase contact of the circuit breaker and the text description, "The C-phase contact temperature of the circuit breaker is 82°C, exceeding the 80°C warning threshold specified in DL / T 596 standard, which may indicate an increase in contact resistance. Note: If operation is required, it must be performed according to the switching operation ticket, and opening or closing under load is prohibited." It should be noted that at least one question text can be extracted from the text description.

[0082] Step S208: Enhance the second image text group to obtain a data sample group.

[0083] The data sample group includes electrical images and problem text.

[0084] Optionally, the terminal extracts at least one question text from the text description in the second image text group, and the terminal enhances the image and the question text in the second image text group respectively to obtain a data sample group.

[0085] In this embodiment, by overlaying image-text groups and enhancing them, a high-quality training dataset can be constructed, providing a highly reliable foundation for subsequent model training.

[0086] In one exemplary embodiment, such as Figure 3 As shown, the first image text group is labeled using electrical rules and a visual language model to determine the second image text group, including steps S302 to S306. Wherein:

[0087] Step S302: Analyze the images in the first image text group using a visual language model to obtain descriptive text.

[0088] Optionally, the terminal inputs the first image text group into a visual language model, and uses a visual language model such as GPT-4o to annotate the images in the first image text group to generate descriptive text, which is the initial description mentioned above. The terminal then receives the descriptive text.

[0089] Step S304: Modify the description text based on at least one of the physical rules, safety constraints, and action timing interlocking logic in the electrical rules, and determine the modified text.

[0090] Optionally, the terminal submits the description text to an expert verification interface or automatically compares it with a pre-set electrical rule knowledge base. Power experts or the rule engine verify and correct the text based on at least one of the following rule categories: Physical rule compliance: Verifying whether parameter changes in the description conform to physical laws such as circuit principles and energy conservation; for example, checking whether the "sudden temperature rise" matches the current change trend. Safety limit constraints: Verifying whether the voltage, current, temperature, and other values ​​involved in the description exceed the warning or danger thresholds specified in the equipment safety operation procedures (such as the DL / T standard), and clearly marking and classifying any exceeding limits. Action sequence interlocking logic: Verifying whether the operation command to be generated violates the "five-prevention" logic or operation ticket sequence of the power system; for example, prohibiting the generation of a "direct tripping" command before confirming "load has been transferred." The corrected text ensures that all technical descriptions are accurate and implicitly include safe operation requirements.

[0091] Step S306: Determine the second image text group based on the corrected text and the images in the first image text group.

[0092] Optionally, the terminal uses the corrected text as the core accurate description of the image. Based on this core description, multiple task-oriented "question-answer" pairs are automatically generated using preset templates or natural language generation technology and bound to the original image. For example, for an image displaying "C-phase contact temperature 82°C" and its corrected text, the following can be generated: a status description task pair (question: "Describe the equipment status", answer: "C-phase contact temperature 82°C, exceeding the warning threshold..."), a fault diagnosis task pair (question: "Is there a fault?", answer: "There is a temperature over-limit alarm..."), and an operation suggestion task pair (question: "How should it be handled?", answer: "Strengthen monitoring; if the temperature continues to rise, apply for load reduction..."). Thus, a high-quality "image-question-answer" triplet sample is constructed, and numerous such samples constitute the second image-text set used for training large models.

[0093] In this embodiment, an expert-model collaborative annotation mechanism is introduced. First, a large visual language model is used to automate the initial annotation process, significantly improving efficiency. Then, key corrections and security reinforcement are performed using the knowledge of power industry experts or a rigid rule base, ensuring the professional accuracy and reliability of the annotation results. This method effectively solves the dual challenges of low efficiency and difficulty in ensuring security when constructing high-quality vertical domain training data, laying a solid data foundation for the subsequent training of a reliable and secure large-scale power industry-specific model.

[0094] In one exemplary embodiment, such as Figure 4 As shown, the second image text group is enhanced to obtain a data sample group, including steps S402 to S406. Wherein:

[0095] Step S402: Perform geometric transformation, noise simulation, or illumination perturbation enhancement on the electrical image in the second image text group to obtain the enhanced electrical image.

[0096] Optionally, the terminal invokes the image processing engine to apply various preset image transformations to the original high-quality images (such as composite electrical feature images and actual equipment photos) in the second image text group that have been verified by experts. Specifically, these include: (1) Geometric transformation: randomly rotating the image by ±5 degrees, or flipping it horizontally or vertically to simulate different installation perspectives and shooting angles of the equipment; (2) Noise simulation: superimposing Gaussian noise or salt-and-pepper noise that conforms to the sensor characteristics onto the image, with noise levels (such as...) =0.01) Refer to the noise models of real industrial cameras and infrared thermal imagers to improve the robustness of the model to noise in actual acquired images; (3) Illumination disturbance: Adjust the brightness and contrast of the image (e.g., ±10%) to simulate the changes in ambient illumination under different time periods and weather conditions in the substation. Each original image can generate multiple enhanced image variants through one or more of the above transformations.

[0097] Step S404: Enhance the text in the second image text group based on the power knowledge graph and reference electrical parameters to obtain the enhanced text.

[0098] Optionally, the terminal synchronously enhances the text descriptions corresponding to the electrical images in the second image text group to ensure the consistency of the semantics of the enhanced images and texts. The enhancement methods include: (1) Parameter value floating: The specific electrical parameters described in the text (such as temperature, current, and voltage values) are randomly floated and adjusted within the allowable error range of ±5%. For example, "C phase contact temperature 82°C" in the original text can be enhanced to "C phase contact temperature 78°C" or "C phase contact temperature 86°C". At the same time, other related descriptions in the text (such as "exceeding the limit by 2°C") also need to be updated synchronously to maintain logical consistency. (2) Semantic expansion guided by knowledge graph: Access the pre-built power equipment knowledge graph, and based on the entity relationships in the graph (such as "circuit breaker - has components - contacts" and "contacts - possible faults - overheating"), automatically generate different descriptive sentences or supplementary background knowledge for the same image phenomenon. For example, in response to the phenomenon of "contact overheating", the knowledge graph can generate supplementary and explanatory text such as "It is recommended to check the contact resistance" or "It may be related to recent load fluctuations", thereby increasing the diversity and knowledge density of the text description.

[0099] Step S406: Based on the enhanced electrical image and the enhanced text, determine the data sample group.

[0100] Optionally, the terminal automatically pairs and combines the enhanced electrical images with their corresponding enhanced text. The terminal uses verification logic to ensure that the paired images and texts maintain consistency in the physical states they describe (e.g., an overheated image with adjusted brightness must be paired with text describing fluctuating temperature values). Ultimately, each original "image-text" pair can be expanded to generate dozens or even hundreds of new, high-quality, and diverse "enhanced image-enhanced text" pairs through this enhancement process. All these newly generated pairings collectively constitute a significantly expanded set of data samples for model fine-tuning.

[0101] In this embodiment, by implementing an image-text joint enhancement strategy, the semantic consistency and physical rationality of the enhanced multimodal data pairs are strictly guaranteed while increasing the scale and diversity of data. This not only effectively simulates the uncertainties of data acquisition in real power grid environments (such as perspective, noise, and lighting variations), but also injects domain knowledge through knowledge graphs. This enhances the generalization ability and cognitive depth of future models when facing complex and variable real-world scenarios at the data level, providing an efficient and reliable technical approach to solving the problem of scarce high-quality training data in the power vertical industry.

[0102] In one exemplary embodiment, such as Figure 5 As shown, the pre-trained model deployed on each set of computing nodes is trained in parallel using each electrical image and corresponding question text in the power switchgear control training dataset, including steps S502 to S506. Wherein:

[0103] Step S502: Divide the power switchgear control training dataset into power switchgear control training sub-datasets.

[0104] In this case, the number of training subsets for power switchgear control is the same as the number of sets of computing nodes, such as... .

[0105] Optionally, the terminal divides the power switchgear control training dataset D into groups according to the number of computing nodes using a data parallel approach. A training dataset for power switchgear control Each compute node loads a unique data shard. This is also known as the training dataset for power switchgear control.

[0106] Each compute device (GPU) within a compute node holds its own unique data shard. A copy of the model and a complete copy of the pre-trained model are used, and independent local training is performed based on this data.

[0107] Step S404: In each group of computing nodes, the pre-trained model is divided into the number of model stages according to the network layer order through pipeline parallelism; the power switch control training subset is divided into multiple power switch control training micro subsets according to the number of model stages.

[0108] Optionally, to address the issue of excessively large model parameters that a single computing node cannot accommodate, the terminal employs a pipelined parallelism strategy within each computing node. This includes dividing the L-layer network of the pre-trained model into P consecutive stages, either evenly or based on computational complexity, along the sequential order of the layers. Here, P represents the pipeline parallelism, equal to the number of computing nodes participating in the pipeline (typically P is less than or equal to the number of nodes in a data parallel group). Each computing device (GPU) is assigned a model stage, responsible for storing and executing the network layers contained within that stage. Simultaneously, to fill the pipeline and reduce device idle time (bubbles), a power switch control training subset is further divided into M power switch control training micro-subsets during training. These power switch control training micro-subsets serve as the basic units of the data flow, sequentially and overlappingly flowing through the P model stages.

[0109] Step S406: Based on the number of computing node groups and the number of model stages, determine the number of parallel tensors for each model stage; according to the number of parallel tensors, use each electrical image and corresponding question text in the power switch equipment control training sub-dataset to train the network layer of the pre-trained model, obtain the calculation results, and pass the calculation results in the order of the network layers.

[0110] Optionally, within each pipeline stage's deployed compute nodes, the terminal further employs a tensor parallel strategy to partition the computation of individual network layers. The number of parallel tensors (i.e., tensor parallelism T) is typically determined based on the number of GPUs within that compute node. For each network layer (e.g., a feedforward neural network layer) within that stage, its weight matrix is ​​partitioned into T parts along the row or column dimensions, with each GPU storing and computing one part. When processing a power switch control training sub-dataset, all samples within this sub-dataset (complete sequences of electrical images and corresponding question text encoded into text) are broadcast to all T GPUs within the compute node. Each GPU uses its local partial weights to compute, obtaining a partial output. Then, an All-Gather communication operation is performed via high-speed interconnects between GPUs (e.g., NVLink) to concatenate the partial outputs of all GPUs into the complete output of that layer. This complete result, after passing through an activation function, serves as the output of that stage. If the current stage is not the last stage, this output (i.e., the "activation value") is passed through an inter-node network (e.g., InfiniBand) to the compute node of the next pipeline stage as its input. During backpropagation, gradient calculation and propagation proceed in the opposite direction, and gradient aggregation and synchronization are achieved through communication operations such as Reduce-Scatter.

[0111] In this embodiment, by organically combining the three parallel strategies of data parallelism, pipeline parallelism and tensor parallelism into a three-dimensional hybrid parallel architecture, the decomposition and acceleration of the training task are realized simultaneously in the three dimensions of data, model layer and time.

[0112] In one exemplary embodiment, such as Figure 6 As shown, each data sample group also includes the answer text corresponding to the question text; based on the gradient equalization mechanism and at least two prediction sequences, the gradient of each set of computation nodes is calculated, including steps S602 to S610. Wherein:

[0113] Step S602: Calculate the task loss between each predicted sequence and the answer text.

[0114] Optionally, the terminal calculates the cross-entropy loss for the same input sample based on its task type (such as status description, fault diagnosis, or operation instruction generation). For example, for the answer text "C-phase contact temperature 82°C, exceeding the warning threshold, it is recommended to strengthen monitoring" in the sample, the model will generate a corresponding prediction sequence (i.e., the probability distribution of each word predicted by the model). The terminal calculates the following losses separately: a first task loss for evaluating the accuracy of the status description (comparing the difference between the model-generated description and "C-phase contact temperature 82°C"), a second task loss for evaluating the accuracy of the fault judgment (comparing the difference with "exceeding the warning threshold"), and a third task loss for evaluating the rationality of the recommendation (comparing the difference with "it is recommended to strengthen monitoring").

[0115] Step S404: Determine the total loss based on the losses of each task, security constraints, and adversarial actions.

[0116] Among them, the security constraint loss represents the difference between the features of the data sample group and the predefined security feature library; the adversarial training loss represents the difference between the answer text output based on the data sample group and the answer text output based on the second image text group, that is, the difference between the answer text output by the image text group before and after enhancement.

[0117] Optionally, the terminal uses contrastive learning to calculate the safety constraint loss, which includes: the terminal extracting intermediate feature representations generated by the model when processing the current input (image and question), calculating the similarity between these features and positive samples (such as known safe operating mode features) in a predefined safety feature library, and the dissimilarity between these features and negative samples (such as known dangerous operating mode features). The safety constraint loss is used to reduce the distance between model features and safe modes, and increase the distance between model features and dangerous modes, thereby internalizing safety knowledge at the feature level.

[0118] Optionally, the terminal applies a tiny, imperceptible random perturbation to the current input electrical image, generating an adversarial example. Then, the original input and the adversarial example are input into the model respectively, yielding two output probability distributions. The adversarial loss is measured by calculating the difference between these two output distributions (e.g., KL divergence), with the aim of making the model insensitive to such harmless perturbations and avoiding the generation of logically conflicting or abrupt instructions due to minute input changes.

[0119] Optionally, the terminal performs a weighted summation of the losses from each task, security constraints, and adversarial actions to determine the total loss.

[0120] Step S406: Calculate the gradient norm of each task loss, security constraint loss, and adversarial loss, and adjust the weights of task loss, security constraint loss, and adversarial loss in the total loss according to each gradient norm.

[0121] Optionally, during backpropagation, the terminal calculates the gradient vector of each task loss (L_task1, L_task2, L_task3, L_safety, L_adv) with respect to the model parameters and calculates its norm, such as the L2 norm.

[0122] The terminal uses a gradient balancing mechanism to adjust the weights of each task's loss, the safety constraint loss, and the adversarial loss in the total loss based on their respective gradient norms. The gradient balancing mechanism dynamically adjusts these weights to prevent tasks with excessively large gradient norms from dominating the training process, ensuring multi-objective collaborative optimization.

[0123] The larger the gradient norm of a loss term, the lower the weight of the task corresponding to that loss term will be in the current training step. For example, the new weight wi' = (initial weight wi) / (gradient norm gi + ... ),in It is a very small constant to prevent division by zero, and then all wi' are normalized so that their sum is 1.

[0124] Step S608: Determine the weighted total loss based on the losses of each task, security constraints, adversarial actions, and their corresponding weights.

[0125] Optionally, the terminal determines the weighted total loss based on the loss of each task and its corresponding adjusted weight, the loss of security constraints and its corresponding adjusted weight, and the loss of adversarial actions and its corresponding adjusted weight.

[0126] Step S610: Calculate the gradient of each set of computation nodes based on the weighted total loss.

[0127] Optionally, the terminal performs backpropagation based on the weighted total loss to calculate the global gradient of all model parameters. In a distributed environment, this gradient calculation process is performed independently on each computing node, and then synchronized across groups through communication operations such as All-Reduce to ensure that all model replicas are updated with consistent gradients.

[0128] In this embodiment, by introducing a multi-task joint optimization and dynamic gradient equalization mechanism, and integrating safety constraint loss and adversarial training loss, a unified improvement in the accuracy, safety and robustness of a single model in the power switch control scenario is achieved.

[0129] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0130] Based on the same inventive concept, this application also provides a power switchgear control large model training device for implementing the above-mentioned power switchgear control large model training method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more power switchgear control large model training device embodiments provided below can be found in the limitations of the power switchgear control large model training method above, and will not be repeated here.

[0131] In one exemplary embodiment, such as Figure 7 As shown, a large-scale model training device for power switchgear control is provided, comprising: an acquisition module 701, a training module 702, a gradient calculation module 703, and an update module 704, wherein:

[0132] The acquisition module 701 is used to acquire the power switchgear control training dataset; wherein each data sample group in the power switchgear control training dataset includes an electrical image and a question text.

[0133] Training module 702 is used to train the pre-trained model deployed on each set of computing nodes in parallel using each electrical image and corresponding question text in the power switch equipment control training dataset, to obtain at least two prediction sequences for each pre-trained model; wherein each prediction sequence corresponds to a prediction task.

[0134] The gradient calculation module 703 is used to calculate the gradient of each set of computing nodes for each pre-trained model, based on the gradient equalization mechanism and at least two prediction sequences.

[0135] The update module 704 is used to synchronize the gradients of each group of computing nodes and use the synchronized gradients to update the parameters of the pre-trained model until the pre-trained model meets the termination condition, thus obtaining the large control model for power switchgear.

[0136] In an exemplary embodiment, the acquisition module 701 is further configured to acquire electrical signals, environmental parameters, and internal state diagrams of the power switchgear; convert the electrical signals into waveform diagrams and the environmental parameters into heat maps; determine composite electrical feature images based on the waveform diagrams overlaid with heat maps; determine a first image-text group based on each composite electrical feature image and its corresponding internal state diagram; annotate the first image-text group using electrical rules and a visual language model to determine a second image-text group; and enhance the second image-text group to obtain a data sample group; wherein the data sample group includes electrical images and question text.

[0137] In an exemplary embodiment, the acquisition module 701 is further configured to analyze the images in the first image text group using a visual language model to obtain descriptive text; modify the descriptive text based on at least one of physical rules, safety constraint limits, and action timing interlocking logic in electrical rules to determine the modified text; and determine the second image text group based on the modified text and the images in the first image text group.

[0138] In an exemplary embodiment, the acquisition module 701 is further configured to perform geometric transformation, noise simulation, or illumination disturbance enhancement on the electrical images in the second image-text group to obtain enhanced electrical images; enhance the text in the second image-text group based on the power knowledge graph and reference electrical parameters to obtain enhanced text; and determine a data sample group based on the enhanced electrical images and enhanced text.

[0139] In an exemplary embodiment, the training module 702 is further configured to divide the power switchgear control training dataset into power switchgear control training subsets; wherein the number of power switchgear control training subsets is the same as the number of computing node groups; in each group of computing nodes, the pre-trained model is divided into a number of model stages according to the network layer order through pipeline parallelism; the power switchgear control training subset is divided into multiple power switchgear control training micro-sub-datasets according to the number of model stages; the number of parallel tensors for each model stage is determined based on the number of computing node groups and the number of model stages; the network layers of the pre-trained model are trained using each electrical image and corresponding question text in the power switchgear control training micro-sub-dataset according to the number of parallel tensors, the calculation results are obtained, and the calculation results are passed in the order of network layers.

[0140] In an exemplary embodiment, each data sample group further includes the answer text corresponding to the question text; the gradient calculation module 703 is further used to calculate the task loss between each predicted sequence and the answer text; determine the total loss based on each task loss, security constraint loss, and adversarial loss; wherein, the security constraint loss represents the difference between the features of the data sample group and the predefined security feature library; the adversarial training loss represents the difference between the answer text output according to the data sample group and the answer text output according to the second image text group; calculate the gradient norm of each task loss, security constraint loss, and adversarial loss respectively, and adjust the weights of the task loss, security constraint loss, and adversarial loss in the total loss according to each gradient norm; determine the weighted total loss based on each task loss, security constraint loss, adversarial loss, and the corresponding weights; calculate the gradient of each set of computing nodes according to the weighted total loss.

[0141] Each module in the aforementioned large-scale power switchgear control training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0142] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to a large-scale control model of power switching equipment. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a training method for a large-scale control model of power switching equipment.

[0143] Those skilled in the art will understand that Figure 8The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0144] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0145] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0146] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0147] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0148] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0149] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A power switchgear control large model training method, characterized by, The method comprises: acquiring a power switch device control training data set; wherein each data sample group in the power switch device control training data set comprises an electrical image and a question text; training a pre-trained model deployed on each group of computing nodes in parallel using each electrical image and the corresponding question text in the power switch device control training data set, obtaining at least two prediction sequences of each pre-trained model; wherein each prediction sequence corresponds to a prediction task; for each pre-trained model, calculating the gradient of each group of computing nodes based on a gradient balancing mechanism and the at least two prediction sequences; synchronize the gradients of each group of computing nodes, and update the parameters of the pre-trained model using the synchronized gradients until the pre-trained model meets the termination condition, obtaining a power switch device control large model.

2. The method of claim 1, wherein, The acquisition of the power switch device control training data set comprises: acquiring electrical signals, environmental parameters and internal state graphs of the power switch device, converting the electrical signals into waveform graphs, and converting the environmental parameters into thermal maps; determining a composite electrical feature image based on the superposition of the waveform graph and the thermal map; determining a first image text group based on each composite electrical feature image and the corresponding internal state graph; annotating the first image text group using electrical rules and visual language models to determine a second image text group; enhancing the second image text group to obtain a data sample group; wherein the data sample group comprises an electrical image and a question text.

3. The method of claim 2, wherein, The annotation of the first image text group using electrical rules and visual language models to determine a second image text group comprises: analyzing the images in the first image text group using a visual language model to obtain description text; correcting the description text based on at least one of the physical rules, safety limit constraints and action timing interlocking logic in the electrical rules to determine corrected text; determining the second image text group based on the corrected text and the images in the first image text group.

4. The method of claim 2, wherein, The enhancement of the second image text group to obtain a data sample group comprises: performing geometric transformation, noise simulation or illumination disturbance enhancement on the electrical images in the second image text group to obtain enhanced electrical images; enhancing the text in the second image text group based on a power knowledge graph and reference electrical parameters to obtain enhanced text; determining a data sample group based on the enhanced electrical images and the enhanced text.

5. The method of claim 1, wherein, The training of a pre-trained model deployed on each group of computing nodes in parallel using each electrical image and the corresponding question text in the power switch device control training data set comprises: dividing the power switch device control training data set into power switch device control training sub-data sets; wherein the number of power switch device control training sub-data sets is the same as the number of groups of computing nodes; dividing the pre-training model into a number of model stages in sequence of network layers by pipeline parallelism in each group of computing nodes; and dividing the power switch device control training sub-data set into a plurality of power switch device control training micro sub-data sets according to the number of model stages; determining the number of parallel tensors for each model stage based on the number of groups of computing nodes and the number of model stages; and training the network layer of the pre-training model using each electrical image and corresponding question text in the power switch device control training micro sub-data set according to the number of parallel tensors, obtaining a calculation result, and passing the calculation result in sequence of network layers.

6. The method of claim 1, wherein, Each of the data sample groups further includes an answer text corresponding to the question text. The gradient of each group of computing nodes is calculated based on the gradient balancing mechanism and the at least two prediction sequences, including: calculating a task loss between each of the prediction sequences and the answer text; determining a total loss based on each of the task losses, a safety constraint loss, and an adversarial loss; wherein the safety constraint loss represents a difference between a feature of the data sample group and a pre-defined safety feature library; and the adversarial training loss represents a difference between an answer text output according to the data sample group and an answer text output according to a second image text group; respectively calculating a gradient norm of each of the task losses, the safety constraint loss, and the adversarial loss, and adjusting a weight of each of the task loss, the safety constraint loss, and the adversarial loss in the total loss according to the corresponding gradient norm; determining a weighted total loss based on each of the task losses, the safety constraint loss, the adversarial loss, and the corresponding weight; calculating the gradient of each group of computing nodes according to the weighted total loss.

7. A power switching device control large model training apparatus characterized by comprising: The apparatus includes: an acquisition module configured to acquire a power switch device control training data set; wherein each data sample group in the power switch device control training data set includes an electrical image and a question text; a training module configured to use each electrical image and corresponding question text in the power switch device control training data set to train a pre-training model deployed on each group of computing nodes in parallel, to obtain at least two prediction sequences of each pre-training model; wherein each of the prediction sequences corresponds to a prediction task; a gradient calculation module configured to, for each pre-training model, calculate the gradient of each group of computing nodes based on a gradient balancing mechanism and the at least two prediction sequences; an update module configured to synchronize the gradients of each group of computing nodes, and update parameters of the pre-training model using the synchronized gradients until the pre-training model meets a termination condition, to obtain a power switch device control large model.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor implements the steps of the method of any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 6.