Multi-mode segmentation method, device and equipment for high-voltage switch design drawing and medium

By fusing multi-scale texture and geometric features of high-voltage switch design drawings through a cross-modal attention mechanism, the problems of visual fatigue and insufficient integration of multi-modal information in traditional manual review are solved, and high-precision target edge segmentation is achieved.

CN120997510APending Publication Date: 2025-11-21中国电气装备集团科学技术研究院有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511129087.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional manual review of high-voltage switch design drawings suffers from visual fatigue and insufficient multimodal information integration capabilities, making it difficult to meet the requirements of millimeter-level tolerances and dynamic specification adaptation, resulting in insufficient drawing segmentation precision and accuracy.

Method used

A multimodal segmentation model based on a cross-modal attention mechanism is adopted, which integrates multi-scale texture features and geometric features from high-voltage switch design drawings. The cross-modal attention mechanism is used to perform feature fusion on CAD images and primitive information to achieve high-precision target edge segmentation.

Benefits of technology

It improves the precision and accuracy of target edge segmentation in high-voltage switch design drawings, reduces single-mode segmentation error, and enhances the segmentation robustness in complex occlusion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997510A_ABST
    Figure CN120997510A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode segmentation method and device for a high-voltage switch design drawing, equipment and a medium. The method comprises the steps that a multi-mode segmentation event responding to a high-voltage switch design drawing is triggered, and a CAD file of a high-voltage switch of power transmission and transformation equipment is acquired; determining CAD primitive information related to the CAD file and a CAD image corresponding to the CAD file; inputting the CAD primitive information and the CAD image into a pre-trained multi-modal segmentation model to obtain a target segmentation image corresponding to the CAD image; wherein the multi-modal segmentation model is a segmentation model for fusing multi-scale texture features in the CAD image and geometric features of CAD primitive information based on a cross-modal attention mechanism. According to the scheme, global-local feature complementation is realized by fusing the global texture information of the drawing image and the standardized geometric data of the CAD primitive, so that the precision of target edge segmentation in the drawing of the high-voltage switch of the power transmission and transformation equipment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power equipment technology, and in particular to a multimodal segmentation method, device, equipment and medium for high-voltage switch design drawings. Background Technology

[0002] The review of high-voltage switch design drawings is a core link in ensuring the safe operation of power systems, but the traditional manual review method has significant shortcomings. Traditional manual drawing review is limited by visual fatigue and insufficient ability to integrate multimodal information, making it difficult to meet the stringent requirements of millimeter-level tolerances and dynamic specification adaptation. In addition, drawings contain multimodal information such as images (surface textures), vector data (geometric coordinates), and text annotations (specification parameters). Single-modal segmentation can easily lead to semantic fragmentation. Therefore, the field of drawing segmentation faces the dual challenges of high boundary accuracy and multimodal accuracy. Summary of the Invention

[0003] This invention provides a multi-modal segmentation method, device, equipment, and medium for high-voltage switch design drawings, which can effectively improve the accuracy of target edge segmentation in high-voltage switch drawings for power transmission and transformation equipment.

[0004] According to one aspect of the present invention, a multi-modal segmentation method for high-voltage switch design drawings is provided, comprising:

[0005] In response to the multimodal segmentation event triggered by the high-voltage switch design drawings, the CAD file of the high-voltage switch of the power transmission and transformation equipment is obtained;

[0006] Determine the CAD element information involved in the CAD file and the CAD image corresponding to the CAD file;

[0007] The CAD primitive information and the CAD image are input into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses multi-scale texture features in the CAD image and geometric features of the CAD primitive information based on a cross-modal attention mechanism.

[0008] According to another aspect of the present invention, a multi-mode segmentation device for high-voltage switch design drawings is provided, comprising:

[0009] The CAD file acquisition module is used to acquire the CAD file of the high-voltage switch of the power transmission and transformation equipment in response to the multimodal segmentation event of the high-voltage switch design drawings.

[0010] The graphic element information and image acquisition module is used to determine the CAD graphic element information involved in the CAD file and the CAD image corresponding to the CAD file;

[0011] The target multimodal segmentation module is used to input the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses the multi-scale texture features in the CAD image and the geometric features of the CAD primitive information based on a cross-modal attention mechanism.

[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0013] At least one processor; and

[0014] A memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the multimodal segmentation method for high-voltage switch design drawings according to any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the multi-modal segmentation method for high-voltage switch design drawings according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the multimodal segmentation method for high-voltage switch design drawings as described in any embodiment of the present invention.

[0018] The multimodal segmentation scheme for high-voltage switch design drawings of this invention, in response to the triggering of a multimodal segmentation event for high-voltage switch design drawings, acquires the CAD file of the high-voltage switch of power transmission and transformation equipment; determines the CAD primitive information involved in the CAD file and the corresponding CAD image; inputs the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses multi-scale texture features in the CAD image and geometric features of the CAD primitive information based on a cross-modal attention mechanism. The technical solution provided by this invention, by fusing global texture information of the drawing image with standardized geometric data of CAD primitives, achieves global-local feature complementarity, thereby effectively improving the accuracy of target edge segmentation in the drawings of high-voltage switches of power transmission and transformation equipment.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a multi-modal segmentation method for high-voltage switch design drawings provided in Embodiment 1 of the present invention;

[0022] Figure 2 This is an architecture diagram of a multimodal segmentation model provided in an embodiment of the present invention;

[0023] Figure 3 This is a structural schematic diagram of a multi-mode segmentation device in a high-voltage switch design drawing provided in Embodiment 2 of the present invention;

[0024] Figure 4 A schematic diagram of the structure of an electronic device for implementing the multi-modal segmentation method of the high-voltage switch design drawings in this embodiment of the invention. Detailed Implementation

[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] Example 1

[0028] Figure 1 This is a flowchart of a multimodal segmentation method for high-voltage switch design drawings provided in Embodiment 1 of the present invention. This embodiment is applicable to the segmentation of targets in high-voltage switch design drawings. The method can be executed by a multimodal segmentation device for high-voltage switch design drawings. This multimodal segmentation device can be implemented in hardware and / or software and can be configured in electronic equipment. Figure 1 As shown, the method includes:

[0029] S110, In response to the multimodal segmentation event triggered by the high-voltage switch design drawings, obtain the CAD file of the high-voltage switch of the power transmission and transformation equipment.

[0030] In this embodiment of the invention, when a multimodal segmentation command for a high-voltage switch design drawing is received from a user, a multimodal segmentation event for the high-voltage switch design drawing is determined to be triggered. In response to the triggering of the multimodal segmentation event, the CAD file of the high-voltage switch of the power transmission and transformation equipment is obtained. The CAD file is a .dwg file containing the design data of the high-voltage switch of the power transmission and transformation equipment. For example, the CAD file storage path is obtained, and the CAD file of the high-voltage switch of the power transmission and transformation equipment is read based on the CAD file storage path; alternatively, the CAD file of the high-voltage switch of the power transmission and transformation equipment transmitted by other terminal devices can also be directly obtained. It should be noted that this embodiment of the invention does not limit the method of obtaining the CAD file of the high-voltage switch of the power transmission and transformation equipment.

[0031] S120. Determine the CAD element information involved in the CAD file and the CAD image corresponding to the CAD file.

[0032] In this embodiment of the invention, since the high-voltage switch of power transmission and transformation equipment may contain multiple components, when drawing CAD drawings of the high-voltage switch of power transmission and transformation equipment, multiple components are involved, and each component can be represented by multiple graphic elements. Therefore, the CAD file is analyzed to determine the CAD graphic element information involved in the CAD file. The CAD graphic element information may include graphic element location information and target information to which the graphic element belongs. For example, the geometric data of the CAD file is parsed to determine the CAD graphic element information, and based on the graphic element information of the CAD file, all graphic element data within the same target are saved with a unified fixed name, for example, it can be labeled and saved in the mode of target type + number. A fixed conversion ratio is set for the parsed CAD file, and a high-precision quality CAD image is exported according to the conversion ratio.

[0033] S130. Input the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses the multi-scale texture features in the CAD image and the geometric features of the CAD primitive information based on a cross-modal attention mechanism.

[0034] In this embodiment of the invention, CAD primitive information and CAD images are input into a pre-trained multimodal segmentation model. The multimodal segmentation model fuses multi-scale texture features and geometric features of CAD primitive information in the CAD image based on a cross-modal attention mechanism, and performs target segmentation on the CAD image based on the fused features, thereby obtaining the target segmentation image corresponding to the CAD image output by the multimodal segmentation model. The segmented target in the CAD image can be one or more complete components, or it can be a local region of a component.

[0035] For example, the multimodal segmentation model includes a dual-branch encoder built on a multimodal data fusion architecture. The dual-branch encoder comprises an image branch encoder and a primitive branch encoder, where the image branch encoder is an encoder built on the mask2transform segmentation model framework. After inputting CAD primitive information and CAD images into the multimodal segmentation model, the image branch encoder processes the CAD image through a multi-scale feature extraction layer, a Transformer decoder, and a mask attention mechanism to extract multi-scale texture features from the CAD image. The CAD branch encoder extracts geometric features from the CAD primitive information. Then, based on a hierarchical coding strategy, the multi-scale texture features and geometric features are mapped to features of the same width. The texture and geometric information are dynamically fused through a cross-modal attention mechanism. Finally, the CAD image is segmented based on the fused features.

[0036] Optionally, the multimodal segmentation model includes a cue encoder, a ViT model, and a pixel decoder. Inputting the CAD primitive information and the CAD image into the pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image includes: inputting the CAD primitive information into the cue encoder to obtain the geometric feature encoding vector output by the cue encoder; inputting the CAD image into the ViT model to obtain the multi-scale texture feature encoding vector output by the ViT model; fusing the geometric feature encoding vector and the multi-scale texture feature encoding vector based on a cross-modal attention mechanism to generate a fused feature encoding vector; decoding the fused feature encoding vector based on the pixel decoder to generate a segmentation mask, and segmenting the target in the CAD image based on the segmentation mask to generate the target segmented image.

[0037] For example, Figure 2 This is an architecture diagram of a multimodal segmentation model provided in an embodiment of the present invention. Figure 2 As shown, since CAD primitive information includes primitive location information and target information to which the primitive belongs, the CAD primitive information can be input into the prompt encoder (i.e., the Prompt Encoder) in the form of point groups to obtain the geometric feature encoding vector output by the Prompt Encoder. The CAD image is then input into the ViT model to obtain the multi-scale texture feature encoding vector output by the ViT model. Simultaneously, to adapt the CAD image to targets of different scales (such as components or local parts of components), the CAD image is input into the ViT (Vision Transformer) model to obtain the multi-scale texture feature encoding vector through the ViT model. Optionally, inputting the CAD image into the ViT model to obtain the multi-scale texture feature encoding vector output by the ViT model includes: inputting the CAD image into the ViT model, dividing the CAD image into at least two non-overlapping blocks using the ViT model, and determining the position encoding corresponding to each non-overlapping block; converting each non-overlapping block into a corresponding non-overlapping block feature vector using a linear mapping; and inputting the non-overlapping block feature vector and the corresponding position encoding into a Transformer encoder to obtain the multi-scale texture feature encoding vector output by the Transformer encoder. Figure 2 As shown, the ViT model divides the CAD image into multiple fixed-size non-overlapping blocks and unfolds them into vectors. Simultaneously, it learns the position corresponding to each non-overlapping block, generating a positional code. Then, it maps each non-overlapping block to a non-overlapping block feature vector through linear projection. The non-overlapping block feature vector and its corresponding positional code are input into a multi-layer standard Transformer encoder, which generates a multi-scale texture feature encoding vector corresponding to the CAD image.

[0038] In this embodiment of the invention, geometric feature encoding vectors and multi-scale texture feature encoding vectors are fused based on a cross-modal attention mechanism to generate a fused feature encoding vector. Before fusing the geometric feature encoding vectors and multi-scale texture feature encoding vectors based on the cross-modal attention mechanism, the method further includes: aligning the geometric feature encoding vectors and multi-scale texture feature encoding vectors based on a cross-attention mechanism; fusing the geometric feature encoding vectors and multi-scale texture feature encoding vectors based on the cross-modal attention mechanism to generate the fused feature encoding vector includes: fusing the aligned geometric feature encoding vectors and multi-scale texture feature encoding vectors based on the cross-modal attention mechanism to generate the fused feature encoding vector. Figure 2As shown, the geometric feature encoding vector output by the Prompt Encoder and the multi-scale texture feature encoding vector output by the ViT model are input into the Cross Attention model. Cross Attention aligns the geometric feature encoding vector and the multi-scale texture feature encoding vector based on the cross attention mechanism, thereby aligning primitive information with image features. Then, the aligned geometric feature encoding vector and the multi-scale texture feature encoding vector are fused based on the cross-modal attention mechanism.

[0039] In the multimodal segmentation model, the pixel decoder decodes the fused feature encoding vector to generate a segmentation mask. Based on this mask, targets in the CAD image are segmented to generate a segmented image. Essentially, the pixel decoder uses the mutual constraints between multi-scale texture and geometric features to process image features and generate the segmentation mask. These mutual constraints are defined as follows: the mask can optimize classification based on the KL divergence with primitives, and the classified primitives have complete boundary representations. Cross-entropy loss is used to truncate the boundaries, achieving a boundary correction effect.

[0040] The multimodal segmentation scheme for high-voltage switch design drawings of this invention, in response to the triggering of a multimodal segmentation event for high-voltage switch design drawings, acquires the CAD file of the high-voltage switch of power transmission and transformation equipment; determines the CAD primitive information involved in the CAD file and the corresponding CAD image; inputs the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses multi-scale texture features in the CAD image and geometric features of the CAD primitive information based on a cross-modal attention mechanism. The technical solution provided by this invention, by fusing global texture information of the drawing image with standardized geometric data of CAD primitives, achieves global-local feature complementarity, thereby effectively improving the accuracy of target edge segmentation in the drawings of high-voltage switches of power transmission and transformation equipment.

[0041] In some embodiments, the loss function used by the multimodal segmentation model is a comprehensive loss function composed of a classification loss function, a regression loss function, and a KL divergence loss function. The classification loss function is used to calculate the difference between the predicted category and the true label; the regression loss function is used to calculate the distance between mask feature points and primitive points; and the KL divergence loss function is used to calculate the KL divergence between the pixel-level segmentation result and the CAD geometric boundary. The advantage of this setup is that by combining three loss functions, the final training loss function is obtained, achieving loss function optimization under multiple considerations of category and segmentation accuracy. This helps to further improve the accuracy of the multimodal segmentation model in segmenting target edges in drawings of high-voltage switches for power transmission and transformation equipment.

[0042] like Figure 2 As shown, the loss function L used in the multimodal segmentation model is: L = L cls +L dice +L reg Among them, L cls The classification loss function can be calculated using the cross-entropy loss function to determine the difference between the predicted class and the true label; L reg L is the regression loss function used to calculate the distance between the mask feature points and the primitive point coordinates; dice The KL divergence loss function is used to calculate the KL divergence between the pixel-level segmentation result and the CAD geometric boundary. By adding the above three loss functions, the final loss function used to train the multimodal segmentation model can be obtained, achieving loss function optimization under multiple considerations of category and segmentation accuracy.

[0043] Optionally, the training strategy used when training the multimodal segmentation model is as follows: end-to-end training is achieved using a detector head + classification head + segmentation head mechanism. Adversarial enhancement can also be performed during training, i.e., injecting noise such as CAD layer misalignment and drawing scanning distortion to improve the model's generalization ability to real-world engineering data. After training the multimodal segmentation model, accuracy verification and lightweight encapsulation can be performed. Accuracy verification can be implemented in high-voltage switch contact segmentation tasks by establishing a high-quality data acceptance database to achieve high-quality and high-precision segmentation results. Lightweight encapsulation can include deploying the model based on different frameworks, supporting multimodal input of drawing images / CAD files / text commands, and achieving millisecond-level inference at the edge through ONNX Runtime.

[0044] The multimodal segmentation method for high-voltage switch design drawings provided in this invention fuses global texture information from the drawing image with standardized geometric data from the CAD drawing. It employs a cross-modal attention mechanism to achieve feature complementarity, with the image branch extracting texture features and the geometric branch resolving vectorized coordinate localization. Multi-scale fusion is achieved through a dual-channel feature pyramid network, reducing segmentation errors compared to single-modal methods and significantly improving robustness in complex occluded scenes, ultimately achieving high-precision, high-edge segmentation. Furthermore, by fusing the multi-scale feature pyramid of the image with the vector topological constraints of the CAD geometric data through the cross-modal attention mechanism, a bidirectional interactive channel is constructed: the image → CAD branch uses a spatial attention mechanism to dynamically filter texture features matching the geometric coordinates; the CAD → image branch uses an edge alignment loss function to enhance the consistency between vector boundaries and pixel-level segmentation results. Combined with bidirectional cross-entropy loss, end-to-end optimization is achieved, effectively improving the robustness of the segmentation algorithm.

[0045] Example 2

[0046] Figure 3 This is a structural schematic diagram of a multi-modal segmentation device for a high-voltage switch design drawing provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes:

[0047] The CAD file acquisition module 310 is used to acquire the CAD file of the high-voltage switch of the power transmission and transformation equipment in response to the multimodal segmentation event of the high-voltage switch design drawing being triggered.

[0048] The graphic element information and image acquisition module 320 is used to determine the CAD graphic element information involved in the CAD file and the CAD image corresponding to the CAD file;

[0049] The target multimodal segmentation module 330 is used to input the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses the multi-scale texture features in the CAD image and the geometric features of the CAD primitive information based on a cross-modal attention mechanism.

[0050] Optionally, the multimodal segmentation model includes a cue encoder, a ViT model, and a pixel decoder;

[0051] The target multimodal segmentation module includes:

[0052] A geometric feature encoding vector acquisition unit is used to input the CAD primitive information into the prompt encoder and acquire the geometric feature encoding vector output by the prompt encoder;

[0053] A multi-scale texture feature encoding vector acquisition unit is used to input the CAD image into the ViT model and acquire the multi-scale texture feature encoding vector output by the ViT model;

[0054] The fusion feature encoding vector generation unit is used to fuse the geometric feature encoding vector and the multi-scale texture feature encoding vector based on a cross-modal attention mechanism to generate a fusion feature encoding vector;

[0055] The target segmentation image generation unit is used to decode the fused feature encoding vector based on the pixel decoder, generate a segmentation mask, and segment the target in the CAD image based on the segmentation mask to generate a target segmentation image.

[0056] Optionally, the multi-scale texture feature encoding vector acquisition unit is used for:

[0057] The CAD image is input into the ViT model, and the ViT model divides the CAD image into at least two non-overlapping blocks, and the position code corresponding to each non-overlapping block is determined.

[0058] Each non-overlapping block is converted into a corresponding non-overlapping block feature vector through linear mapping; and the non-overlapping block feature vector and the corresponding position code are input into the Transformer encoder to obtain the multi-scale texture feature encoding vector output by the Transformer encoder.

[0059] Optional, also includes:

[0060] The feature alignment unit is used to align the geometric feature encoding vector and the multi-scale texture feature encoding vector based on a cross-attention mechanism before fusing them based on a cross-modal attention mechanism.

[0061] The fusion feature encoding vector generation unit is used for:

[0062] The aligned geometric feature encoding vector and the multi-scale texture feature encoding vector are fused based on a cross-modal attention mechanism to generate a fused feature encoding vector.

[0063] Optionally, the loss function used in the multimodal segmentation model is a comprehensive loss function composed of a classification loss function, a regression loss function, and a KL divergence loss function; wherein, the classification loss function is used to calculate the difference between the predicted category and the true label; the regression loss function is used to calculate the distance between the mask feature point and the primitive point; and the KL divergence loss function is used to calculate the KL divergence between the pixel-level segmentation result and the CAD geometric boundary.

[0064] Optionally, the CAD element information includes element location information and target information to which the element belongs.

[0065] The multimodal segmentation device for high-voltage switch design drawings provided in this embodiment of the invention can execute the multimodal segmentation method for high-voltage switch design drawings provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0066] Example 3

[0067] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0068] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0069] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0070] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the multimodal segmentation method in high-voltage switch design drawings.

[0071] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0072] In some embodiments, the multimodal segmentation method for high-voltage switchgear design drawings can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the multimodal segmentation method for high-voltage switchgear design drawings described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the multimodal segmentation method for high-voltage switchgear design drawings by any other suitable means (e.g., by means of firmware).

[0073] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0074] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0075] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0077] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0078] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0079] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0080] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A multi-modal segmentation method for high-voltage switch design drawings, characterized in that, include: In response to the multimodal segmentation event triggered by the high-voltage switch design drawings, the CAD file of the high-voltage switch of the power transmission and transformation equipment is obtained; Determine the CAD element information involved in the CAD file and the CAD image corresponding to the CAD file; The CAD primitive information and the CAD image are input into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses multi-scale texture features in the CAD image and geometric features of the CAD primitive information based on a cross-modal attention mechanism.

2. The method according to claim 1, characterized in that, The multimodal segmentation model includes a cue encoder, a ViT model, and a pixel decoder; Inputting the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image includes: The CAD primitive information is input into the prompt encoder to obtain the geometric feature encoding vector output by the prompt encoder; The CAD image is input into the ViT model to obtain the multi-scale texture feature encoding vector output by the ViT model; The geometric feature encoding vector and the multi-scale texture feature encoding vector are fused based on a cross-modal attention mechanism to generate a fused feature encoding vector; The pixel decoder decodes the fused feature encoding vector to generate a segmentation mask, and the target in the CAD image is segmented based on the segmentation mask to generate a target segmentation image.

3. The method according to claim 2, characterized in that, The CAD image is input into the ViT model to obtain the multi-scale texture feature encoding vector output by the ViT model, including: The CAD image is input into the ViT model, and the ViT model divides the CAD image into at least two non-overlapping blocks, and the position code corresponding to each non-overlapping block is determined. Each non-overlapping block is converted into a corresponding non-overlapping block feature vector through linear mapping; and the non-overlapping block feature vector and the corresponding position code are input into the Transformer encoder to obtain the multi-scale texture feature encoding vector output by the Transformer encoder.

4. The method according to claim 2, characterized in that, Before fusing the geometric feature encoding vector and the multi-scale texture feature encoding vector based on the cross-modal attention mechanism, the following steps are also included: The geometric feature encoding vector and the multi-scale texture feature encoding vector are aligned based on a cross-attention mechanism; The geometric feature encoding vector and the multi-scale texture feature encoding vector are fused based on a cross-modal attention mechanism to generate a fused feature encoding vector, including: The aligned geometric feature encoding vector and the multi-scale texture feature encoding vector are fused based on a cross-modal attention mechanism to generate a fused feature encoding vector.

5. The method according to claim 1, characterized in that, The loss function used in the multimodal segmentation model is a comprehensive loss function consisting of a classification loss function, a regression loss function, and a KL divergence loss function. The classification loss function is used to calculate the difference between the predicted category and the true label. The regression loss function is used to calculate the distance between the mask feature point and the primitive point. The KL divergence loss function is used to calculate the KL divergence between the pixel-level segmentation result and the CAD geometric boundary.

6. The method according to claim 1, characterized in that, The CAD element information includes element location information and the target information to which the element belongs.

7. A multi-mode segmentation device for high-voltage switch design drawings, characterized in that, include: The CAD file acquisition module is used to acquire the CAD file of the high-voltage switch of the power transmission and transformation equipment in response to the multimodal segmentation event of the high-voltage switch design drawings. The graphic element information and image acquisition module is used to determine the CAD graphic element information involved in the CAD file and the CAD image corresponding to the CAD file; The target multimodal segmentation module is used to input the CAD primitive information and the CAD image into a pre-trained multimodal segmentation model to obtain the target segmented image corresponding to the CAD image; wherein, the multimodal segmentation model is a segmentation model that fuses the multi-scale texture features in the CAD image and the geometric features of the CAD primitive information based on a cross-modal attention mechanism.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal segmentation method for high-voltage switch design drawings according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multimodal segmentation method for the high-voltage switch design drawings according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the multimodal segmentation method for high-voltage switch design drawings according to any one of claims 1-6.