A yard wiring diagram link relationship automatic identification method, system, medium and equipment
By constructing and training a multimodal large model, and combining text detection and manual annotation, the problem of insufficient accuracy and flexibility of existing electrical drawing connection relationship recognition methods is solved. This enables efficient recognition and automatic generation of element categories and connection relationships in electrical drawings, supporting the design and maintenance of electrical systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUADIAN ELECTRIC POWER SCI INST CO LTD
- Filing Date
- 2025-01-24
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods for identifying connection relationships in electrical drawings rely on image processing technology, ignoring textual information and domain knowledge in the drawings. This results in insufficient accuracy and completeness, making it difficult to handle complex and ever-changing electrical drawings, and lacking flexibility and adaptability.
We construct a primitive category recognition dataset and a primitive relationship recognition dataset, train a multimodal large model, generate a list containing text boxes and text content through text detection and recognition, perform manual annotation and fine-tuning using an annotation platform, generate a language instruction dataset and train the model, and optimize the model's recognition capabilities by combining manual review and iterative methods.
It significantly improves the accuracy and flexibility of identifying element categories and connection relationships in electrical drawings, can automatically generate power grid and substation wiring structure diagrams, supports the design and maintenance of electrical systems, and enhances the adaptability and robustness of the model.
Smart Images

Figure CN119964189B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method, system, medium, and equipment for automatically identifying connection relationships in a site wiring diagram, belonging to the field of electrical engineering drawing digitization technology. Background Technology
[0002] Existing methods for identifying connection relationships in electrical drawings primarily rely on machine learning techniques to initially identify line segments, typically through image processing and feature extraction. These methods can accurately identify basic elements within the drawing. After identifying the line segments, existing methods then shift to using various non-machine learning, non-intelligent methods to construct the connection relationships between these line segments.
[0003] However, the aforementioned methods are often based on preset rules or templates, lacking flexibility and adaptability, and are ill-suited to handling complex and ever-changing electrical drawings. They exhibit several significant shortcomings in identifying connection relationships in electrical drawings. First, existing methods primarily rely on image processing techniques, neglecting the rich textual information and domain knowledge within the drawings. This results in an inability to fully utilize all the information in the drawings when identifying connection relationships, thus affecting the accuracy and completeness of the identification. Second, electrical drawings typically contain complex connection relationships and a large amount of detailed information. Traditional methods are easily interfered with by complex scenarios such as noise and occlusion, leading to a decrease in recognition accuracy. Furthermore, traditional methods struggle to deeply understand the complex logical topology and functional relationships between devices in electrical drawings, limiting their application in complex electrical systems. On one hand, existing methods mainly rely on image features for recognition, neglecting the importance of domain knowledge, resulting in recognition results that often lack depth and accuracy. On the other hand, electrical drawings contain various types of information, including images, text, and symbols. Existing methods often struggle to effectively integrate this heterogeneous data, thus affecting the comprehensiveness and accuracy of the identification. These shortcomings limit the application and development of existing methods in identifying connection relationships in electrical drawings.
[0004] In summary, existing methods generally exhibit insufficient data-driven capabilities and low levels of intelligence, failing to accurately identify the element categories and connection relationships in CAD station wiring images and automatically generate power grid station wiring structure diagrams. Summary of the Invention
[0005] The purpose of this invention is to provide a method, system, medium, and device for automatically identifying the connection relationships in power grid station wiring diagrams. By constructing a primitive category identification dataset and a primitive relationship identification dataset, and further constructing a primitive category fine-tuning dataset and a primitive connection relationship fine-tuning dataset, a multimodal large model is trained to generate power grid station wiring structure diagrams. This solves the problem that existing technologies cannot accurately identify the primitive categories and connection relationships in CAD station wiring images, and automatically generates power grid station wiring structure diagrams.
[0006] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.
[0007] In a first aspect, the present invention provides a method for automatically identifying the connection relationships in a station wiring diagram, comprising:
[0008] Obtain the CAD site wiring image to be identified, the CAD site wiring image containing multiple electrical components and connecting wires;
[0009] The CAD station wiring images to be identified are input into the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships, respectively. The identification is performed, and the element types and the connection relationships between elements are output respectively, generating a power grid station wiring structure diagram.
[0010] Specifically, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset, and a primitive category fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive categories.
[0011] Based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets are generated, and a primitive link relationship fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive relationships.
[0012] The methods for constructing the primitive category recognition dataset and the primitive relationship recognition dataset include:
[0013] A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images.
[0014] Based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools.
[0015] Furthermore, it also includes performing image preprocessing on the CAD station wiring images after acquiring a large number of CAD station wiring images, including grayscale conversion, Gaussian smoothing, sharpening, and data enhancement operations.
[0016] Furthermore, it also includes, after generating a list containing text boxes and text content, judging the accuracy of text detection and recognition; if the accuracy is lower than a preset accuracy threshold, then using an annotation platform to annotate the text boxes and text content and fine-tune them, the annotation platform including the Label Studio annotation platform.
[0017] Furthermore, based on the list containing text boxes and text content, a primitive category recognition dataset is constructed using annotation tools, including:
[0018] Input the CAD site wiring diagram, which has undergone text detection and recognition, into the annotation platform;
[0019] On the annotation platform, text boxes in CAD station wiring images are associated with corresponding graphic element categories to obtain a graphic element category recognition dataset, whereby the graphic element categories include electrical component types.
[0020] Furthermore, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset, and a primitive category fine-tuning dataset is constructed to train a multimodal large model for primitive category recognition, including:
[0021] Based on the primitive category recognition dataset, generate the first language instruction dataset and the first CAD station wiring image instruction dataset;
[0022] The first language instruction dataset includes an input first language instruction and a first output language instruction. The first input language instruction is a question that asks for the category of the primitive to be classified, and the first output language instruction is an answer that includes the category of the primitive to be classified.
[0023] The first CAD station wiring image instruction dataset includes a first global image instruction and a first local image instruction;
[0024] The first global image instruction is a global CAD site wiring image containing only one text box to be classified;
[0025] The first local image instruction is a local CAD site wiring image centered on the text box to be classified, wherein the local CAD site wiring image centered on the text box to be classified can be randomly segmented;
[0026] The instructions in the first language instruction dataset and the first CAD site wiring image instruction dataset are randomly paired to ensure that each set of instructions contains one input first language instruction, one first global image instruction, and one first local image instruction. The output of each set of instructions corresponds to one first output language instruction.
[0027] Furthermore, based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets are generated, and a primitive link relationship fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive relationships, including:
[0028] Based on the primitive relationships, identify the dataset and generate a second language instruction dataset;
[0029] The second language instruction dataset includes input second language instructions and second output language instructions. The second input language instructions are questions that ask whether the primitive connection relationships are correct, and the second output language instructions are answers that include the primitive connection relationships.
[0030] Based on the primitive relationship recognition dataset, a second CAD station wiring image instruction dataset is generated.
[0031] The second CAD station wiring image instruction dataset includes a second global image instruction and a second local image instruction;
[0032] The second global image instruction is a global CAD site wiring image containing two text boxes;
[0033] The second local image instruction is a local CAD site wiring image centered on a text box, wherein the local CAD site wiring image centered on the text box can be randomly segmented;
[0034] The instructions in the second language instruction dataset and the second CAD station wiring image instruction dataset are randomly paired to ensure that each set of instructions contains one input language instruction, one global image instruction and two local image instructions. The output of each set of instructions corresponds to a second output language instruction, and the output value of the second output language instruction is "yes" or "no".
[0035] Specifically, if the electrical components corresponding to the two text boxes are connected, the second CAD station wiring image instruction dataset is a positive CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "yes". If the electrical components corresponding to the two text boxes are not connected, the second CAD station wiring image instruction dataset is a negative CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "no".
[0036] Furthermore, before generating the power grid station wiring structure diagram, the multimodal large model used to identify element categories and the multimodal large model used to identify element relationships output element types and connection relationships between elements are submitted to the annotation review platform for manual review, confirmation, and correction. The corrected dataset fed back by the annotation review platform is obtained, and the multimodal large model used to identify element categories and the multimodal large model used to identify element relationships are input again. The multimodal large model used to identify element categories and the multimodal large model used to identify element relationships are trained using an iterative method.
[0037] Secondly, the present invention provides an automatic identification system for connection relationships in a station wiring diagram, comprising:
[0038] The data acquisition module is used to acquire the CAD site wiring image to be identified, the CAD site wiring image containing multiple electrical components and connecting wires;
[0039] The image recognition module is used to input the CAD station wiring image to be recognized into the multimodal large model for recognizing the category of graphic elements and the multimodal large model for recognizing the relationship between graphic elements, and then perform recognition.
[0040] The result output module is used to output the element type and the connection relationship between elements respectively, and generate the power grid station wiring structure diagram.
[0041] Specifically, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset, and a primitive category fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive categories.
[0042] Based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets are generated, and a primitive link relationship fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive relationships.
[0043] The methods for constructing the primitive category recognition dataset and the primitive relationship recognition dataset include:
[0044] A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images.
[0045] Based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools.
[0046] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the automatic identification method for connection relationships in a site wiring diagram as described in the first aspect.
[0047] Fourthly, the present invention provides a computer device, comprising:
[0048] Memory, used to store instructions;
[0049] A processor is configured to execute the instructions, causing the device to perform operations that implement the automatic identification method for connection relationships in the site wiring diagram as described in the first aspect.
[0050] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0051] 1. This invention achieves automatic identification of electrical components and their connection relationships in CAD station wiring images by constructing and training a multimodal large model for identifying element categories and element relationships. It does not rely on preset rules or templates, significantly improving the flexibility and adaptability of identification. It can not only accurately identify element categories and connection relationships in complex electrical drawings, but also automatically generate power grid station wiring structure diagrams, providing an efficient and accurate tool for the design and maintenance of electrical systems.
[0052] 2. This invention employs a rigorous annotation process when constructing the primitive category recognition dataset and the primitive relationship recognition dataset. First, a list containing text boxes and text content is generated through text detection and recognition, and its accuracy is assessed. If the accuracy is insufficient, manual annotation and fine-tuning are performed using an annotation platform such as Label Studio to ensure the accuracy and completeness of the data, greatly improving the quality of the dataset and providing strong support for training an efficient and accurate recognition model. Simultaneously, language instruction and image instruction datasets generated from the dataset are further optimized through random pairing and fine-tuning to enhance the model's recognition capabilities.
[0053] 3. Before generating the wiring structure diagram of the power grid station, the present invention adds a manual review and correction step. Through the annotation review platform, the primitive types and connection relationships between primitives output by the multimodal large model are manually reviewed to obtain corrected data. This not only ensures the accuracy and reliability of the recognition results, but also uses an iterative method to re-input the corrected dataset into the model for training, continuously improving the model's recognition accuracy and generalization ability. This makes the method of the present invention more adaptable and robust when dealing with complex and ever-changing electrical drawings.
[0054] 4. This invention generates a first language instruction dataset and a first CAD station wiring image instruction dataset, and constructs a primitive category fine-tuning dataset, thereby achieving multimodal information fusion training. It not only makes full use of the visual features in the CAD station wiring images, but also integrates the semantic information in the language instructions, thus significantly improving the ability of the multimodal large model used to identify primitive types to identify primitive categories.
[0055] 5. This invention generates a second language instruction dataset and positive and negative CAD site wiring image instruction datasets, and constructs a primitive link relationship fine-tuning dataset, realizing paired training of positive and negative samples. This not only improves the recognition accuracy of the multimodal large model used to identify primitive relationships for correct connections, but also enhances the ability of the multimodal large model used to identify primitive relationships to distinguish incorrect connections, thereby improving the overall recognition accuracy and reliability. Attached Figure Description
[0056] Figure 1The diagram shown is a flowchart illustrating an automatic identification method for connection relationships in a site wiring diagram according to an embodiment of the present invention.
[0057] Figure 2 The diagram shown is a schematic representation of the text detection and recognition process of a conventional method provided in this embodiment of the invention.
[0058] Figure 3 The diagram shown illustrates the working principle of the multimodal large model for identifying primitive categories provided in an embodiment of the present invention.
[0059] Figure 4 The diagram shown is an application example of the multimodal large model for identifying primitive categories provided in this embodiment of the invention.
[0060] Figure 5 The diagram shown illustrates the working principle of a multimodal large model for identifying primitive relationships provided in an embodiment of the present invention.
[0061] Figure 6 The diagram shown is an application example of a multimodal large model for identifying primitive relationships provided by an embodiment of the present invention.
[0062] Figure 7 The figure shown is a schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0063] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0064] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0065] Example 1
[0066] like Figure 1 As shown in the figure, this embodiment introduces a method for automatically identifying the connection relationships in a station wiring diagram, including:
[0067] Step 1: Obtain the CAD site wiring image to be identified, which contains multiple electrical components and connecting wires.
[0068] Step 2: Input the CAD station wiring image to be identified into the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships, respectively, and perform identification.
[0069] Step 2.1: Construct the primitive category recognition dataset and the primitive relationship recognition dataset.
[0070] In this embodiment, the method for constructing the primitive category recognition dataset and the primitive relationship recognition dataset includes:
[0071] A flowchart illustrating the traditional text detection and recognition process is shown below. Figure 2 As shown, this embodiment utilizes mature OCR tools such as PaddleOCR for text detection and recognition, including the following steps:
[0072] A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images.
[0073] In some embodiments, the method further includes, after acquiring a large number of CAD site wiring images, performing image preprocessing on the CAD site wiring images, including grayscale conversion, Gaussian smoothing, sharpening, and data enhancement operations.
[0074] Image preprocessing is an essential step in the digitization process. Basic image processing techniques such as grayscale conversion, Gaussian smoothing, and sharpening remove noise and highlight edges, improving subsequent recognition accuracy. Furthermore, data augmentation techniques enhance the robustness of the model by randomly transforming and augmenting the image, enabling it to adapt to more image variations.
[0075] In some embodiments, based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools, including:
[0076] Input the CAD site wiring diagram, which has undergone text detection and recognition, into the annotation platform;
[0077] On the annotation platform, text boxes in CAD station wiring images are associated with corresponding graphic element categories to obtain a graphic element category recognition dataset. The graphic element categories include electrical component types, such as circuit breakers, busbars, and transformers.
[0078] In some embodiments, the method further includes, after generating a list containing text boxes and text content, determining the accuracy of text detection and recognition; if the accuracy is lower than a preset accuracy threshold, then using an annotation platform to annotate the text boxes and text content and make fine adjustments, wherein the annotation platform includes the Label Studio annotation platform.
[0079] Step 2.2: Generate the first language instruction dataset and the first CAD station wiring image instruction dataset based on the primitive category recognition dataset, and construct a primitive category fine-tuning dataset to train a multimodal large model for primitive category recognition. The working principle of the multimodal large model for primitive category recognition is as follows: Figure 3 As shown.
[0080] Step 2.2.1: Based on the primitive category identification dataset, generate the first language instruction dataset and the first CAD station wiring image instruction dataset;
[0081] The first language instruction dataset includes an input first language instruction and a first output language instruction. The first input language instruction is a question that asks for the category of the primitive to be classified, and the first output language instruction is an answer that includes the category of the primitive to be classified.
[0082] The first CAD station wiring image instruction dataset includes a first global image instruction and a first local image instruction;
[0083] The first global image instruction is a global CAD site wiring image containing only one text box to be classified;
[0084] The first local image instruction is a local CAD site wiring image centered on the text box to be classified, wherein the local CAD site wiring image centered on the text box to be classified can be randomly segmented;
[0085] Step 2.2.2: Randomly pair the instructions in the first language instruction dataset and the first CAD site wiring image instruction dataset to ensure that each set of instructions contains one input first language instruction, one first global image instruction, and one first local image instruction. The output of each set of instructions corresponds to one first output language instruction.
[0086] This embodiment provides an application example diagram of a multimodal large model for identifying primitive categories, as shown in the following diagram. Figure 4 As shown.
[0087] Step 2.3: Generate a second language instruction dataset and positive and negative CAD site wiring image instruction datasets based on the primitive relationship recognition dataset, and construct a primitive link relationship fine-tuning dataset to train a multimodal large model for recognizing primitive relationships. The working principle of the multimodal large model for recognizing primitive relationships is as follows: Figure 5 As shown.
[0088] Step 2.3.1: Identify the dataset based on the primitive relationships and generate a second language instruction dataset;
[0089] The second language instruction dataset includes input second language instructions and second output language instructions. The second input language instructions are questions that ask whether the primitive connection relationships are correct, and the second output language instructions are answers that include the primitive connection relationships.
[0090] Step 2.3.2: Based on the primitive relationships, identify the dataset and generate the second CAD station wiring image instruction dataset;
[0091] The second CAD station wiring image instruction dataset includes a second global image instruction and a second local image instruction;
[0092] The second global image instruction is a global CAD site wiring image containing two text boxes;
[0093] The second local image instruction is a local CAD site wiring image centered on a text box, wherein the local CAD site wiring image centered on the text box can be randomly segmented;
[0094] Step 2.3.3: Randomly pair the instructions in the second language instruction dataset and the second CAD site wiring image instruction dataset to ensure that each set of instructions contains one input language instruction, one global image instruction and two local image instructions. The output of each set of instructions corresponds to one second output language instruction, and the output value of the second output language instruction is "yes" or "no".
[0095] Specifically, if the electrical components corresponding to the two text boxes are connected, the second CAD station wiring image instruction dataset is a positive CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "yes". If the electrical components corresponding to the two text boxes are not connected, the second CAD station wiring image instruction dataset is a negative CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "no".
[0096] This embodiment provides an application example diagram of a multimodal large model for identifying primitive relationships, as shown in the following diagram. Figure 6 As shown.
[0097] Step 3: Output the element types and the connection relationships between elements respectively, and generate the power grid station wiring structure diagram.
[0098] In some embodiments, before generating the power grid station wiring structure diagram, the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships output by the multimodal large model are submitted to the annotation review platform for manual review, confirmation and correction, to obtain the corrected dataset fed back by the annotation review platform, and the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships are input again, and the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships are trained using an iterative method.
[0099] In some embodiments, based on the correction data fed back by the annotation review platform, the multimodal large model is continuously optimized and fine-tuned to achieve data-driven model iteration. Through multiple iterations, the model continuously improves its recognition accuracy, gradually achieving automated recognition and accurate processing of complex wiring diagrams.
[0100] The power grid station wiring diagram provided by this invention can provide dispatching and maintenance personnel with accurate graphic element information and topological relationships, support real-time dispatching decisions and production management, and can be applied to specific scenarios such as equipment maintenance planning, operation mode changes, and fault analysis. Technical personnel in related fields can also use this invention to combine the power grid station wiring diagram with specific business needs, support the intelligent construction of power plants / power grids, and apply it to specific scenarios such as primary equipment information management, secondary circuit diagram analysis, and equipment ledger management.
[0101] Example 2
[0102] Based on the same inventive concept as Embodiment 1, this embodiment introduces an automatic identification system for connection relationships in a site wiring diagram, comprising:
[0103] The data acquisition module is used to acquire the CAD site wiring image to be identified, the CAD site wiring image containing multiple electrical components and connecting wires;
[0104] The image recognition module is used to input the CAD station wiring image to be recognized into the multimodal large model for recognizing the category of graphic elements and the multimodal large model for recognizing the relationship between graphic elements, and then perform recognition.
[0105] The result output module is used to output the element type and the connection relationship between elements respectively, and generate the power grid station wiring structure diagram.
[0106] Specifically, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset, and a primitive category fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive categories.
[0107] Based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets are generated, and a primitive link relationship fine-tuning dataset is constructed to train a multimodal large model for recognizing primitive relationships.
[0108] The methods for constructing the primitive category recognition dataset and the primitive relationship recognition dataset include:
[0109] A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images.
[0110] Based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools.
[0111] The specific functions of each module described above are explained in the relevant content of the method in Embodiment 1, and will not be repeated here.
[0112] Example 3
[0113] Based on the same inventive concept as other embodiments, and based on the above... Figure 1 Accordingly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the big data-driven consumer preference analysis method of any of the above embodiments.
[0114] Example 4
[0115] Based on the same inventive concept as other embodiments, and based on the above... Figure 1 The method shown in the embodiment of the present invention also provides a physical structure diagram of a computer device, such as... Figure 7 As shown, the computer device may include a communication bus, a processor, a memory, and a communication interface. It may also include input / output interfaces and a display device. The various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor executes the program stored in the memory to perform the steps of the automatic identification method for the connection relationship of the site wiring diagram described in the above embodiments.
[0116] In summary, by constructing and training a multimodal large model for identifying element categories and element relationships, this invention achieves automatic identification of electrical components and their connection relationships in CAD station wiring images. This is done without relying on preset rules or templates, significantly improving the flexibility and adaptability of identification. It can not only accurately identify element categories and connection relationships in complex electrical drawings, but also automatically generate power grid station wiring structure diagrams, providing an efficient and accurate tool for the design and maintenance of electrical systems.
[0117] This invention employs a rigorous annotation process when constructing the primitive category recognition dataset and the primitive relationship recognition dataset. First, a list containing text boxes and text content is generated through text detection and recognition, and its accuracy is assessed. If the accuracy is insufficient, manual annotation and fine-tuning are performed using an annotation platform such as Label Studio to ensure the accuracy and completeness of the data, significantly improving the quality of the dataset and providing strong support for training an efficient and accurate recognition model. Simultaneously, language instruction and image instruction datasets generated from the dataset are further optimized through random pairing and fine-tuning to enhance the model's recognition capabilities.
[0118] Before generating the wiring structure diagram of the power grid station, this invention adds a manual review and correction step. Through the annotation review platform, the types of graphic elements output by the multimodal large model and the connection relationships between graphic elements are manually reviewed to obtain corrected data. This not only ensures the accuracy and reliability of the recognition results, but also uses an iterative method to re-input the corrected dataset into the model for training, continuously improving the model's recognition accuracy and generalization ability. This makes the method of this invention more adaptable and robust when dealing with complex and ever-changing electrical drawings.
[0119] This invention generates a first language instruction dataset and a first CAD station wiring image instruction dataset, and constructs a primitive category fine-tuning dataset, thereby achieving multimodal information fusion training. It not only makes full use of the visual features in the CAD station wiring images, but also integrates the semantic information in the language instructions, thus significantly improving the ability of the multimodal large model used to identify primitive types to identify primitive categories.
[0120] This invention generates a second language instruction dataset and positive and negative CAD site wiring image instruction datasets, and constructs a primitive link relationship fine-tuning dataset. It achieves paired training of positive and negative samples, which not only improves the recognition accuracy of the multimodal large model for recognizing primitive relationships for correct connections, but also enhances the ability of the multimodal large model for recognizing primitive relationships for erroneous connections, thereby improving the overall recognition accuracy and reliability.
[0121] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0125] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for automatically identifying connection relationships in a station wiring diagram, characterized in that, include: Obtain the CAD site wiring image to be identified, the CAD site wiring image containing multiple electrical components and connecting wires; The CAD station wiring images to be identified are input into the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships, respectively. The identification is performed, and the element types and the connection relationships between elements are output respectively, generating a power grid station wiring structure diagram. Specifically, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset. A primitive category fine-tuning dataset is constructed, and a multimodal large model for primitive category recognition is trained, including: Based on the primitive category recognition dataset, generate the first language instruction dataset and the first CAD station wiring image instruction dataset; The first language instruction dataset includes a first input language instruction and a first output language instruction. The first input language instruction is a question that asks for the category of the primitive to be classified, and the first output language instruction is an answer that includes the category of the primitive to be classified. The first CAD station wiring image instruction dataset includes a first global image instruction and a first local image instruction; The first global image instruction is a global CAD site wiring image containing only one text box to be classified; The first local image instruction is a local CAD site wiring image centered on the text box to be classified, wherein the local CAD site wiring image centered on the text box to be classified is randomly segmented; The instructions in the first language instruction dataset and the first CAD station wiring image instruction dataset are randomly paired to ensure that a set of instructions contains a first input language instruction, a first global image instruction and a first local image instruction, wherein the output of each set of instructions corresponds to a first output language instruction. Based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets were generated. A primitive link relationship fine-tuning dataset was constructed, and a multimodal large-scale model for recognizing primitive relationships was trained, including: Based on the primitive relationships, identify the dataset and generate a second language instruction dataset; The second language instruction dataset includes a second input language instruction and a second output language instruction. The second input language instruction is a question that asks whether the primitive connection relationship is correct, and the second output language instruction is an answer that includes the primitive connection relationship. Based on the primitive relationship recognition dataset, a second CAD station wiring image instruction dataset is generated. The second CAD station wiring image instruction dataset includes a second global image instruction and a second local image instruction; The second global image instruction is a global CAD site wiring image containing two text boxes; The second local image instruction is a local CAD site wiring image centered on a text box, wherein the local CAD site wiring image centered on the text box is randomly segmented; The instructions in the second language instruction dataset and the second CAD site wiring image instruction dataset are randomly paired to ensure that each set of instructions contains one input language instruction, one global image instruction and two local image instructions. The output of each set of instructions corresponds to a second output language instruction, and the output value of the second output language instruction is "yes" or "no". Wherein, if the electrical components corresponding to the two text boxes are connected, the second CAD station wiring image instruction dataset is a positive CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "yes"; if the electrical components corresponding to the two text boxes are not connected, the second CAD station wiring image instruction dataset is a negative CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "no". The methods for constructing the primitive category recognition dataset and the primitive relationship recognition dataset include: A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images. Based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools.
2. The method for automatically identifying connection relationships in a station wiring diagram according to claim 1, characterized in that, It also includes performing image preprocessing on the CAD site wiring images after acquiring a large number of CAD site wiring images, including grayscale conversion, Gaussian smoothing, sharpening, and data enhancement operations.
3. The method for automatically identifying connection relationships in a station wiring diagram according to claim 1, characterized in that, It also includes, after generating a list containing text boxes and text content, judging the accuracy of text detection and recognition; if the accuracy is lower than a preset accuracy threshold, then using an annotation platform to annotate the text boxes and text content and make fine adjustments, the annotation platform including the Label Studio annotation platform.
4. The method for automatically identifying connection relationships in a station wiring diagram according to claim 1, characterized in that, Based on the list containing text boxes and text content, a primitive category recognition dataset is constructed using annotation tools, including: Input the CAD site wiring diagram, which has undergone text detection and recognition, into the annotation platform; On the annotation platform, text boxes in CAD station wiring images are associated with corresponding graphic element categories to obtain a graphic element category recognition dataset, whereby the graphic element categories include electrical component types.
5. The method for automatically identifying connection relationships in a station wiring diagram according to claim 1, characterized in that, It also includes submitting the outputs of the multimodal large model for identifying element categories and the multimodal large model for identifying element relationships to the annotation review platform before generating the power grid station wiring structure diagram. The outputs of these two models are then manually reviewed, confirmed, and corrected to obtain the corrected dataset from the annotation review platform. The multimodal large models for identifying element categories and identifying element relationships are then input again, and the multimodal large models for identifying element categories and identifying element relationships are trained using an iterative method.
6. An automatic identification system for connection relationships in a station wiring diagram, characterized in that, include: The data acquisition module is used to acquire the CAD site wiring image to be identified, the CAD site wiring image containing multiple electrical components and connecting wires; The image recognition module is used to input the CAD station wiring image to be recognized into the multimodal large model for recognizing the category of graphic elements and the multimodal large model for recognizing the relationship between graphic elements, and then perform recognition. The result output module is used to output the element type and the connection relationship between elements respectively, and generate the power grid station wiring structure diagram. Specifically, a first language instruction dataset and a first CAD station wiring image instruction dataset are generated based on the primitive category recognition dataset. A primitive category fine-tuning dataset is constructed, and a multimodal large model for primitive category recognition is trained, including: Based on the primitive category recognition dataset, generate the first language instruction dataset and the first CAD station wiring image instruction dataset; The first language instruction dataset includes a first input language instruction and a first output language instruction. The first input language instruction is a question that asks for the category of the primitive to be classified, and the first output language instruction is an answer that includes the category of the primitive to be classified. The first CAD station wiring image instruction dataset includes a first global image instruction and a first local image instruction; The first global image instruction is a global CAD site wiring image containing only one text box to be classified; The first local image instruction is a local CAD site wiring image centered on the text box to be classified, wherein the local CAD site wiring image centered on the text box to be classified is randomly segmented; The instructions in the first language instruction dataset and the first CAD station wiring image instruction dataset are randomly paired to ensure that a set of instructions contains a first input language instruction, a first global image instruction and a first local image instruction, wherein the output of each set of instructions corresponds to a first output language instruction. Based on the primitive relationship recognition dataset, a second language instruction dataset and positive and negative CAD station wiring image instruction datasets were generated. A primitive link relationship fine-tuning dataset was constructed, and a multimodal large-scale model for recognizing primitive relationships was trained, including: Based on the primitive relationships, identify the dataset and generate a second language instruction dataset; The second language instruction dataset includes a second input language instruction and a second output language instruction. The second input language instruction is a question that asks whether the primitive connection relationship is correct, and the second output language instruction is an answer that includes the primitive connection relationship. Based on the primitive relationship recognition dataset, a second CAD station wiring image instruction dataset is generated. The second CAD station wiring image instruction dataset includes a second global image instruction and a second local image instruction; The second global image instruction is a global CAD site wiring image containing two text boxes; The second local image instruction is a local CAD site wiring image centered on a text box, wherein the local CAD site wiring image centered on the text box is randomly segmented; The instructions in the second language instruction dataset and the second CAD station wiring image instruction dataset are randomly paired to ensure that each set of instructions contains one input language instruction, one global image instruction and two local image instructions. The output of each set of instructions corresponds to one second output language instruction, and the output value of the second output language instruction is "yes" or "no". If the electrical components corresponding to the two text boxes are connected, the second CAD station wiring image instruction dataset is a positive CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "yes". If the electrical components corresponding to the two text boxes are not connected, the second CAD station wiring image instruction dataset is a negative CAD station wiring image instruction dataset, and the corresponding second output language instruction output value is "no". The methods for constructing the primitive category recognition dataset and the primitive relationship recognition dataset include: A large number of CAD site wiring images are acquired for text detection and recognition, generating a list containing text boxes and text content. The text boxes mark the position of the text content for each electrical component in the CAD site wiring images. Based on the list containing text boxes and text content, a primitive category recognition dataset and a primitive relationship recognition dataset are constructed using annotation tools.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the automatic identification method for connection relationships in the station wiring diagram as described in any one of claims 1-5.
8. A computer device, characterized in that, include: Memory, used to store instructions; A processor is configured to execute the instructions, causing the device to perform operations that implement the automatic identification method for connection relationships in a site wiring diagram as described in any one of claims 1-5.