Electric power communication construction drawing information extraction method and system based on large model

By building efficient large models and CRNN models, text recognition of power communication construction drawings, and fine-tuning the model using improved LoRA technology, the problems of low identification accuracy and efficiency in the existing technology are solved, and automated and high-precision information extraction is achieved.

CN120107983AActive Publication Date: 2025-06-06NARI INFORMATION & COMM TECH
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510585932.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The prior art has problems with low recognition accuracy and efficiency in the extraction of power communication construction drawing information. Convolutional neural networks have weak understanding of the global context, and BiLSTMs have weakened their capture capabilities when processing long sequences.

Method used

The large model-based power communication construction drawing information extraction method is adopted, and the large model is optimized by building an efficient large model, combining computer vision technology, using CRNN models for text recognition, and fine-tuning the large model through improved LoRA technology, enhancing the model's expression ability and training flexibility.

Benefits of technology

It realizes automated and high-precision identification of key information in power communication construction drawings, improves identification accuracy and efficiency, and reduces manual identification errors and dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107983A_ABST
    Figure CN120107983A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power communication construction drawing information extraction method and system based on a large model, and the method comprises the steps: obtaining an image of a to-be-recognized electric power communication construction drawing, and carrying out the preprocessing of the image; performing text recognition on the preprocessed construction drawing image by using a constructed and trained CRNN model, performing post-processing on the recognized text, and converting the post-processed text into supervision instruction data; performing fine tuning on the large model based on the supervision instruction data to obtain a fine-tuned large model; and performing key information extraction on a text identified based on the CRNN model by using the fine-tuned large model. According to the method, the large model is finely adjusted, the computer vision technology is combined, automatic and high-precision recognition of key information in the electric power communication construction drawing is achieved, and the recognition accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric power communication construction, and in particular relates to a method and system for extracting information from electric power communication construction drawings based on a large model. Background Art

[0002] Power communication construction drawings contain a large amount of complex information, such as equipment layout, line direction, connection relationships, etc. The traditional manual identification method is not only time-consuming and labor-intensive, but also prone to errors. The recognition accuracy and efficiency need to be improved.

[0003] Prior art document 1 (CN118015648A) discloses a drawing information extraction method, device, equipment and storage medium, including obtaining drawings to be processed; extracting primitive information from the drawings to be processed based on a pre-trained primitive extraction model; extracting relationship information between primitives in the drawings to be processed based on the primitive information and a pre-trained relationship extraction model; determining the component name corresponding to each primitive based on the primitive information, the relationship information and a pre-trained component recognition model.

[0004] Prior art document 2 (CN111860348A) discloses a weakly supervised OCR recognition method for electric power drawings based on deep learning, including using a pre-trained text detection model to detect the image to be recognized and predicting the text area box at the entire word level; performing text recognition on the predicted text area box, using character cutting to obtain single character text for vertical text, and directly using text lines for horizontal text, and then recognizing it through a CNN+BiLSTM+CTC model; post-processing the recognition results, and judging and modifying the results through prior knowledge to improve the accuracy.

[0005] However, the disadvantage of the prior art document 1 is that the convolutional neural network is used to extract the information of the network element, and the CNN has a weak understanding of the global context. The convolution layer in AlexNet uses a larger convolution kernel and stride, which may lose some fine-grained local features, especially for fine graphics extraction. The effect may not be good.

[0006] The shortcoming of the prior art document 2 is that although BiLSTM can capture the bidirectional dependency of the sequence, its ability to capture long-distance dependencies will gradually weaken when processing long sequences. This may cause the model to be unable to accurately understand the overall semantics of the text, thereby affecting the accuracy of recognition. Summary of the invention

[0007] In order to solve the deficiencies in the prior art, the present invention provides a method and system for extracting information from power communication construction drawings based on a large model. By constructing an efficient large model and combining it with computer vision technology, automatic and high-precision recognition of key information in power communication construction drawings is achieved, thereby improving recognition accuracy and efficiency.

[0008] The present invention adopts the following technical solution.

[0009] A first aspect of the present invention provides a method for extracting information from power communication construction drawings based on a large model, comprising: S1: Acquire the image of the power communication construction drawing to be identified and perform preprocessing; S2: Use the constructed and trained CRNN model to perform text recognition on the construction drawing images preprocessed in S1 and post-process the recognized text, and then convert the post-processed text into supervision instruction data; S3: fine-tune the large model based on the supervised instruction data to obtain a fine-tuned large model; S4: Use the fine-tuned large model to extract key information from the text recognized by the CRNN model.

[0010] Optionally, in S1, the preprocessing includes part or all of binarization, grayscale, denoising, and rotation correction.

[0011] Optionally, in S2, training the CRNN model includes: Acquire various types of power communication construction drawing images, annotate each character in each power communication construction drawing image, and obtain original power communication construction drawing images and annotated power communication construction drawing images; Performing image preprocessing on the original power communication construction drawing images and the annotated power communication construction drawing images to form a data set, and dividing the data set into a training set, a test set, and a validation set according to a preset ratio; Input the training set into the constructed CRNN model for training, and construct the CRNN model loss function based on the output result of the CRNN model and the training label; Use the validation set to judge the performance of the CRNN model on unseen data through the first evaluation indicator during the CRNN model training process, adjust the hyperparameters according to the performance, and obtain the optimal CRNN model; The test set is used to evaluate the recognition accuracy of the trained CRNN model based on the second evaluation indicator.

[0012] Optionally, image preprocessing is performed on the power communication construction drawings, including: The original image of the electric power communication construction drawing is converted into a grayscale image, and the grayscale image is converted into a binary image and then the noise in the binary image is removed; The straight line in the denoised binary image is identified using an edge algorithm, the straight line parameters in the denoised binary image are extracted using Hough transform, and the tilt angle of the original image of the power communication construction drawing is estimated based on the straight line parameters; The original image of the electric power communication construction drawing is rotationally corrected according to the estimated tilt angle so that the text lines in the electric power communication construction drawing are aligned.

[0013] Optionally, construct the CRNN model loss function according to the following formula:

[0014] in, is the loss value, is the target label sequence All permutations of is the time step The feature vector corresponding to the preprocessed construction drawing image on is the time step The predicted label, For in time The predicted label is The probability of , T is the length of the input sequence, is the target label sequence.

[0015] Optionally, the validation set is used to judge the performance of the CRNN model on unseen data by a first evaluation indicator during the CRNN model training process, and the hyperparameters are adjusted according to the performance to obtain the optimal CRNN model, including: Constructing an objective function according to the first evaluation indicator and its corresponding weight; Random sampling is performed in the hyperparameter space to construct multiple CRNN models. The validation set is used to train multiple CRNN models to obtain the objective functions corresponding to different hyperparameter combinations, and the objective functions corresponding to different hyperparameter combinations are sorted from small to large. Based on the Bayesian optimization algorithm, the models corresponding to the hyperparameter combinations corresponding to the top k objective functions are trained, and the model corresponding to the hyperparameter combination corresponding to the minimum objective function is selected as the optimal CRNN model.

[0016] Optionally, the constructed CRNN model includes CNN layer, feature fusion layer, two-head self-attention mechanism, bidirectional LSTM and CTC layer; The CNN layer is used to extract features from the image text lines of the power communication construction drawings to obtain a feature map, wherein the features in the feature map include the shape of the characters, the edges of the characters, the texture of the characters, the layout between the characters, and the structure of the characters; the CNN layer includes a plurality of convolutional layers; The feature fusion layer connects multiple convolutional layers and is used to fuse the feature maps extracted by the multiple convolutional layers; The fused feature map is converted into a feature sequence and then input into the dual-head attention mechanism. The output of the two heads is weighted summed to obtain the final feature sequence. The forward LSTM in the bidirectional LSTM reads the forward information of the final feature sequence, and the reverse LSTM reads the reverse information of the final feature sequence. The outputs of the forward LSTM and the reverse LSTM are concatenated at each time step to obtain the feature vector of each time step and output the probability distribution of all characters. The CTC layer aligns the sequence according to the probability distribution of characters and calculates the loss, and the final character sequence is extracted from the output of the CTC layer through a decoding strategy.

[0017] Optionally, the feature maps extracted from multiple convolutional layers are fused in the feature fusion layer according to the following formula:

[0018] in, Indicates The feature map after layer fusion, Indicates The feature map of the layer, represents a channel attention mechanism, For the general The processed feature map is concatenated with the feature map of the previous layer. Indicates the maximum pooling operation on the concatenated feature map. , H and W represent the height and width of the feature map respectively, and C is the number of channels; Among them, the channel attention mechanism learns the importance of each channel and dynamically adjusts the weight of the channel. Feature map of the layer Multiply to obtain the weighted feature map .

[0019] Optionally, the feature fusion layer includes multiple fully connected layers, and the channel attention mechanism dynamically adjusts the weight of each channel by learning the importance of each channel, specifically including: Use channel attention mechanism to process the feature map of layer i , perform global average pooling on each channel to obtain a channel description vector; The channel description vector is compressed through the first fully connected layer, and then the compressed channel description vector is input into the ReLU activation function; The output result of the ReLU activation function is input into the second fully connected layer, and the output of the second fully connected layer is passed through the sigmoid activation function to obtain the weight of each channel.

[0020] Optionally, the large model is fine-tuned using the improved LoRA technology. During the fine-tuning process, the matrix A and the sub-matrices of the matrix B are used to perform Hadamard products, the parameter matrix is ​​obtained according to the Hadamard product, and the number of Hadamard products is dynamically adjusted.

[0021] Optionally, during the fine-tuning process, a sub-matrix of matrix A and matrix B is used to perform a Hadamard product, a parameter matrix is ​​obtained according to the Hadamard product, and the number of Hadamard products is dynamically adjusted, including: Slice the matrix B by columns to obtain submatrices of B. When r / s is an integer, each submatrix of B contains s columns, and r / s submatrices of B are obtained. When r / s is not an integer, r / s+1 submatrices of B are obtained. Slice the matrix A by rows to obtain submatrices of A. When r / s is an integer, each submatrix of A contains s rows, and r / s submatrices of A are obtained. When r / s is not an integer, r / s+1 submatrices of A are obtained, where r is the rank of the A matrix and the B matrix, and s is the step size. Perform matrix product of the corresponding sub-matrices of B and A:

[0022] in, , Indicates the sub-matrices, Indicates the sub-matrices, N represents the number of sub-matrices; N Perform the Hadamard product in order to obtain the parameter matrix:

[0023] in, The parameter matrix representing the parameter update during fine-tuning, Denotes the Hadamard product.

[0024] Optionally, the method further includes: The key information extracted from the large model is converted into structured data, and the converted key information and power communication construction drawings are stored in the database in a set format.

[0025] A second aspect of the present invention provides a large model-based power communication construction drawing information extraction system, comprising: A data acquisition module, used to obtain images of power communication construction drawings to be identified; A drawing recognition module is used to use the constructed and trained CRNN model to perform text recognition on images of power communication construction drawings and post-process the recognized text, and then convert the post-processed text into supervision instruction data; A fine-tuning module is used to fine-tune the large model based on the supervision instruction data to obtain a fine-tuned large model; The large model information extraction module is used to extract key information from the text recognized by the CRNN model using the fine-tuned large model.

[0026] Optionally, the system also includes an output module and a data storage module, the output module is used to convert the key information output by the large model information extraction module into structured data, and the data storage module is used to store the structured data and power communication construction drawings in a set format in a database.

[0027] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the method for extracting information from power communication construction drawings based on a large model is implemented.

[0028] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned method for extracting information from power communication construction drawings based on a large model.

[0029] Compared with the prior art, the beneficial effects of the present invention include at least: The present invention uses visual technology combined with large model technology to automatically and accurately extract information from power communication construction drawings, reduce manual recognition errors, and improve recognition accuracy. The automated recognition process not only significantly shortens construction preparation time and improves construction efficiency, but also reduces dependence on professionals and reduces labor costs.

[0030] The present invention adds a downsampling branch on the basis of the convolution layer to fuse the features from different convolution layers, dynamically adjusts the learning rate, adds a dual-head self-attention mechanism before BiLSTM, uses more stringent evaluation indicators to evaluate the drawing recognition effect, and uses the improved LoRA algorithm to fine-tune the large model; the present invention improves the feature extraction and multi-scale perception capabilities, so that the model can more comprehensively capture the low-level and high-level features of the image and adapt to complex drawing recognition tasks. The LoRA improved fine-tuning algorithm enhances the expression ability of the model, and the step size mechanism improves the flexibility of training and helps the model converge more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them: Figure 1 A schematic diagram of the architecture of a large-model-based power communication construction drawing information extraction system provided in an embodiment of the present invention; Figure 2 A schematic diagram of a large-model-based power communication construction drawing information extraction process provided by an embodiment of the present invention; Figure 3 A schematic diagram of constructing a CRNN model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The embodiments described in this application are only embodiments of a part of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the protection scope of the present invention.

[0033] Combination Figures 1 to 3 As shown, embodiment 1 of the present invention provides a method for extracting information from power communication construction drawings based on a large model, comprising the following steps: S1: Acquire an image of a power communication construction drawing to be identified and perform preprocessing.

[0034] The image format of the power communication construction drawing in the embodiment of the present invention may be JPEG, BMP or PNG, which is not limited here.

[0035] In an exemplary implementation of an embodiment of the present invention, the drawing information extraction system may obtain drawings to be identified uploaded by a user through a preset drawing upload interface, wherein the drawings to be identified may be multiple drawings of a construction object.

[0036] Optionally, obtaining the image of the power communication construction drawing includes: obtaining the original image of the power communication construction drawing; performing a preprocessing operation on the original image to obtain the image of the power communication construction drawing to be identified, and the preprocessing operation includes part or all of grayscale, binarization, denoising, and rotation correction.

[0037] In the disclosed embodiment, the image quality can be improved by processing the collected original image of the power communication construction drawing.

[0038] S2: Use the constructed and trained CRNN model to perform text recognition on the construction drawing graphics preprocessed in S1 and post-process the recognized text, and then convert the post-processed text into supervision instruction data.

[0039] Optionally, the CRNN model includes a feature extraction model and a drawing recognition model. The feature extraction model may be a convolutional neural network (CNN), and the drawing recognition model may be a recurrent neural network (RNN), which are trained by massive power communication construction drawing data. The CRNN model can learn the features of Chinese and English characters, symbols, numbers, etc. in the construction drawings, so as to accurately recognize the Chinese and English characters, symbols, numbers, etc. in the drawings.

[0040] Specifically, the model is pre-trained using public data sets in other fields to initialize the model parameters, and then a large number of annotated power communication construction drawing data sets are used to train the pre-trained model. By adjusting the model parameters, optimizing the network structure, and introducing the attention mechanism, the recognition accuracy and generalization ability of the model are improved. The training of the CRNN model in S2 includes: S2.1: Acquire various types of electric power communication construction drawing images, annotate each character in each electric power communication construction drawing image, and obtain the original electric power communication construction drawing image and the annotated electric power communication construction drawing image text.

[0041] Specifically, the annotated data is saved in a specific format, which can be json, xml, or txt, and the name is consistent with the drawing name.

[0042] S2.2: Perform image preprocessing on the original power communication construction images, combine the processed images with the corresponding annotated texts to form a data set, and divide the data set into a training set, a test set, and a validation set according to a preset ratio.

[0043] Specifically, the data set is divided into training set, test set, and validation set in a ratio of 8:1:1.

[0044] Optional, combined Figure 2 As shown, image preprocessing is performed on the power communication construction drawings, including: The original image of the electric power communication construction drawing is converted into a grayscale image, and the grayscale image is converted into a binary image and then the noise in the binary image is removed; The straight line in the denoised binary image is identified using an edge algorithm, the straight line parameters in the denoised binary image are extracted using Hough transform, and the tilt angle of the original image of the power communication construction drawing is estimated based on the straight line parameters; The original image of the electric power communication construction drawing is rotationally corrected according to the estimated tilt angle so that the text lines in the electric power communication construction drawing are aligned.

[0045] Specifically, image preprocessing of power communication construction drawings includes the following steps: S2.2.1: Convert the original image of the power communication construction drawing to be identified into a grayscale image. The grayscale image only contains brightness information and no color information, making subsequent processing more efficient. The grayscale conversion adopts the weighted average method.

[0046] S2.2.2: Use the adaptive threshold method to convert the grayscale image into a binary image to highlight the text area for subsequent text recognition. Different thresholds are calculated based on the local mean of each small block of the image. For example, in a small window, the average value of all pixels is used as the threshold of the area.

[0047] S2.2.3: Remove noise from the binarized image and retain useful text information. The denoising method may be a filter-based method.

[0048] S2.2.4: If the image is not horizontal or vertical, correct the tilt in the image so that the text in the image is horizontally aligned. The specific steps are as follows: (1) Use the Canny edge detection algorithm to determine the edges in the image by finding the gradient changes in the image, and obtain an edge image in which the edge positions are marked.

[0049] (2) Use Hough transform to detect straight lines in edge images. In image space, each edge point ( ) can be expressed as parameterized polar coordinate line equations:

[0050] in( is a point in image space, is the distance from the line to the origin, is the inclination angle of the line, which ranges from [0,180]. For each point in the image, traverse all possible , calculate the corresponding Value, count all ( ), when the number reaches a threshold (e.g. 10% of the total number of edge points), it is considered as a straight line in space, and all straight lines in the image are extracted.

[0051] (3) Calculate the inclination angle of the text area and obtain the inclination angles of all straight lines according to (2) , statistics all The most frequent The text tilt angle.

[0052] (4) Image rotation correction: Use the rotation matrix to rotate the image. The rotation matrix can rotate a point around the center point by a certain angle. The form of the rotation matrix is:

[0053] in Rotate the image by the rotation matrix to the angle you want to rotate. Angle returns the text area to a horizontal state.

[0054] S2.3: Input the training set into the constructed CRNN model for training, and construct the CRNN model loss function based on the output results of the CRNN model and the training labels.

[0055] The CRNN model is trained based on the preprocessed training set, and the validation set is used to evaluate the model performance during the training process. Specifically: The constructed loss function is used in S2.3 to optimize the model, which can handle the problem of inconsistent lengths of input and output sequences.

[0056]

[0057] in, is the loss value, is the target label sequence All possible alignments of, that is, all possible permutations of, the labels, is the time step The feature vector corresponding to the preprocessed construction drawing image is is the time step The predicted label, It's in time The predicted label is The probability of , T is the length of the input sequence, It is a target label sequence, which can be text, symbols, etc. of time in the construction drawings.

[0058] Combination Figure 3 As shown in Figure 2, the CRNN network structure constructed in S2.3 specifically includes CNN layer, feature fusion layer, dual-head self-attention mechanism, bidirectional LSTM and CTC layer; among them, The CNN layer is used to extract features from the image text lines of the power communication construction drawings to obtain feature maps. The hot evidence in the feature maps includes the shape of the characters, the edges of the characters, the texture features of the characters, the layout and structure between the characters, etc. The present invention uses four convolutional layers for feature extraction, and inputs the preprocessed drawings, whose size is (H, W, C) = (32, 150, 1). After four convolution operations, the sizes are (3, 3, 64), (3, 3, 128), (3, 3, 256), and (3, 3, 256), respectively.

[0059] The feature fusion layer connects multiple convolutional layers to fuse the feature maps extracted by multiple convolutional layers. Specifically, a downsampling branch is added on the basis of the convolutional layer to fuse the feature maps from different convolutional layers, which is expressed as the following formula:

[0060] in, Indicates The feature map after layer fusion, Indicates The feature map of the layer, represents a channel attention mechanism, , H and W represent the height and width of the feature map respectively, C is the number of channels, and the feature map selects channels that are more helpful for recognition through a channel attention mechanism SE-Layer. The channel refers to the depth dimension of the feature map output by the convolutional layer, thereby reducing the number of unnecessary channels during feature fusion. The maximum pooling operation is performed on the concatenated feature maps to reduce the spatial dimension of the feature maps. The specific steps of feature fusion are described below.

[0061] (1) Use the channel attention mechanism SE-Layer to process the Feature map of the layer , perform global average pooling on each channel to obtain the channel description vector.

[0062]

[0063] Where H and W represent the height and width of the feature map respectively. is the eigenvalue of the i-th row, j-th column, and c-th channel.

[0064] (2) The feature fusion layer includes a fully connected layer, which learns channel weights through a fully connected layer and a nonlinear activation function. The channel description vector is compressed, then passed through the ReLU activation function, and then through a fully connected layer To expand, and finally get the weight of each channel through the sigmoid activation function.

[0065]

[0066] in, is the weight vector of all channels, represents the weight of channel c, is the sigmoid function, and is the weight matrix of the fully connected layer, and z is the matrix describing all channels.

[0067] The channel attention mechanism learns the importance of each channel and dynamically adjusts the channel weights, thereby enhancing channel features that are helpful to the task and suppressing unimportant channel features.

[0068] (3) Combine the learned weights with the original feature map Multiply them together to get the weighted feature map.

[0069]

[0070] in represents the weight of channel c, Represents the characteristics of channel c.

[0071] (4) and Concatenate in the channel dimension and get .

[0072] (5) Perform the maximum pooling operation on the concatenated feature map to reduce the spatial dimension of the feature map and obtain the fused feature map .

[0073] The embodiment of the present invention can combine low-level detail information and high-level abstract information by fusing features from different convolutional layers, thereby improving the model's ability to understand the overall structure of the data.

[0074] Two-headed self-attention mechanism: The fused sequence is input into the two-headed attention mechanism, and the output of the two heads is weighted to obtain the final feature sequence; in this way, the two attention heads are used for parallel calculation to capture global dependencies and enhance the model's ability to model long-term dependencies on input data. This sequence is then fed into the two-headed self-attention mechanism as input. The features at each position will calculate the attention score through the query, key, and value. In the two-headed self-attention mechanism, important contextual information in the image is captured from different angles. The output of the two heads is weighted and summed as the final feature.

[0075] Bidirectional LSTM: The forward LSTM reads the forward information of the sequence, and the reverse LSTM reads the reverse information of the sequence. The outputs of the two are concatenated at each time step to obtain the feature vector of each time step, and finally the softmax probability distribution of all characters is output.

[0076] It can be understood that the bidirectional LSTM belongs to the RNN layer in the CRNN model.

[0077] CTC layer: The CTC layer aligns the sequence and calculates the loss according to the probability distribution of the characters, and extracts the final character sequence from the output of CTC through a decoding strategy.

[0078] The decoding strategy includes one or more of greedy decoding, beam set search decoding, prefix beam set search decoding, Viterbi decoding, and non-autoregressive decoding.

[0079] In the data iteration training process in S2.3, the preprocessed drawings and corresponding label data are input, the training rounds are set to 60 times, the initial learning rate is 0.001, the learning rate is adjusted dynamically, and the learning rate decays to half of the original every 6 rounds. The training batch size is 16.

[0080] S2.4: Use the validation set to judge the performance of the CRNN model on unseen data through the first evaluation indicator during the CRNN model training process, adjust the hyperparameters according to the performance, and obtain the optimal CRNN model; The first evaluation index includes a loss function and a character accuracy rate. The first evaluation index reflects the generalization ability of the model.

[0081] In this embodiment, the following objective function can be set according to the first evaluation index: , in, , Represent the weights of the CRNN model loss function and character accuracy, respectively. represents the CRNN model loss function, In this way, the CRNN model is trained by combining the loss function and the character accuracy to minimize the objective function and obtain the optimal CRNN model, thereby improving the recognition accuracy.

[0082] Specifically, the hyperparameters include learning rate, batch size, number of LSTM hidden units, etc. First, random sampling is performed in the hyperparameter space to obtain the objective functions corresponding to different hyperparameter combinations, and the objective functions corresponding to different hyperparameter combinations are sorted from small to large. Then, based on the Bayesian optimization algorithm, the models corresponding to the hyperparameter combinations corresponding to the first k objective functions are further fine-tuned, and the model corresponding to the hyperparameter combination corresponding to the minimum objective function is selected as the optimal CRNN model. In this way, the hyperparameters of the CRNN model are adjusted by combining random search and the Bayesian optimization algorithm to obtain the optimal CRNN model more quickly and accurately.

[0083] S2.5: Use the test set to evaluate the trained CRNN model, and evaluate the recognition accuracy of the CRNN model based on the second evaluation indicator.

[0084] The second evaluation index includes character accuracy and text accuracy. The second evaluation index is used to reflect the accuracy of model recognition.

[0085] Specifically, the test set is used to evaluate the performance of the trained model in order to further improve the algorithm, accuracy and generalization ability. Multiple indicators are used, including character accuracy based on edit distance. , text accuracy .

[0086]

[0087] in, Predicted text and real text The edit distance between Represents real text The number of characters, N is the total number of texts.

[0088]

[0089] in, is the predicted text, is the real text, and N is the total number of texts.

[0090] Specifically, when the character accuracy is over 95% but the text accuracy is less than 60%, it indicates that the model is prone to errors in certain characters. You can increase the training samples of specific characters and improve the diversity of training data. For example, you can use data augmentation technology to increase the diversity of training data and increase the training weights of these error-prone character samples. When the character accuracy is less than 60%, increase the number of LSTM hidden units to improve the ability of sequence modeling.

[0091] In S2, post-processing is the optimization and proofreading of the recognition results, including removing redundant spaces, correcting typos, etc., to improve the quality of the generated supervisory instruction data. Based on the CRNN model, the extracted text is converted into instruction data, and the instruction data is in the following form: [{"instruction":"Extract the cable naming service information from the text given below and output it in a specific format.", "input": " 1. This project is designed according to the access system and preliminary design review opinions, relevant meeting spirit, etc. The design scope of this volume includes the relevant contents of optical fiber communication of photovoltaic power stations. The installed capacity of the starting photovoltaic power station is 100MW, and it is connected to the 110kV side of the target substation through a 110kV grid-connected line with a line length of 18km. 2. System communication plan 2.1 Optical fiber communication plan (1) Optical cable construction plan Along the 110kV line from the starting photovoltaic power station to the target substation, two 48-core OPGW optical cables are installed, with a fiber core of G.652D and a cable length of 2*18km. ……", "output":{ "A-side site": "Starting photovoltaic power station", "Z-end site": "target substation" "Primary line voltage level": "110kV", "Number of cores": "48 cores", "Optical cable type": "OPGW optical cable", "Core Model": "G.652D", "cable length": "18km"} } ] S3: Fine-tune the large model based on the supervised instruction data to obtain a fine-tuned large model.

[0092] The process of model fine-tuning involves adjusting the parameters of the model to make it better suited to the needs of a specific task. This process can not only improve the performance of the model on specific tasks, but also help the model understand and generate more accurate text content.

[0093] Optionally, the large model is fine-tuned using an improved LoRA (Low-Rank Adaptation) technique. The improvement of LoRA is that during the fine-tuning process, sub-matrices of matrix A and matrix B are used to perform Hadamard products, a parameter matrix is ​​obtained based on the Hadamard product, and the number of Hadamard products is dynamically adjusted.

[0094] Specifically, the improved LoRA technology is used to fine-tune the large model. The process is as follows: S3.1: For the pre-trained weight matrix , LoRA limits its update method, that is, the incremental parameter matrix of all parameter fine-tuning Represented as two matrices with smaller parameters and A low-rank approximation of :

[0095] in, or Represents the pre-trained weight matrix, whose dimension is , this matrix represents the weights in the original model. and They are the two low-rank matrices introduced in LoRA fine-tuning. and ,rank Much smaller than .

[0096] S3.2: During training, initialize the low-rank adaptation matrices B and A: The matrix A is initialized by a Gaussian function, , the matrix B is initialized to all 0s, .

[0097] S3.3: Dynamically adjust the number of Hadamard products, starting from the initial B and A, and use their sub-matrices to perform Hadamard products. In order to make the number of Hadamard products more flexible, introduce a step size s, for example s=2, r=8.

[0098] Matrix B is sliced ​​by columns to obtain submatrices. When r / s is an integer, each submatrix contains s columns, and r / s submatrices are obtained. When r / s is not an integer, r / s+1 submatrices are obtained. The number of columns of the last submatrix is ​​less than s. The number of submatrices is recorded as N. Similarly, matrix A is sliced ​​by rows to obtain the same number of submatrices.

[0099] Perform matrix product of the corresponding submatrices of B and A to obtain:

[0100] in, Indicates the sub-matrices, Indicates the sub-matrices, , N represents the number of sub-matrices. Perform the Hadamard product in sequence to obtain the final parameter matrix:

[0101] S3.4: Given input , output after adding LoRA :

[0102] in, represents the pre-trained weight matrix, The parameter matrix representing the parameter update during fine-tuning, Denotes the Hadamard product.

[0103] S4: Use the fine-tuned large model to extract key information from the text recognized by the CRNN model.

[0104] Specifically, when a new power communication construction drawing is received, the text information in the drawing needs to be recognized by the CRNN model first. The recognized text is then passed as input to the fine-tuned large model. The large model understands the user's intention based on the prompt template and extracts key information from it. The required text output is generated based on the extracted information. This process not only improves the efficiency of information processing, but also ensures that the generated text content is accurate and meets actual needs.

[0105] Specifically, the prompt template can be: Please extract the optical cable naming service information from the following text, including the A-end site, the Z-end site, the primary line voltage level, the number of cores, the optical cable type, the fiber core model, and the optical cable length. If no specific information is found in the provided text, the value of the information should be set to None and output in JSON format.

[0106] Specifically, the key information of the power communication construction drawings includes construction project information and business information. The construction project information includes the construction project name and the construction project address. The business information includes optical cable naming, networking information, automation or scheduling IAD information, protection business information, telephone number and construction site information.

[0107] Specifically, the key information is explained as follows: (1) Optical cable naming includes basic information such as the type, number of cores, length, name, and the names of the starting and ending sites. This part is crucial to the configuration and management of optical cables, ensuring that the identification of the optical cable is consistent with the actual physical location.

[0108] (2) Networking information involves the overall structure of the network, including the names of the starting and ending sites, the usage of the equipment at the starting and ending sites, the transmission system, the usage of optical cables, etc., to help design and implement network connections.

[0109] (3) Automation / Dispatching IAD information mainly includes basic information such as the transmission system, the names of the starting and ending sites, the type of transmission channel, and the usage of intermediate channel resources.

[0110] (4) Protection business information involves protection measures in the power communication system, mainly including protection device model, transmission system, single-port / dual-port, transmission channel type, etc. This information is very important for ensuring the safety and reliability of the system.

[0111] (5) The telephone number shall include the starting and ending site names, names, application instructions and other information.

[0112] S5: Convert the key information extracted from the large model into structured data, and store the converted key information and power communication construction drawings in a set format in the database.

[0113] Specifically, power communication construction drawings usually exist in image format. To ensure the integrity and accessibility of these files, the drawing files are first stored in the file system. Each file should be assigned a unique file name and path to facilitate subsequent retrieval and management. For the extracted structured data, a table is directly created to store it.

[0114] By converting the unstructured text information output by the large model into structured data, it can be easily stored and managed. The power communication construction drawings and the extracted structured data can be stored in the database efficiently and securely.

[0115] Combination Figure 1 As shown, embodiment 2 of the present invention provides a large model-based power communication construction drawing information extraction system, which runs the method described in embodiment 1, and the system includes: A data acquisition module, used to obtain images of power communication construction drawings to be identified; A drawing recognition module is used to use the constructed and trained CRNN model to perform text recognition on images of power communication construction drawings and post-process the recognized text, and then convert the post-processed text into supervision instruction data; A fine-tuning module is used to fine-tune the large model based on the supervision instruction data to obtain a fine-tuned large model; The large model information extraction module is used to extract key information from the text recognized by the CRNN model using the fine-tuned large model.

[0116] Optionally, the system also includes an output module and a data storage module, the output module is used to convert the key information output by the large model information extraction module into structured data, and the data storage module is used to store the structured data and power communication construction drawings in a set format in a database.

[0117] Optionally, the system further includes a preprocessing module, which preprocesses the original image of the power communication construction drawing, including part or all of binarization, grayscale, denoising, and rotation correction.

[0118] Embodiment 3 of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the method for extracting information from electric power communication construction drawings based on a large model as described in Embodiment 1 is implemented.

[0119] Embodiment 4 of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for extracting information from power communication construction drawings based on a large model according to Embodiment 1 is implemented.

[0120] Compared with the prior art, the beneficial effects of the present invention include at least: The present invention uses visual technology combined with large model technology to automatically and accurately extract information from power communication construction drawings, reduce manual recognition errors, and improve recognition accuracy. The automated recognition process not only significantly shortens construction preparation time and improves construction efficiency, but also reduces dependence on professionals and reduces labor costs.

[0121] The present invention adds a downsampling branch on the basis of the convolution layer to fuse the features from different convolution layers, dynamically adjusts the learning rate, adds a dual-head self-attention mechanism before BiLSTM, uses more stringent evaluation indicators to evaluate the drawing recognition effect, and uses improved LoRA to fine-tune the large model; the present invention improves the feature extraction and multi-scale perception capabilities, so that the model can more comprehensively capture the low-level and high-level features of the image and adapt to complex drawing recognition tasks. The LoRA improved fine-tuning algorithm enhances the expression ability of the model, and the step size mechanism improves the flexibility of training and helps the model converge more efficiently.

[0122] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0123] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0124] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the above. The computer-readable storage medium used herein is not to be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0125] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0126] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for extracting information from power communication construction drawings based on a large model, characterized in that: include: S1: Acquire the image of the power communication construction drawing to be identified and perform preprocessing; S2: Use the constructed and trained CRNN model to perform text recognition on the construction drawing images preprocessed in S1 and post-process the recognized text, and then convert the post-processed text into supervision instruction data; S3: Fine-tune the large model based on the supervised instruction data to obtain a fine-tuned large model; S4: Use the fine-tuned large model to extract key information from the text recognized by the CRNN model.

2. The method for extracting information from power communication construction drawings based on a large model according to claim 1 is characterized in that: In S1, the preprocessing includes part or all of binarization, grayscale, denoising, and rotation correction.

3. The method for extracting information from power communication construction drawings based on a large model according to claim 1 is characterized in that: In S2, training the CRNN model includes: Acquire various types of power communication construction drawing images, annotate each character in each power communication construction drawing image, and obtain original power communication construction drawing images and annotated power communication construction drawing images; Performing image preprocessing on the original power communication construction drawing images and the annotated power communication construction drawing images to form a data set, and dividing the data set into a training set, a test set, and a validation set according to a preset ratio; Input the training set into the constructed CRNN model for training, and construct the CRNN model loss function based on the output result of the CRNN model and the training label; Use the validation set to judge the performance of the CRNN model on unseen data through the first evaluation indicator during the CRNN model training process, adjust the hyperparameters according to the performance, and obtain the optimal CRNN model; The test set is used to evaluate the recognition accuracy of the trained CRNN model based on the second evaluation indicator.

4. The method for extracting information from power communication construction drawings based on a large model according to claim 3 is characterized in that: Image preprocessing of power communication construction drawings, including: The original image of the electric power communication construction drawing is converted into a grayscale image, and the grayscale image is converted into a binary image and then the noise in the binary image is removed; The straight line in the denoised binary image is identified using an edge algorithm, the straight line parameters in the denoised binary image are extracted using Hough transform, and the tilt angle of the original image of the power communication construction drawing is estimated based on the straight line parameters; The original image of the electric power communication construction drawing is rotationally corrected according to the estimated tilt angle so that the text lines in the electric power communication construction drawing are aligned.

5. The method for extracting information from power communication construction drawings based on a large model according to claim 3 is characterized in that: Construct the CRNN model loss function according to the following formula: in, is the loss value, is the target label sequence All permutations of is the time step The feature vector corresponding to the preprocessed construction drawing image on is the time step The predicted label, For in time The predicted label is The probability of , T is the length of the input sequence, is the target label sequence.

6. The method for extracting information from power communication construction drawings based on a large model according to claim 3 is characterized in that: During the CRNN model training process, the validation set is used to judge the performance of the CRNN model on unseen data through the first evaluation indicator, and the hyperparameters are adjusted according to the performance to obtain the optimal CRNN model, including: Constructing an objective function according to the first evaluation indicator and its corresponding weight; Random sampling is performed in the hyperparameter space to construct multiple CRNN models. The validation set is used to train multiple CRNN models to obtain the objective functions corresponding to different hyperparameter combinations, and the objective functions corresponding to different hyperparameter combinations are sorted from small to large. Based on the Bayesian optimization algorithm, the models corresponding to the hyperparameter combinations corresponding to the top k objective functions are trained, and the model corresponding to the hyperparameter combination corresponding to the minimum objective function is selected as the optimal CRNN model.

7. The method for extracting information from power communication construction drawings based on a large model according to claim 1 is characterized in that: The constructed CRNN model includes CNN layer, feature fusion layer, dual-head self-attention mechanism, bidirectional LSTM and CTC layer; The CNN layer is used to extract features from the image text lines of the power communication construction drawings to obtain a feature map, wherein the features in the feature map include the shape of the characters, the edges of the characters, the texture of the characters, the layout between the characters, and the structure of the characters; the CNN layer includes a plurality of convolutional layers; The feature fusion layer connects multiple convolutional layers and is used to fuse the feature maps extracted by the multiple convolutional layers; The fused feature map is converted into a feature sequence and then input into the dual-head attention mechanism. The output of the two heads is weighted summed to obtain the final feature sequence. The forward LSTM in the bidirectional LSTM reads the forward information of the final feature sequence, and the reverse LSTM reads the reverse information of the final feature sequence. The outputs of the forward LSTM and the reverse LSTM are concatenated at each time step to obtain the feature vector of each time step and output the probability distribution of all characters. The CTC layer aligns the sequence according to the probability distribution of characters and calculates the loss, and the final character sequence is extracted from the output of the CTC layer through a decoding strategy.

8. The method for extracting information from power communication construction drawings based on a large model according to claim 7 is characterized in that: The feature maps extracted from multiple convolutional layers are fused in the feature fusion layer according to the following formula: in, Indicates The feature map after layer fusion, Indicates The feature map of the layer, represents a channel attention mechanism, For the general The processed feature map is concatenated with the feature map of the previous layer. Indicates the maximum pooling operation on the concatenated feature map. , H and W represent the height and width of the feature map respectively, and C is the number of channels; Among them, the channel attention mechanism learns the importance of each channel and dynamically adjusts the weight of the channel. Feature map of the layer Multiply to obtain the weighted feature map .

9. The method for extracting information from power communication construction drawings based on a large model according to claim 8 is characterized in that: The feature fusion layer includes multiple fully connected layers. The channel attention mechanism dynamically adjusts the weight of each channel by learning the importance of each channel, specifically including: Use channel attention mechanism to process the feature map of layer i , perform global average pooling on each channel to obtain a channel description vector; The channel description vector is compressed through the first fully connected layer, and then the compressed channel description vector is input into the ReLU activation function; The output result of the ReLU activation function is input into the second fully connected layer, and the output of the second fully connected layer is passed through the sigmoid activation function to obtain the weight of each channel.

10. The method for extracting information from power communication construction drawings based on a large model according to claim 1, characterized in that: The improved LoRA technology is used to fine-tune the large model. During the fine-tuning process, the sub-matrices of matrix A and matrix B are used to perform Hadamard products. The parameter matrix is ​​obtained based on the Hadamard product, and the number of Hadamard products is dynamically adjusted.

11. The method for extracting information from power communication construction drawings based on a large model according to claim 10, characterized in that: During the fine-tuning process, the sub-matrices of matrix A and matrix B are used to perform Hadamard products, the parameter matrix is ​​obtained based on the Hadamard product, and the number of Hadamard products is dynamically adjusted, including: Slice the matrix B by columns to obtain submatrices of B. When r / s is an integer, each submatrix of B contains s columns, and r / s submatrices of B are obtained. When r / s is not an integer, r / s+1 submatrices of B are obtained. Slice the matrix A by rows to obtain submatrices of A. When r / s is an integer, each submatrix of A contains s rows, and r / s submatrices of A are obtained. When r / s is not an integer, r / s+1 submatrices of A are obtained, where r is the rank of the A matrix and the B matrix, and s is the step size. Perform matrix product of the corresponding sub-matrices of B and A: in, , Indicates the sub-matrices, Indicates the sub-matrices, N represents the number of sub-matrices; N Perform the Hadamard product in order to obtain the parameter matrix: in, The parameter matrix representing the parameter update during fine-tuning, Denotes the Hadamard product.

12. The method for extracting information from power communication construction drawings based on a large model according to any one of claims 1 to 11, characterized in that: The method further comprises: The key information extracted from the large model is converted into structured data, and the converted key information and power communication construction drawings are stored in the database in a set format.

13. A large-model-based power communication construction drawing information extraction system, characterized in that: include: A data acquisition module, used to obtain images of power communication construction drawings to be identified; A drawing recognition module is used to use the constructed and trained CRNN model to perform text recognition on images of power communication construction drawings and post-process the recognized text, and then convert the post-processed text into supervision instruction data; A fine-tuning module is used to fine-tune the large model based on the supervision instruction data to obtain a fine-tuned large model; The large model information extraction module is used to extract key information from the text recognized by the CRNN model using the fine-tuned large model.

14. The large model-based power communication construction drawing information extraction system according to claim 13, characterized in that: The system also includes an output module and a data storage module. The output module is used to convert the key information output by the large model information extraction module into structured data, and the data storage module is used to store the structured data and power communication construction drawings in a set format in a database.

15. An electronic device, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method for extracting information from power communication construction drawings based on a large model according to any one of claims 1-12.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for extracting information from power communication construction drawings based on a large model as described in any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Weak supervision electric power drawing OCR identification method based on deep learning

    CN111860348A

  • Drawing information extraction method and device, equipment and storage medium

    CN118015648A

  • Text recognition method and device based on deep learning, equipment and storage medium

    CN113569608A

  • Transformer substation secondary drawing identification method and system and transformer substation secondary drawing retrieval method and system

    CN115223189A

  • Large model fine tuning method, device and equipment based on Hadamard product and medium

    CN117910599A