Image preprocessing method for recording archives of power supply system
By conducting quality assessment and correction of photos entered into the power supply system archives, and combining deep learning models for text recognition and semantic segmentation, the problems of inconsistent photo quality and numerous interfering factors were solved. This enabled efficient extraction and storage of key information, improving the automation and accuracy of archive entry.
Patent Information
- Application Number
- CN202511635482.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-01-09
AI Technical Summary
During the process of entering new archives into the power supply system, the inconsistent quality of photos, the diversity of content, and the numerous interfering factors make it difficult for image preprocessing technology to accurately locate and segment the target area, thus affecting the accuracy and efficiency of information extraction.
An image quality assessment model is used for brightness adjustment, deblurring, and tilt correction. A deep learning model is combined for text detection and recognition. A semantic segmentation network is used to separate the foreground and background. Key information locations are determined through feature extraction and matching. Data is then normalized and stored.
It improves the automation and accuracy of photo file entry, ensuring a high-quality data foundation for subsequent information entry and supporting image retrieval and analysis.
Smart Images

Figure CN121305579A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power supply, and more specifically to an image preprocessing method for power supply system file entry. Background Technology
[0002] Images used for new power supply system record entry refer to image data used for the automated creation and updating of power equipment records under the background of new technologies such as AI recognition, mobile acquisition, and digital twins. Compared with traditional paper record scanning, the new record entry places greater emphasis on image recognition capabilities, real-time performance, rich metadata, and system compatibility.
[0003] In the image preprocessing process of new photo archives, extraction technology faces several technical challenges. Differences in the shooting environment, angle, and lighting of the original photos result in varying quality. Some photos may suffer from overexposure, underexposure, blurring, or tilting, making subsequent information extraction difficult. Furthermore, text in photos may exhibit inconsistent fonts, sizes, and orientations, increasing the complexity of text recognition. Simultaneously, different types of photos differ significantly in content, each with distinct characteristics in composition, content, and layout. This necessitates that extraction technology adapt to the characteristics of different photo types, flexibly adjusting extraction strategies and parameters to achieve optimal extraction results. Another issue is the presence of interfering information in the photos, such as backgrounds, watermarks, and stains, which can affect the location and extraction of key information. Accurately locating and segmenting the target area against complex backgrounds is a pressing technical challenge that needs to be addressed.
[0004] In summary, image preprocessing technology in the new photo archive entry faces challenges such as inconsistent photo quality, diverse content, and numerous interfering factors. In-depth research is needed in areas such as algorithm models, strategy control, and parameter optimization to continuously improve the accuracy and efficiency of extraction, thereby providing a high-quality data foundation for subsequent information entry. Summary of the Invention
[0005] The purpose of this invention is to solve the above-mentioned problems and provide an image preprocessing method for power supply system file entry.
[0006] The technical solution adopted by this invention to solve its technical problem is: An image preprocessing method for power supply system file entry includes the following steps: S101 Acquire the original photo, determine whether the photo has quality problems through the image quality assessment model, improve the quality of the problematic photo, and obtain the improved photo; S102 locates the text area in the improved photo, corrects the text orientation, and then recognizes the text content through optical character recognition technology to obtain the text-corrected photo. S103 determines the type of the text-corrected photo based on the photo type classification model, and adjusts the feature extraction parameters and region segmentation strategy for different types of photos. S104 Separates the foreground and background in the photo, removes watermarks and stains, retains the target area, and obtains the segmented target area; S105 performs feature extraction on the segmented target region and compares it with a pre-established photo type template to determine the location of key information; S106 Based on the determined key information location, extract the target region, and accurately locate the information boundary through edge detection and contour analysis to obtain the extracted target region; S107 Normalizes the extracted target region to build a data foundation; S108 stores the normalized photos and extracted information into the archive database, establishing indexes and relationships for subsequent image retrieval and analysis.
[0007] Further, step S101 includes: Obtain the original photo and input it into a pre-built image quality assessment model to determine whether the original photo has quality problems such as overexposure, underexposure, blurring, and tilting. If the original photo has overexposure or underexposure issues, an adaptive histogram equalization algorithm is used to adjust the brightness of the original photo to obtain a brightness-adjusted photo. If the original photo has a blur problem, the Laplacian variance of the original photo is calculated by the Laplacian operator to determine the blurry area, and a deblurring algorithm is used to process it to obtain the deblurred photo; If the original photo has a tilt problem, the tilt angle of the original photo is detected by Hough transform, and the original photo is rotated to correct the tilt, resulting in a rotated corrected photo. The brightness-adjusted photo, the deblurred photo, and the rotation-corrected photo are fused together to obtain a photo with improved quality.
[0008] Further, step S102 includes: After acquiring the improved image quality, a convolutional neural network model is used to extract image features, resulting in an image feature vector. Based on the image feature vector, the CTPN model is used to detect and locate the text region in the image, and the coordinates of the text region are obtained. For the coordinates of the text region, a direction classification model is used to determine the text direction. If the text direction is incorrect, direction correction is performed to obtain the text region image after direction correction. The image of the text region after the orientation correction is subjected to image enhancement processing. The enhanced text region image is then input into the CRNN model for text recognition to obtain the recognition result. The recognition results are processed using a language model to correct errors, resulting in the final text recognition result. The recognized text content is then embedded into the corresponding area of the original photo, and the text-corrected photo is output.
[0009] Further, step S103 includes: The target photo after text correction is obtained, and the target photo is determined to belong to a certain category based on a pre-established photo type classification model. Based on the type of the target photo, a corresponding region segmentation strategy is adopted to divide the target photo into different regions; For different regions after segmentation, extract the corresponding feature descriptors to obtain facial feature vectors and text feature vectors or main body feature vectors; The extracted feature vectors are combined into a comprehensive feature vector, which serves as the feature representation of the target photo.
[0010] Further, step S104 includes: The target image to be processed is obtained and input into a pre-trained semantic segmentation network model for foreground-background separation to obtain foreground and background masks. For the foreground mask, morphological opening operation is used to remove the small watermark area, and morphological closing operation is used to fill the stain holes in the target area to obtain the foreground mask after removing the watermark and stains. Based on the foreground mask after removing watermarks and stains, extract the corresponding foreground target region image from the target photo; For the foreground target region image, its brightness, contrast, and color are adjusted using an image enhancement algorithm to obtain an enhanced foreground target region image; The enhanced foreground target area image is fused with the background area of the target photo to obtain a fused photo; An image restoration algorithm is used to optimize the fused photo, eliminating the abruptness of the foreground target edges, and obtaining the restored target photo.
[0011] Further, step S105 includes: Obtain pre-established photo type templates, extract image feature vectors for each photo type, and build a feature template library; The image to be processed is acquired, and the image is segmented using an image segmentation algorithm to obtain several target regions; For each target region, the BIC algorithm is used to extract the image feature vector of that region; The feature vector of the target region is matched with the feature vector in the feature template library, and the similarity is calculated. Based on the similarity, determine which photo type the target area belongs to, and obtain the location of key information corresponding to that photo type; If the similarity between the target area and a certain photo type template exceeds a preset threshold, the area is marked as that photo type, and the key information location coordinates are recorded. Based on the type label and key information location coordinates of each target region, the location of key information in the image to be processed is determined.
[0012] Further, step S106 includes: Based on the pre-established key information location model, the location of key information in the image to be processed is determined, and the coordinates of the key information location are obtained; Based on the obtained key information location coordinates, an adaptive threshold segmentation algorithm is used to segment the image to be processed, resulting in a segmented target region image. Edge detection processing is performed on the segmented target region image to extract the edge information of the target region and obtain the edge coordinates of the target region; Based on the obtained edge coordinates of the target area, contour analysis is performed to accurately locate the boundary position of the target area and obtain the precise boundary coordinates of the target area. Determine whether the boundary coordinates of the target area match the location coordinates of the key information; if they match, obtain the extracted target area image. If there is a mismatch, the parameters of the adaptive threshold segmentation algorithm are adjusted, and the target region is extracted again until the extracted target region is accurate.
[0013] Further, step S107 includes: Extract the target region image based on the preset target region coordinates; The target region image is normalized to unify the image size and resolution; Based on preset data augmentation parameters, transformation parameters are randomly generated; The normalized target region image is transformed according to the transformation parameters to obtain the transformed target region image. Random noise is added to the transformed target region image according to the preset noise type and noise intensity to obtain the target region image with added noise. The target region image and multiple target region images with added noise after data augmentation are combined to construct a target region image dataset.
[0014] Further, step S108 includes: The normalized photos and key information are packaged according to a preset data format to obtain structured archive data; Based on the preset database table structure, the structured archive data is stored in the archive database, an image feature index is created for the normalized photos, and a text index is created for the key information.
[0015] The beneficial effects of this invention are: This invention first performs quality assessment and correction on the original photograph, including brightness adjustment, deblurring, and tilt correction. Then, text detection and recognition are performed, followed by feature extraction and region segmentation based on the photograph type. Next, a semantic segmentation network is used to separate the foreground and background, removing watermarks and stains. Feature extraction and matching are performed on the segmented target regions to determine the location of key information. Adaptive threshold segmentation and edge detection are used to accurately locate information boundaries. Finally, the extracted target regions are normalized and data augmented, and the processed photographs and information are stored in an archival database. This invention solves the problems of image quality, text recognition, photograph classification, and key information extraction in the photographic archiving process, improving the automation and accuracy of photographic archiving and laying the foundation for subsequent image retrieval and analysis. Attached Figure Description
[0016] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0017] like Figure 1 As shown, an image preprocessing method for power supply system file entry includes the following steps: S101 acquires the original photo, uses an image quality assessment model to determine whether the photo has overexposure, underexposure, blur, or tilt issues, and for the photo with quality issues, uses an adaptive histogram equalization algorithm to adjust the brightness of the overexposed or underexposed photo, uses Laplacian variance to detect the blurred area and applies a deblurring algorithm, uses Hough transform to detect the tilt angle and performs rotation correction to obtain the corrected photo. S102 For the corrected photo, a deep learning-based text detection model is used to locate the text region, correct the text direction, and then the text content is recognized through optical character recognition technology to obtain the text-corrected photo. S103 determines whether the text-corrected photo belongs to a front-facing portrait, ID photo, or lifestyle photo based on the photo type classification model. For different types of photos, the feature extraction parameters and region segmentation strategies are adjusted. Specifically, for front-facing portraits, the focus is on extracting facial features; for ID photos, the focus is on text information and standardized layout; and for lifestyle photos, the focus is on scene and subject recognition. S104 uses a semantic segmentation network to separate the foreground and background in the photo, removes watermarks and stains through morphological operations, and retains the target region to obtain the segmented target region. S105 performs feature extraction on the segmented target region, uses BIC to extract image features, and compares them with a pre-established photo type template through a feature matching algorithm to determine the location of key information; S106 Based on the determined key information location, an adaptive threshold segmentation algorithm is used to extract the target region, and the information boundary is accurately located through edge detection and contour analysis to obtain the extracted target region; S107 normalizes the extracted target region to unify the image size and resolution, and expands the sample size by using data augmentation techniques such as rotation, translation, scaling, flipping, and adding noise to build a data foundation. S108 stores the normalized photos and extracted information into the archive database, establishing indexes and relationships for subsequent image retrieval and analysis.
[0018] Step S101 includes: The original photo is acquired and input into a pre-built image quality assessment model to determine if it has overexposure, underexposure, blur, or tilt issues. The image quality assessment model is trained using machine learning algorithms and can automatically identify various quality problems in the photo. Overexposure manifests as an overall washed-out image and loss of detail in highlight areas; underexposure results in an overall dark image and blurred details in shadow areas; blur is commonly caused by motion blur or out-of-focus issues, making image edges and textures unclear. Tilting issues manifest as tilted horizontal or vertical lines in the image.
[0019] If the original photo is overexposed or underexposed, an adaptive histogram equalization algorithm is used to adjust the brightness of the original photo, resulting in a brightness-adjusted photo. The adaptive histogram equalization algorithm effectively improves the brightness and contrast of the photo. This algorithm divides the image into several small blocks, performs histogram equalization on each block separately, and then stitches the processed blocks together. This method can improve the overall contrast while avoiding noise amplification caused by over-enhancement.
[0020] If the original photo is blurry, the Laplacian variance of the original photo is calculated using the Laplacian operator to determine the blurry areas. A preset deep learning-based deblurring algorithm, such as a convolutional neural network, is then used to process these blurry areas to obtain the deblurred photo. The Laplacian operator is used to measure image sharpness. By calculating the Laplacian variance, blurry areas in the photo can be located.
[0021] If the original photo has a tilt problem, the tilt angle of the original photo is detected by Hough transform, and the original photo is rotated and corrected according to the tilt angle to obtain a rotated and corrected photo.
[0022] The brightness-adjusted photo, the deblurred photo, and the rotation-corrected photo are merged to obtain a photo with improved quality. Image fusion is the process of combining multiple photos that have undergone different processing into a single final photo. In this process, methods such as weighted averaging or multi-resolution analysis can be used to retain the best parts of each photo.
[0023] The image quality of the improved photo is re-evaluated to determine whether the improved photo meets a preset image quality threshold. If the quality of the improved photo reaches the image quality threshold, then the improved photo will be used as the final output result. Otherwise, return to the step of determining whether the original photo has quality problems such as overexposure, underexposure, blurring, and tilt, and continue to improve the quality of the original photo until the quality requirements are met.
[0024] Image quality reassessment can be achieved by setting objective metrics such as Peak Signal-to-Noise Ratio (PSNR) or Structural Similarity (SSIM) to quantify image quality. For example, the processed image might be required to have a PSNR of at least 30 dB and an SSIM of at least 0.9. If these thresholds are not met, quality improvement processing needs to be repeated, which may require adjusting algorithm parameters or trying other processing methods. This iterative image quality improvement process can effectively improve the overall quality of photos. By comprehensively utilizing multiple image processing techniques, different types of image quality problems can be solved, ultimately resulting in clear, bright, and correctly composed high-quality photos.
[0025] Step S102 includes: acquiring the corrected photo, extracting image features using a convolutional neural network model, which can use pre-trained deep learning models such as ResNet or VGGNet, and extracting high-dimensional feature vectors from the image through multi-layer convolution and pooling operations.
[0026] Based on the image feature vector, the CTPN model is used to detect and locate text regions in the image and obtain the coordinates of the text regions. The CTPN model combines convolutional neural networks and recurrent neural networks, which can effectively detect text lines in various complex backgrounds and output their coordinate information.
[0027] For the coordinates of the text region, an orientation classification model is used to determine the text orientation, such as a CNN classifier, which divides the text region image into four categories: 0°, 90°, 180° and 270°. If the text orientation is incorrect, orientation correction is performed to obtain the text region image after orientation correction.
[0028] Image enhancement processing is performed on the text region image after orientation correction, including contrast adjustment, sharpening and other operations, to improve the clarity and readability of the text. The enhanced text region image is then input into the CRNN model for text recognition. The CRNN model combines the feature extraction capability of CNN and the sequence modeling capability of RNN, and can effectively recognize text sequences of variable length to obtain recognition results.
[0029] The recognition results are then processed using a language model for error correction. This can be achieved using an N-gram model or a more advanced pre-trained language model such as BERT, which corrects potential recognition errors based on contextual information to obtain the final text recognition result. Based on this final result, the recognized text content is embedded into the corresponding region of the original image. This requires precisely locating and replacing the original text in the original image using the previously obtained text region coordinates, resulting in a text-corrected image. This series of steps constitutes a complete text recognition and correction process, effectively improving the accuracy and readability of text in images.
[0030] Step S103 includes: acquiring the target photo after text correction, and determining which category the target photo belongs to based on a pre-established photo type classification model; in practical applications, problems such as uneven lighting and occlusion may be encountered, and at this time, a multi-scale pyramid strategy can be combined to improve the detection accuracy.
[0031] If the target photo is a frontal portrait, a facial key point localization algorithm is used to obtain the position coordinates of the facial area and extract facial texture and contour features. If the target photo is an ID photo, optical character recognition technology is used to extract the text content in the target photo, and the target photo is judged to meet the standardized layout requirements according to a preset ID photo layout template. If the target photo is a work photo, object detection and image segmentation algorithms are used to identify the main scene and subject in the target photo, and the scene classification model is used to determine the scene category to which the target photo belongs.
[0032] Based on the type of the target photo, an appropriate region segmentation strategy is adopted to divide the target photo into different regions. For each segmented region, corresponding feature descriptors are extracted to obtain facial feature vectors and text feature vectors or subject feature vectors. The extracted feature vectors are combined into a comprehensive feature vector as the feature representation of the target photo. Facial features can use Local Binary Pattern (LBP) or Gabor filters to capture texture information, combined with shape context descriptors to represent contour features. Text features can use the TF-IDF (Term Frequency-Inverse Document Frequency) method to convert text into numerical vectors. Subject features can utilize pre-trained deep convolutional networks, such as VGG or Inception, to extract high-dimensional semantic features. The construction of the comprehensive feature vector needs to consider the scale and importance of different features. Feature selection techniques, such as Principal Component Analysis (PCA) for dimensionality reduction, or attention mechanisms can be used to adaptively adjust the weights of each feature.
[0033] The comprehensive feature vector of the target photo is input into a pre-trained photo analysis model to obtain the high-level semantic information and attribute labels of the target photo. This multimodal feature fusion method can effectively improve the accuracy and robustness of photo analysis.
[0034] Step S104 includes: acquiring the target photo to be processed and inputting it into a pre-trained semantic segmentation network model for foreground-background separation. The semantic segmentation network model is a deep learning technique used to classify each pixel in an image into a specific category. In photo processing, this model can effectively separate the foreground object from the background, obtaining foreground and background masks.
[0035] For the foreground mask, morphological opening operation is used to remove small watermark areas, and morphological closing operation is used to fill the dirt holes in the target area, resulting in a foreground mask after removing watermarks and dirt. Opening operation can remove small noise areas, such as watermarks; closing operation is used to repair the holes caused by removing watermarks, making the portrait more complete.
[0036] Based on the foreground mask after removing watermarks and stains, the corresponding foreground target area image is extracted from the target photo.
[0037] For the foreground target area image, its brightness, contrast, and color are adjusted using an image enhancement algorithm to obtain an enhanced foreground target area image. For an indoor photo with insufficient light, increasing the brightness can make the facial features of the person clearer, increasing the contrast can enhance the sense of depth in the image, and adjusting the color can make the skin tone more natural.
[0038] By fusing the enhanced foreground target area image with the background area of the target photo, the authenticity of the background can be preserved while highlighting the subject, resulting in a fused photo.
[0039] Image inpainting algorithms are used to optimize the fused photo, eliminating abrupt edges of the foreground object to obtain the restored target photo. For a portrait photo that has undergone foreground extraction and background replacement, the inpainting algorithm can eliminate unnatural edges around the subject's outline, making the subject appear more integrated into the background. The purpose of this series of processing steps is to improve the overall quality and visual appeal of the photo. The restored target photo is then output.
[0040] Foreground-background separation allows for targeted processing of both the subject and background; watermark and stain removal cleans up interfering elements in the image; image enhancement improves the visual effect of the subject; and fusion and restoration ensure a natural and harmonious final image. This comprehensive processing method can be widely applied in various fields such as personal photo enhancement and commercial product display image processing, effectively improving image quality and visual expressiveness.
[0041] Step S105 includes: obtaining a pre-established photo type template, which is a pre-established set of image features used to identify and classify different types of photos; extracting the image feature vector for each photo type; constructing a feature template library; and providing a reference benchmark for subsequent photo type identification.
[0042] The image to be processed is acquired, and an image segmentation algorithm is used to segment the image to obtain several target regions.
[0043] For each target region, the BIC algorithm is used to extract the image feature vector of that region. These feature vectors contain the color and texture information of the target region and can effectively characterize the visual features of that region.
[0044] The feature vector of the target region is matched with the feature vectors in the feature template library to calculate the similarity. Common similarity metrics include Euclidean distance and cosine similarity. If Euclidean distance is used, a smaller distance indicates a higher similarity.
[0045] Based on the similarity, the target area is determined to be of a certain photo type, and the key information location corresponding to that photo type is obtained. If the similarity between the target area and a certain photo type template exceeds a preset threshold, the area is marked as that photo type, and the key information location coordinates are recorded. By setting an appropriate similarity threshold, it is possible to determine whether the target area belongs to a specific photo type.
[0046] Based on the type label and key information location coordinates of each target region, the location of key information in the image to be processed is determined, and the result is output.
[0047] Photos may differ from standard templates due to factors such as shooting angle, lighting conditions, or image quality. To improve the robustness of recognition, a multi-template matching strategy can be employed, creating multiple variant templates for each photo type to adapt to different shooting conditions. This template-matching-based method for photo type recognition and key information localization can significantly improve the automation and efficiency of document processing. Simultaneously, this method provides an important foundation for subsequent OCR text recognition and information verification, contributing to the construction of more intelligent and secure information processing systems.
[0048] Step S106 includes: acquiring an image to be processed, wherein the image to be processed contains key information; determining the location of the key information in the image to be processed according to a pre-established key information location model, and acquiring the coordinates of the key information location; the pre-established key information location model is derived from statistical analysis of a large number of sample images, which can quickly locate the approximate area of the key information and improve the efficiency of subsequent processing.
[0049] Based on the obtained key information location coordinates, an adaptive threshold segmentation algorithm is used to segment the image to be processed, resulting in a segmented target region image. Compared with a fixed threshold, the adaptive threshold can dynamically adjust the segmentation threshold according to the brightness characteristics of the local area of the image, thus better handling images under different lighting conditions.
[0050] Edge detection processing is performed on the segmented target region image. Edge detection algorithms such as Sobel and Canny can be used to extract the edge information of the target region and obtain the edge coordinates of the target region, providing a basis for subsequent contour analysis.
[0051] Based on the obtained edge coordinates of the target area, contour analysis is performed. Contour tracking algorithms, such as Moore's neighborhood tracking, can be used to accurately locate the boundary position of the target area and obtain the precise boundary coordinates of the target area.
[0052] Determine whether the boundary coordinates of the target area match the location coordinates of the key information; if they match, obtain the extracted target area image. If a mismatch occurs, it may be due to image quality issues or improper algorithm parameter settings. In this case, adjust the parameters of the adaptive threshold segmentation algorithm and re-extract the target region, such as adjusting the local region size or contrast threshold, until the extracted target region is accurate. Then, perform subsequent information recognition processing on the extracted target region image to obtain key information within the target region.
[0053] This entire process is designed to improve the accuracy and robustness of key information extraction. Through meticulous processing in multiple steps, it can effectively handle various complex situations, such as poor image quality and uneven lighting, ultimately achieving efficient and accurate information extraction.
[0054] Step S107 includes: acquiring the original image, and extracting the target region image from the original image according to the preset target region coordinates; The target region image is normalized to obtain a normalized target region image. This normalization process includes standardizing the image size to a preset fixed size and the image resolution to a preset fixed value. This contributes to consistency in subsequent processing. This ensures that the data input to the model has the same format, improving the model's generalization ability.
[0055] Based on preset data augmentation parameters, transformation parameters are randomly generated, including rotation angle, translation distance, scaling ratio, and flip type. By randomly generating transformation parameters, various real-world scenarios can be simulated.
[0056] Based on the transformation parameters, the normalized target region image is transformed to obtain a transformed target region image. The transformation process includes rotation, translation, scaling, and flipping. These transformations can simulate license plate images under different angles, distances, and lighting conditions.
[0057] Random noise is added to the transformed target region image according to preset noise type and intensity, resulting in a noisy target region image. Adding random noise can further enhance the model's anti-interference ability. Common noise types include Gaussian noise and salt-and-pepper noise.
[0058] The target region image and multiple noisy target region images after data augmentation are combined to construct a target region image dataset. By combining the original image and the data-augmented images, a diverse dataset can be obtained. For example, for each original image, 10 images that have been rotated, translated, scaled, flipped, and have noise added to varying degrees can be generated, thereby expanding the dataset by 10 times.
[0059] A convolutional neural network algorithm is used to train a target detection model using the target region image dataset as training data, resulting in a trained target detection model.
[0060] The trained target detection model is used to perform target detection on the input image, thereby achieving accurate localization and recognition of the target region.
[0061] Step S108 includes: acquiring the original photo, and performing normalization processing on the original photo according to a preset image normalization rule to obtain a normalized photo.
[0062] Image analysis algorithms are used to extract key information from the normalized photos. This key information includes facial features and scene labels. Facial feature extraction can utilize deep learning models, such as convolutional neural networks, to identify facial key points, age, gender, and other attributes. Scene labels can be generated using pre-trained image classification models, such as "indoor," "office," and "natural scenery." This information provides crucial information for subsequent retrieval and analysis.
[0063] The normalized photos and key information are encapsulated according to a preset data format to obtain structured archive data; the encapsulation of structured archive data can use JSON format. Each record contains fields such as the normalized photo file path, facial feature vector, and scene label list. This format facilitates storage and retrieval while maintaining data integrity and relevance.
[0064] According to the preset database table structure, the structured archival data is stored in the archival database. In the archival database, an image feature index is created for the normalized photos, and a text index is created for the key information. The image feature index uses techniques such as Local Sensitive Hash (LSH) to map high-dimensional feature vectors to a low-dimensional space, enabling fast similarity queries. The text index adopts an inverted index structure, supporting fast keyword matching.
[0065] By analyzing the similarity between the normalized photos and the correlation between the key information, the system establishes associations between photos and between photos and key information in the archive database. Cosine similarity is used to compare feature vectors; when the similarity exceeds 0.9, two photos are considered highly correlated. The association between photos and key information can be established through co-occurrence analysis; for example, facial features frequently appearing under the same scene label may indicate a close relationship. Upon receiving an image retrieval request, the system first parses the request content. If it is an image query, the system extracts the feature vector of the query image and searches for similar photos in the image feature index. If it is a text query, the system matches keywords in the text index to find relevant photos and information.
[0066] When an image retrieval request is received, the image or text information contained in the request is obtained. By querying the image feature index and the text index, and using the association relationship, relevant photos and key information are retrieved from the archive database and returned to the requester.
[0067] The system can also leverage established relationships to expand search results and provide more comprehensive information. This archival management system not only improves search efficiency but also uncovers potential connections between photographs, providing strong support for archival analysis and utilization. For example, in historical research, combinations of personal characteristics and scene tags can quickly locate relevant photographs from specific periods and occasions, helping researchers better reconstruct historical scenes and relationships between individuals.
Claims
1. An image preprocessing method for power system archive entry, characterized by, The method comprises the following steps: S101: obtaining an original photo, judging whether the photo has quality problems through an image quality evaluation model, improving the quality of the photo with problems, and obtaining a photo with improved quality; S102: positioning a text region for the photo with improved quality, correcting the direction of the text, and then identifying the text content through an optical character recognition technology to obtain a photo with corrected text; S103: judging the type of the photo with corrected text according to a photo type classification model, adjusting feature extraction parameters and region segmentation strategies for different types of photos; S104: separating the foreground and background of the photo, removing watermarks and stains, and retaining a target region to obtain a segmented target region; S105: performing feature extraction on the segmented target region and comparing the feature extraction with a pre-established photo type template to determine the position of key information; S106: extracting a target region according to the determined position of key information, accurately positioning the information boundary through edge detection and contour analysis, and obtaining the extracted target region; S107: performing normalization processing on the extracted target region to construct a data basis; S108: storing the photo after normalization processing and the extracted information in an archive database, establishing an index and a correlation relationship, and used for subsequent image retrieval and analysis.
2. The image pre-processing method for power system archive entry according to claim 1, characterized in that, Step S101 comprises: obtaining an original photo, inputting the original photo into a pre-established image quality evaluation model, and judging whether the original photo has overexposure, underexposure, blur, and tilt quality problems; if the original photo has overexposure or underexposure problems, performing brightness adjustment on the original photo through an adaptive histogram equalization algorithm to obtain a photo with adjusted brightness; if the original photo has blur problems, calculating the Laplacian variance of the original photo through a Laplacian operator, determining a blur region, and processing through a deblurring algorithm to obtain a photo after deblurring; if the original photo has tilt problems, detecting the tilt angle of the original photo through a Hough transform, and performing rotation correction on the original photo to obtain a photo after rotation correction; fusing the photo with adjusted brightness, the photo after deblurring, and the photo after rotation correction to obtain a photo with improved quality.
3. The image pre-processing method for power system file entry according to claim 1, wherein, Step S102 comprises: obtaining a photo with improved quality, extracting image features through a convolutional neural network model to obtain an image feature vector; detecting and positioning the text region in the image through a CTPN model according to the image feature vector to obtain text region coordinates; judging the direction of the text through a direction classification model for the text region coordinates, correcting the direction if the direction of the text is not correct, and obtaining a text region image after direction correction; performing image enhancement processing on the text region image after direction correction, inputting the enhanced text region image into a CRNN model for text recognition to obtain a recognition result; performing error correction processing on the recognition result through a language model to obtain a final text recognition result, embedding the recognized text content into a corresponding region of the original photo, and outputting a photo with corrected text.
4. The image pre-processing method for power system file entry according to claim 1, wherein, Step S103 comprises: Obtaining the target photo after text correction, judging which type the target photo belongs to according to the pre-established photo type classification model; According to the type of the target photo, the target photo is divided into different regions by using the corresponding region segmentation strategy; For different regions after segmentation, the corresponding feature descriptors are extracted to obtain the face feature vector and the text feature vector or the subject feature vector; The feature vectors extracted are combined into a comprehensive feature vector as the feature representation of the target photo.
5. The image pre-processing method for power system file entry according to claim 1, wherein, Step S104 includes: Obtaining the target photo to be processed, inputting it into the pre-trained semantic segmentation network model for foreground and background separation to obtain foreground and background masks; For the foreground mask, the small area watermark region is removed by using morphological opening operation, and the stain hole in the target region is filled by using morphological closing operation to obtain the foreground mask after removing the watermark and the stain; According to the foreground mask after removing the watermark and the stain, the corresponding foreground target region image is extracted from the target photo; For the foreground target region image, the brightness, contrast and color are adjusted by using the image enhancement algorithm to obtain the image enhanced foreground target region image; The image enhanced foreground target region image is fused with the background region of the target photo to obtain the fused photo; The fused photo is optimized by using the image inpainting algorithm to eliminate the abrupt feeling of the foreground target edge to obtain the repaired target photo.
6. The image pre-processing method for power system file entry according to claim 1, wherein, Step S105 includes: Obtaining the pre-established photo type template, extracting the image feature vector for each photo type, and constructing a feature template library; Obtaining the image to be processed, segmenting the image to be processed by using the image segmentation algorithm to obtain a plurality of target regions; For each target region, the image feature vector of the region is extracted by using the BIC algorithm; The feature vector of the target region is matched with the feature vector in the feature template library to calculate the similarity; According to the similarity, it is judged which photo type the target region belongs to, and the key information position corresponding to the photo type is obtained; If the similarity between the target region and a photo type template exceeds a preset threshold, the region is marked as the photo type, and the key information position coordinates are recorded; According to the type mark and the key information position coordinates of each target region, the key information position in the image to be processed is determined.
7. The image pre-processing method for power system file entry according to claim 1, wherein, Step S106 includes: According to the pre-established key information position model, the key information position in the image to be processed is determined, and the key information position coordinates are obtained; For the obtained key information position coordinates, the adaptive threshold segmentation algorithm is used to segment the image to be processed to obtain the segmented target region image; The segmented target region image is subjected to edge detection processing to extract the edge information of the target region and obtain the edge coordinates of the target region; According to the obtained edge coordinates of the target region, contour analysis processing is performed to accurately position the boundary position of the target region to obtain accurate target region boundary coordinates; determining whether the target region boundary coordinates match the key information position coordinates, and if so, obtaining the extracted target region image; if not, adjusting parameters of the adaptive threshold segmentation algorithm, re-performing target region extraction, and repeating until the extracted target region is accurate.
8. The image pre-processing method for power system file entry according to claim 1, wherein, Step S107 includes: extracting a target region image according to preset target region coordinates; performing normalization processing on the target region image to unify image size and resolution; randomly generating transformation parameters according to preset data enhancement parameters; performing transformation processing on the normalized target region image according to the transformation parameters to obtain a transformed target region image; adding random noise to the transformed target region image according to preset noise types and noise intensities to obtain a target region image with added noise; combining the target region image and the multiple target region images with added noise after data enhancement processing to construct a target region image dataset.
9. The image pre-processing method for power system file entry according to claim 1, wherein, Step S108 includes: packaging the normalized photo and the key information according to a preset data format to obtain structured archive data; storing the structured archive data into an archive database according to a preset database table structure, establishing an image feature index for the normalized photo, and establishing a text index for the key information.