A method, system and device for intelligent text recognition in complex power scenarios

Through frequency domain conversion and masking processing, text blurring in complex power scenarios is simulated, and text structure adjustment weight is calculated, which solves the problem of inaccurate positioning of text recognition in complex power scenarios and improves the accuracy of text recognition.

CN118711199BActive Publication Date: 2025-05-13STATE GRID INFORMATION & TELECOMM BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410916324.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2025-05-13
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

In complex power scenarios, the installation position of the surveillance camera is not fixed, resulting in blurred text acquisition of electrical equipment, and existing machine learning algorithms are difficult to accurately locate text recognition, resulting in low recognition accuracy.

Method used

By inputting the initial frame-extracting test images into the frequency domain conversion model, the frequency domain map to be analyzed is obtained, and the frequency domain masking process of different sizes is performed to obtain the test fuzzy grayscale map of different blur levels. Clustering and connectivity domain analysis is performed according to the degree of loss of font details, the text structure under different fuzzy degrees is obtained, and the text structure adjustment weight is calculated to correct the positioning box during the recognition process.

Benefits of technology

By simulating the text structure under different blur levels, the fuzzy effect is accurately quantified, and the inaccurate positioning of the text recognition positioning box caused by structural blur is avoided, and the accuracy of text recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118711199B_ABST
    Figure CN118711199B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of text recognition, and specifically to a method, system and device for intelligent text recognition in complex power scenarios, including: obtaining a frequency domain map of an image to be analyzed by an initial frame extraction test; performing frequency domain masking processing on the frequency domain map of the image to be analyzed to obtain test fuzzy grayscale maps with different blur levels, comparing the test fuzzy grayscale maps with different blur levels with the initial frame extraction test image, and obtaining the degree of font detail loss of the test fuzzy grayscale maps with different blur levels; obtaining different text structures with the same blur level, and the degree of structural detail loss corresponding to different text structures with different blur levels, obtaining text structure adjustment weights for text recognition of electrical equipment in complex power scenarios with different blur levels, and performing positioning frame correction and positioning on text with different blur levels in an actual recognition process according to different text structure adjustment weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text recognition, and in particular to a method, system and device for intelligent text recognition in complex power scenarios. Background Art

[0002] In the power system, it is necessary to understand the status of electrical equipment and the parameter information of the equipment in real time. However, in complex indoor scenes, not all staff have the authority to enter the electrical equipment storage scene, and not all electrical equipment storage scenes are absolutely safe. Therefore, text recognition technology is an extremely effective technical method for real-time understanding of the status of electrical equipment and the parameter information of the equipment. It can realize contactless recognition of the status of electrical equipment in complex power scenes and text recognition in the parameter information of the equipment.

[0003] In the existing technology, machine learning algorithm is a commonly used text recognition algorithm, and its main logic is to extract and train the morphological features and shape features of the font for text recognition. However, in complex power scenes, because the installation position of the monitoring camera is not fixed, the text acquisition of the electrical equipment is fuzzy. In the process of using machine learning algorithms for text recognition, the text positioning box cannot be effectively positioned accurately due to the fuzzy text structure, and the text recognition accuracy is low, which cannot meet the effectiveness of intelligent text recognition in complex electronic scenes. Summary of the invention

[0004] In order to solve the above problems, the present invention provides a method, system and device for intelligent text recognition in complex power scenarios.

[0005] The present invention provides a method, system and device for intelligent text recognition in complex power scenarios using the following technical solutions:

[0006] The present invention proposes a method for intelligent text recognition in complex power scenarios, which includes the following steps:

[0007] Input the initial frame test image into the frequency domain conversion model to obtain the frequency domain image of the image to be analyzed;

[0008] Perform frequency domain masking of different sizes according to the obtained frequency domain image of the image to be analyzed to obtain test fuzzy grayscale images with different blur levels, compare the test fuzzy grayscale images with different blur levels with the initial frame extraction test image, and obtain the degree of font detail loss of the test fuzzy grayscale images with different blur levels;

[0009] According to the font detail loss degree of the test fuzzy grayscale image, clustering and connected domain analysis are performed at the same blur level to obtain different text structures at the same blur level, and the structural detail loss degree corresponding to different text structures at different blur levels;

[0010] According to the degree of loss of structural details corresponding to the same text structure under different blur levels, the text structure adjustment weights for text recognition of electrical equipment in complex power scenarios under different blur levels are obtained. According to the text structure adjustment weights for different text recognition, the positioning frame is corrected for the text under different blur levels in the actual recognition process.

[0011] Optionally, the specific steps of obtaining the frequency domain image of the image to be analyzed are as follows:

[0012] The surveillance video data collected by any one of the randomly selected cameras is frame-decomposed using a video decoder, and the entire video stream is parsed into a series of image frames;

[0013] Then, the first frame image is used as a reference image, and any frame image is used as a comparison image to perform peak signal-to-noise ratio calculation with the first frame image, so as to obtain peak signal-to-noise ratios of a series of image frames except the first frame image;

[0014] Count all the peak signal-to-noise ratios, and select the image frame corresponding to the maximum peak signal-to-noise ratio as the high-quality image, and use the high-quality image as the subsequent initial frame extraction test image;

[0015] The initial frame extraction test image is taken as the original image, and the frequency domain transform of the initial frame extraction test image is performed using a two-dimensional Fourier algorithm to obtain a frequency domain image of the image to be analyzed.

[0016] Optionally, the step of performing frequency domain mask processing of different sizes according to the obtained frequency domain image of the image to be analyzed to obtain a test blurred grayscale image with different blur levels includes the following specific steps:

[0017] The frequency domain image of the image to be analyzed is denoted as W, and the size of W is denoted as A×B pixels;

[0018] The centroid O of W is obtained by the rectangular centroid coordinate calculation formula;

[0019] Using O as the positioning point, frequency domain mask processing is performed on W, wherein the mask is a rectangular mask of (γ×A)×(γ×B) pixels whose size is adjusted by the scaling parameter γ, and the four sides of the rectangular mask are parallel to the four sides of W; thereby obtaining a frequency domain mask image of W with a mask size of γ;

[0020] Then, the frequency domain mask image with a mask size of γ of W is subjected to Fourier inverse operation to convert the frequency domain mask image with a mask size of γ of W into a spatial domain image. The spatial domain image is the S of the initial frame extraction test image. γ Fuzz test grayscale image of the degree of blur;

[0021] Finally, we traverse γ between [0,1] and take the traversal interval as T to obtain Grayscale images of fuzzy tests at different blur levels.

[0022] Optionally, the step of comparing the test fuzzy grayscale images with different blur levels with the initial frame extraction test image to obtain the degree of font detail loss of the test fuzzy grayscale images with different blur levels includes the following specific steps:

[0023] The initial frame test image is recorded as S, and the machine learning classifier is used to mark the text boxes of the electrical equipment text locations on S to obtain text boxes of different text areas;

[0024] For each text area text box obtained on S, the S of the initial frame test image is γ On the fuzzy test grayscale image of the fuzzy degree, mark the same position to obtain multiple marked areas;

[0025] Calculate the grayscale gradient value of a font pixel point G(x,y) in a fixed direction in each text area text box in S, and record the gradient value as the original grayscale gradient value of the pixel point G(x,y);

[0026] Calculate the initial frame test image S γ In each marked area of ​​the grayscale image, a font pixel G γ The grayscale gradient value in the same direction as the original grayscale gradient value of (x, y) is recorded as the fuzzy gradient value F γ (x,y);

[0027] Calculate the absolute value of the difference between F(x,y) and G(x,y) Δ γ (x,y),Δ γ (x, y) is the font pixel G(x, y) in S with a blur level of S γ The degree of loss of font detail under

[0028] Optionally, the specific steps of obtaining different text structures at the same blur level are as follows:

[0029] Use the position of each font pixel in the text area of ​​S to obtain the γ The pixel points at the same position in the fuzzy test grayscale image of the fuzzy degree are denoted as S γ Font pixels at different blur levels;

[0030] S γ The degree of loss of font details of each pixel under the blur level is used as a clustering parameter. All clustering parameters are clustered using adaptive k-means clustering to obtain k clusters with different clustering parameters.

[0031] For S γ Fuzzy test grayscale image of fuzziness S γ The font pixels under the blur level are detected by using morphological algorithm to obtain multiple font structure connected domains;

[0032] The font structure connectivity domain and clusters are comprehensively compared. For the clustering parameters in the same cluster, the S γ The font pixels at different blur levels and in the same font structure connected domain are denoted as S γ A text structure in the fuzzy test grayscale image with different blur levels;

[0033] Get S γ Fuzz test grayscale image of all different text structures with different blur levels.

[0034] Optionally, the specific steps for obtaining the degree of loss of structural details corresponding to different text structures at different blur levels are as follows:

[0035] Get an S of S γ Any text structure in the fuzzy test grayscale image of the fuzziness level is used as an example structure;

[0036] Calculate the average value of the font detail loss of all blurred font pixels constituting the example structure, denoted as S of the example structure. γ The degree of loss of structural details under blur level;

[0037] Obtain the degree of structural detail loss corresponding to different text structures under different blur levels.

[0038] Optionally, obtaining the text structure adjustment weights for text recognition of electrical equipment in complex power scenarios at different fuzziness levels according to the degree of structural detail loss corresponding to the same text structure at different fuzziness levels includes the following specific steps:

[0039] Get S of S γ The degree of blurriness of the fuzzy test is the degree of loss of structural details of all text structures in the grayscale image, and the S γ In the fuzzy test grayscale image of the fuzzy degree, the degree of loss of structural details of all text structures is normalized to obtain the normalized parameters of the degree of loss of structural details of all text structures. The normalized parameter of the degree of loss of structural details of each text structure is S γ The text structure adjustment weight of this text structure in the fuzzy test grayscale image of the fuzziness level.

[0040] Optionally, the text structure adjustment weights according to different text recognition are used to correct the positioning of the text with different blur levels in the actual recognition process, and the specific steps include the following:

[0041] The machine learning algorithm is used to perform text recognition and text positioning frame positioning on the detection frame. When the structure recognition is performed during the positioning process, during the text structure connected domain detection process in the machine learning algorithm, when the detected connected domain is the same connected domain as the text structure under the current blur level, the possibility of the detected connected domain as the text structure is adjusted. The adjustment method is the possible judgment of the inherent machine learning algorithm that the current connected domain is a text structure plus the text structure adjustment weight of this text structure under the blur level;

[0042] Finally, the adjusted probability is used to detect text recognition and text positioning in the frame.

[0043] The present invention also proposes a system for intelligent text recognition in complex power scenarios, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned method for intelligent text recognition in complex power scenarios.

[0044] The present invention also proposes a device for intelligent text recognition in complex power scenarios, including an image acquisition module, an image analysis and processing module, and a text positioning module, so as to implement the steps of the aforementioned method for intelligent text recognition in complex power scenarios.

[0045] The beneficial effects of the technical solution of the present invention are as follows: the present invention obtains an initial frame extraction test image input frequency domain conversion model in a monitoring video to obtain a frequency domain map of the image to be analyzed; by performing frequency domain masking of different sizes on the obtained frequency domain map of the image to be analyzed, test fuzzy grayscale maps with different blur levels are obtained, and the actual fuzzy image is simulated by the fuzzy image under the mask, and the text blur effect of the electrical equipment can be intuitively displayed by adjusting the mask size; then the test fuzzy grayscale map with different blur levels is compared with the initial frame extraction test image to obtain the degree of font detail loss of the test fuzzy grayscale map with different blur levels; clustering and connected domain analysis are performed on the degree of font detail loss of the test fuzzy grayscale map at the same blur level to obtain different text structures at the same blur level, as well as the degree of structural detail loss corresponding to different text structures at different blur levels, so as to accurately quantify the performance of different text structures under different blur effects, so that when the fuzzy text is extracted later, inaccurate positioning of the text recognition positioning frame due to structural blur is avoided. According to the performance of different text structures under different blur effects, the text positioning frame is corrected and adjusted in the process of intelligent text recognition in complex power scenarios based on machine learning to ensure the accuracy of the text positioning frame. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 A flowchart of the steps of a method for intelligent text recognition in complex power scenarios provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the method, system and device for intelligent text recognition in complex power scenarios proposed by the present invention, its specific implementation method, structure, characteristics and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0049] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0050] The text recognition algorithm in general scenarios is to usually use a machine learning target detection algorithm to locate and identify the position and bounding box of the text in the image, and then recognize the detected text area to convert the text in the image into text that can be processed by a computer.

[0051] However, text recognition of electric energy equipment in power scenarios mainly involves text recognition in videos captured by the monitoring system. Due to the randomness of the positions of different monitoring cameras and the characteristics of the text size of electric energy equipment, the text of different electric energy equipment has different expressions. For example, the camera is far away and the text of the electric energy equipment is small, resulting in blurry areas for text recognition. Using machine learning target detection algorithms to locate or mark the position and bounding box of text in an image will result in obvious anomalies, resulting in the text being marked as multiple marking boxes or unable to be effectively marked. Therefore, in order to accurately and comprehensively perform intelligent text recognition in complex power scenarios, it is not possible to use only a single machine learning algorithm to locate and mark text in power scenarios. Instead, it is necessary to fully simulate the complex environment in the power scenario and then perform accurate text recognition.

[0052] The following is a detailed description of a method, system and device for intelligent text recognition in complex power scenarios provided by the present invention in conjunction with the accompanying drawings.

[0053] See also Figure 1 , which shows a flowchart of a method for intelligent text recognition in a complex power scenario provided by an embodiment of the present invention, the method comprising the following steps:

[0054] Step S1, inputting the initial frame test image into the frequency domain conversion model to obtain the frequency domain image of the image to be analyzed;

[0055] Collect surveillance video data in power scenarios. These surveillance videos contain multiple frames of images.

[0056] When performing text recognition on the initial frame extraction test image, the surveillance video data collected by a random camera among several cameras is selected as the data acquisition source. According to the data source, high-quality images are extracted and the extracted high-quality images are used as the initial frame extraction test image; then, based on the initial frame extraction image, the image frequency domain map to be analyzed of the initial frame extraction image is obtained.

[0057] It should be noted that the purpose of randomly selecting the surveillance video data collected by any one of the several cameras as the data acquisition source in this embodiment is that: considering that different cameras correspond to different power scenes, the blurriness of the text on the electrical equipment in the collected videos is different, and the amount of calculation is too large when all the data corresponding to the cameras are collected, processed and analyzed. The main purpose of this embodiment is to perform machine learning training under different blurriness for the initial frame test images in the surveillance video, so any camera can be used as an example for explanation.

[0058] As an example, the method for obtaining the initial frame test image is as follows:

[0059] The surveillance video data collected by any one of the randomly selected cameras is frame-decomposed using a video decoder, and the entire video stream is parsed into a series of image frames;

[0060] Then, the first frame image is used as a reference image, and any frame image is used as a comparison image to perform peak signal-to-noise ratio calculation with the first frame image, so as to obtain peak signal-to-noise ratios of a series of image frames except the first frame image;

[0061] Count all the peak signal-to-noise ratios, and select the image frame corresponding to the maximum peak signal-to-noise ratio as the high-quality image, and use the high-quality image as the subsequent initial frame extraction test image;

[0062] It should be noted that the viewing angle of the surveillance camera in the power scenario is generally fixed, that is, the series of image frames analyzed from the surveillance video data corresponding to the camera are highly similar. Without considering human factors, this series of image frames is only affected by electromagnetic noise. Therefore, this embodiment selects a frame with the largest peak signal-to-noise ratio as the initial frame test image, the purpose of which is to reduce the impact of noise on subsequent analysis and processing.

[0063] As an example, the method for obtaining the frequency domain image of the image to be analyzed is as follows:

[0064] The initial frame extraction test image is taken as the original image, and the frequency domain transform of the initial frame extraction test image is performed using a two-dimensional Fourier algorithm to obtain a frequency domain image of the image to be analyzed.

[0065] Step S2, perform frequency domain masking of different sizes according to the obtained frequency domain image of the image to be analyzed to obtain test blurred grayscale images with different blur levels, compare the test blurred grayscale images with different blur levels with the initial frame test image, and obtain the degree of font detail loss of the test blurred grayscale images under different blur levels.

[0066] In the above, this embodiment has obtained the frequency domain graph of the image to be analyzed of the initial frame test image. It should be noted that the blur of the text image in the surveillance video is uncontrollable, and is often accompanied by the loss of high-frequency and low-frequency details of the image. The frequency domain image can reflect the specific high- and low-frequency information of the image, and after the image is frequency-domain masked, it is converted into a time-domain image. The frequency information of the part masked in the frequency domain will be permanently lost in the time domain graph, which can simulate the blur of the text image in the surveillance video.

[0067] Therefore, this embodiment performs frequency domain masking of different sizes on the frequency domain image of the image to be analyzed, and converts the masked frequency domain image back into a spatial domain image, and obtains test fuzzy grayscale images with different blur levels due to different blurs of the initial frame extraction test image caused by frequency domain masks of different sizes, which are used to simulate the different blur levels of the text of electrical equipment collected by the monitoring video in complex power scenes. And by comparing and analyzing the test fuzzy grayscale images with different blur levels generated by frequency domain masks of different sizes with the initial frame extraction test image, it is determined which details are more blurred under the current mask.

[0068] As a preferred example, the method for obtaining the test blur grayscale images with different blur levels is as follows:

[0069] The frequency domain image of the image to be analyzed is denoted as W, and the size of W is denoted as A×B pixels;

[0070] The centroid O of W is obtained by the rectangular centroid coordinate calculation formula;

[0071] Using O as the positioning point, frequency domain mask processing is performed on W, wherein the mask is a rectangular mask of (γ×A)×(γ×B) pixels whose size is adjusted by the scaling parameter γ, and the four sides of the rectangular mask are parallel to the four sides of W; thereby obtaining a frequency domain mask image of W with a mask size of γ;

[0072] Then, the frequency domain mask image with a mask size of γ of W is subjected to Fourier inverse operation to convert the frequency domain mask image with a mask size of γ of W into a spatial domain image. The spatial domain image is the S of the initial frame extraction test image. γ Fuzz test grayscale image of the degree of blur;

[0073] For γ, traverse between [0,1] and take the traversal interval as T, we get Fuzzy test grayscale images under different blur levels; the empirical value of T is T=0.1.

[0074] It should be noted that this embodiment is aimed at text recognition in complex scenes, and the text in the unclear and blurred images captured by the surveillance video contains a large number of edges, which is reflected in the frequency domain of the image containing a large amount of high-frequency information. In W, the transition process from high-frequency information to low-frequency information starts from O and radiates to the surroundings. Therefore, in this embodiment, the centroid O of W is used as the starting position of the mask to perform a gradually expanding frequency domain mask. The purpose is to blur the image with this simulated blur technology, which can gradually blur the details from the edge to the inside of the text, and more intuitively analyze the specific blur range of the blur test grayscale image of the initial frame test image under different blur levels.

[0075] As an example, the method for obtaining the degree of font detail loss of the test blurred grayscale image at different blur levels is as follows:

[0076] The initial frame test image is recorded as S, and the machine learning classifier is used to mark the text boxes of the electrical equipment text locations on S to obtain text boxes of different text areas;

[0077] For each text area text box obtained on S, the S of the initial frame test image is γ On the fuzzy test grayscale image of the fuzzy degree, mark the same position to obtain multiple marked areas;

[0078] Calculate the grayscale gradient value of a font pixel point G(x,y) in a fixed direction in each text area text box in S, and record the gradient value as the original grayscale gradient value of the pixel point G(x,y);

[0079] Calculate the initial frame test image S γ In each marked area of ​​the grayscale image, a font pixel Gγ The grayscale gradient value in the same direction as the original grayscale gradient value is (x, y), and the grayscale gradient value is recorded as the fuzzy gradient value F γ (x,y);

[0080] Calculate the absolute value of the difference between F(x,y) and G(x,y) Δ γ (x,y), then Δ γ (x, y) is the font pixel G(x, y) in S with a blur level of S γ The degree of loss of font detail under;

[0081] It should be noted that the initial frame test image S is used as a comparison image and is assumed to be a normal, non-blurred image. The existing machine learning classifier technology can be used to locate the text in S and mark the text box. γ Although the blur test grayscale image of the blur degree is processed by frequency domain masking, the actual size of the image in the time domain is the same as S, that is, every font pixel in S is the same as S. γ The font pixels at the same position in the grayscale image of the fuzzy test are one-to-one corresponding; then there is a text box located on S, γ The same position in the fuzzy test grayscale image is also the text box corresponding to the same text; after masking, S γ The fuzzy test grayscale image of the fuzzy degree must be fuzzy compared to S, that is, both the high-frequency and low-frequency information have changed. Specifically, it is manifested in the image as the gradient value of the pixel point, so S γ Fuzzy test grayscale image G γ The gradient value of (x, y) and G(x, y) in S in the same direction must have changed. The greater the degree of change, the greater the change in G(x, y) in S after passing through S. γ After blurring, the corresponding G γ (x,y) The greater the loss of font detail.

[0082] Furthermore, by using the above method to calculate each pixel point of the text area text box in S and the fuzzy test grayscale image with different blur levels, the degree of font detail loss in the fuzzy test grayscale image with different blur levels for each font pixel point in the text area text box of S can be obtained.

[0083] Step S3: performing clustering and connected domain analysis according to the font detail loss degree of the test fuzzy grayscale image at the same fuzziness level to obtain different text structures at the same fuzziness level, and the structural detail loss degree corresponding to different text structures at different fuzziness levels.

[0084] In the above logic, the degree of font detail loss obtained is for a single font pixel. However, in the process of using machine learning algorithms to locate text, the specific structure of the text is identified. The overall calculation amount for locating a single pixel is too large, and the recognition effect is not accurate enough. Therefore, it is necessary to analyze the pixels with font detail loss in the test fuzzy grayscale images with different blur levels to determine whether some pixels with font detail loss belong to the same structure of the font, and then conduct subsequent analysis.

[0085] It should be additionally explained that not all pixels with font detail loss in the test fuzzy grayscale image at the same blur level have the same degree of font detail loss, because the test fuzzy grayscale image at the same blur level is obtained by converting the frequency domain mask image into the spatial domain, and the frequency spectrum distribution of different electrical equipment characters in the initial frame extraction test image is different. If the frequency domain spectrum corresponding to some characters belongs to the partial spectrum processed by the frequency domain mask, then the font detail loss degree of the corresponding pixel in the test fuzzy grayscale image is greater, otherwise the corresponding font detail loss degree is smaller.

[0086] In addition, in the test blurred grayscale image at the same blur level, pixels with similar degree of font detail loss, although the spectral distribution in the frequency domain is masked, these pixels may belong to different fonts in the test blurred grayscale image and thus do not necessarily represent the same structure of the font.

[0087] Therefore, in order to avoid misjudging the different text structures of a font, this embodiment clusters the font detail loss degrees corresponding to different font pixels under the same blur degree, and then performs a connected domain analysis on the font pixels with similar font detail loss degrees in the same category. The different text structures of the font are obtained by determining whether the pixels with similar font detail loss degrees belong to the same connected domain.

[0088] As an example, the method of obtaining different text structures at the same blur level is as follows:

[0089] Use the position of each font pixel in the text area of ​​S to obtain the γ The pixel points at the same position in the fuzzy test grayscale image of the fuzzy degree are denoted as S γ Font pixels at different blur levels;

[0090] S γThe degree of loss of font details of each pixel under the blur level is used as a clustering parameter. All clustering parameters are clustered using adaptive k-means clustering to obtain k clusters with different clustering parameters.

[0091] For S γ Fuzzy test grayscale image of fuzziness S γ The font pixels under the blur level are detected by using morphological algorithm to obtain multiple font structure connected domains;

[0092] The font structure connectivity domain and clusters are comprehensively compared. For the clustering parameters in the same cluster, the S γ The font pixels at different blur levels and in the same font structure connected domain are denoted as S γ A text structure in the fuzzy test grayscale image with different blur levels;

[0093] Similarly, you can get S γ Fuzz test grayscale image of all different text structures with different blur levels.

[0094] As another example, the method for obtaining the degree of structural detail loss corresponding to different text structures at different blur levels is as follows:

[0095] Get an S of S γ Any text structure in the fuzzy test grayscale image of the fuzziness level is used as an example structure;

[0096] Calculate the average value of the font detail loss of all blurred font pixels constituting the example structure, denoted as S of the example structure. γ The degree of loss of structural details under blur level;

[0097] Similarly, the degree of loss of structural details corresponding to different text structures under different blur levels can be obtained.

[0098] Step S4: According to the degree of loss of structural details corresponding to the same text structure under different blur levels, the text structure adjustment weights for text recognition of electrical equipment in complex power scenarios under different blur levels are obtained, and the positioning frame is corrected and positioned for the text under different blur levels in the actual recognition process according to the text structure adjustment weights for different text recognition.

[0099] In history, the degree of loss of structural details corresponding to different text structures under different degrees of blur was obtained. However, in the process of positioning the text box in the intelligent recognition of text in complex power scenes using machine learning algorithms, it is not that the blurred structure cannot be recognized, but that it is judged as a low-value area after recognition, thereby causing the original text structure to be ignored, resulting in disconnection of the positioning box and inaccurate positioning. Therefore, this embodiment needs to perform an adjustment weight calculation based on the degree of loss of structural details for the font structure in the recognition process, and adjust the weight to correct the low-value structural judgment error in the text recognition process of the machine learning algorithm, so that the blurred text structure is considered to be a high-value area, and then the text positioning box under different degrees of blur is corrected and positioned.

[0100] As an example, the text structure adjustment weight method for text recognition of electrical equipment in complex power scenarios with different fuzziness levels is obtained as follows:

[0101] Get S of S γ The degree of blurriness of the fuzzy test is the degree of loss of structural details of all text structures in the grayscale image, and the S γ In the fuzzy test grayscale image of the fuzzy degree, the degree of loss of structural details of all text structures is normalized to obtain the normalized parameters of the degree of loss of structural details of all text structures. The normalized parameter of the degree of loss of structural details of each text structure is S γ The text structure adjustment weight of this text structure in the fuzzy test grayscale image of the fuzziness level.

[0102] As an example: the specific method of adjusting the weights according to different text structures of text recognition to correct the positioning frame of text with different blur levels in the actual recognition process is as follows:

[0103] Select any frame of surveillance video data captured by a surveillance camera in any complex power scene as a detection frame, and use spectrum calculation to obtain the pixel degree of the spectrum of the detection frame at different blur levels to determine the blur level of the detection frame;

[0104] The machine learning algorithm is used to locate the text positioning frame for text recognition in the detection frame. During the positioning process, the structure recognition is performed. During the text structure connected domain detection process in the machine learning algorithm, when the detected connected domain is the same connected domain as the text structure under the current blur level, the possibility of the detected connected domain as the text structure is adjusted. The adjustment method is the inherent machine learning algorithm's possible judgment that the current connected domain is a text structure plus the text structure adjustment weight of this text structure under the blur level. Finally, the adjusted possibility is used to locate the text recognition text in the detection frame.

[0105] At this point, this embodiment is completed. Compared with the existing intelligent text recognition in complex power scenes based on machine learning, this embodiment avoids the interference of fuzzy image recognition caused by the complexity of the monitoring video scene to a certain extent.

[0106] The present invention also proposes a system for intelligent text recognition in complex power scenarios, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned method for intelligent text recognition in complex power scenarios.

[0107] The present invention also proposes a device for intelligent text recognition in complex power scenarios, including an image acquisition module, an image analysis and processing module, and a text positioning module, so as to implement the steps of the aforementioned method for intelligent text recognition in complex power scenarios.

[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for intelligent text recognition in complex power scenarios, characterized in that: The method comprises the following steps: Input the initial frame test image into the frequency domain conversion model to obtain the frequency domain image of the image to be analyzed; Perform frequency domain masking of different sizes according to the obtained frequency domain image of the image to be analyzed to obtain test fuzzy grayscale images with different blur levels, compare the test fuzzy grayscale images with different blur levels with the initial frame extraction test image, and obtain the degree of font detail loss of the test fuzzy grayscale images with different blur levels; According to the font detail loss degree of the test fuzzy grayscale image, clustering and connected domain analysis are performed at the same blur level to obtain different text structures at the same blur level, and the structural detail loss degree corresponding to different text structures at different blur levels; According to the degree of structural detail loss corresponding to the same text structure under different blur levels, the text structure adjustment weights of text recognition of electrical equipment in complex power scenarios under different blur levels are obtained, and the text under different blur levels in the actual recognition process is positioned according to the text structure adjustment weights of different text recognitions. The specific steps of adjusting the weights according to the text structures of different text recognitions to correct the positioning of the texts with different blur levels in the actual recognition process include the following: The machine learning algorithm is used to perform text recognition and text positioning frame positioning on the detection frame. When the structure recognition is performed during the positioning process, during the text structure connected domain detection process in the machine learning algorithm, when the detected connected domain is the same connected domain as the text structure under the current blur level, the possibility of the detected connected domain as the text structure is adjusted. The adjustment method is the possible judgment of the inherent machine learning algorithm that the current connected domain is a text structure plus the text structure adjustment weight of this text structure under the blur level; Finally, the adjusted probability is used to detect text recognition and text positioning in the frame.

2. According to the method for intelligent text recognition in complex power scenarios described in claim 1, it is characterized in that: The specific steps of obtaining the frequency domain image of the image to be analyzed are as follows: The surveillance video data collected by any one of the randomly selected cameras is frame-decomposed using a video decoder, and the entire video stream is parsed into a series of image frames; Then, the first frame image is used as a reference image, and any frame image is used as a comparison image to perform peak signal-to-noise ratio calculation with the first frame image, so as to obtain peak signal-to-noise ratios of a series of image frames except the first frame image; Count all the peak signal-to-noise ratios, and select the image frame corresponding to the maximum peak signal-to-noise ratio as the high-quality image, and use the high-quality image as the subsequent initial frame extraction test image; The initial frame extraction test image is taken as the original image, and the frequency domain transform of the initial frame extraction test image is performed using a two-dimensional Fourier algorithm to obtain a frequency domain image of the image to be analyzed.

3. According to the method for intelligent text recognition in complex power scenarios described in claim 1, it is characterized in that: The specific steps of performing frequency domain mask processing of different sizes according to the obtained frequency domain image of the image to be analyzed to obtain the test blurred grayscale image with different blur levels are as follows: The frequency domain image of the image to be analyzed is denoted as W, and the size of W is denoted as A×B pixels; The centroid O of W is obtained by the rectangular centroid coordinate calculation formula; Using O as the positioning point, frequency domain mask processing is performed on W, wherein the mask is a rectangular mask of (γ×A)×(γ×B) pixels whose size is adjusted by the scaling parameter γ, and the four sides of the rectangular mask are parallel to the four sides of W; thereby obtaining a frequency domain mask image of W with a mask size of γ; Then, the frequency domain mask image with a mask size of γ of W is subjected to Fourier inverse operation to convert the frequency domain mask image with a mask size of γ of W into a spatial domain image. The spatial domain image is the S of the initial frame extraction test image. γ Fuzz test grayscale image of the blur level; Finally, we traverse γ between [0,1] and take the traversal interval as T to obtain Grayscale images of fuzzy tests at different blur levels.

4. According to the method for intelligent text recognition in complex power scenarios described in claim 1, it is characterized in that: The specific steps of comparing the test fuzzy grayscale images with different blur levels with the initial frame extraction test image to obtain the font detail loss degree of the test fuzzy grayscale images with different blur levels are as follows: The initial frame test image is recorded as S, and the machine learning classifier is used to mark the text boxes of the electrical equipment text locations on S to obtain text boxes of different text areas; For each text area text box obtained on S, the S of the initial frame test image is γ On the fuzzy test grayscale image of the fuzzy degree, mark the same position to obtain multiple marked areas; Calculate the grayscale gradient value of a font pixel point G(x,y) in a fixed direction in each text area text box in S, and record the gradient value as the original grayscale gradient value of the pixel point G(x,y); Calculate the initial frame test image S γ In each marked area of ​​the grayscale image, a font pixel G γ The grayscale gradient value in the same direction as the original grayscale gradient value of (x, y) is recorded as the fuzzy gradient value F γ (x,y); Calculate the absolute value of the difference between F(x,y) and G(x,y) Δ γ (x,y),Δ γ (x, y) is the font pixel G(x, y) in S with a blur level of S γ The degree of loss of font detail under 5. According to the method for intelligent text recognition in complex power scenarios described in claim 1, it is characterized by: The specific steps for obtaining different text structures under the same blur level are as follows: Use the position of each font pixel in the text area of ​​S to obtain the γ The pixel points at the same position in the fuzzy test grayscale image of the fuzzy degree are denoted as S γ Font pixels at different blur levels; S γ The degree of loss of font details of each pixel under the blur level is used as a clustering parameter. All clustering parameters are clustered using adaptive k-means clustering to obtain k clusters with different clustering parameters. For S γ Fuzzy test grayscale image of fuzziness S γ The font pixels under the blur level are detected by using morphological algorithm to obtain multiple font structure connected domains; The font structure connectivity domain and clusters are comprehensively compared. For the clustering parameters in the same cluster, the S γ The font pixels at different blur levels and in the same font structure connected domain are denoted as S γ A text structure in the fuzzy test grayscale image with different blur levels; Get S γ Fuzz test grayscale image of all different text structures with different blur levels.

6. According to the method of intelligent text recognition in complex power scenarios described in claim 1, it is characterized by: The specific steps for obtaining the degree of structural detail loss corresponding to different text structures under different blur levels are as follows: Get an S of S γ Any text structure in the fuzzy test grayscale image of the fuzziness level is used as an example structure; Calculate the average value of the font detail loss of all blurred font pixels constituting the example structure, denoted as S of the example structure. γ The degree of loss of structural details under blur level; Obtain the degree of structural detail loss corresponding to different text structures under different blur levels.

7. According to the method for intelligent text recognition in complex power scenarios described in claim 1, it is characterized by: The text structure adjustment weights for text recognition of electrical equipment in complex power scenarios at different fuzziness levels are obtained according to the degree of structural detail loss corresponding to the same text structure at different fuzziness levels, and the specific steps include the following: Get S of S γ The degree of blurriness of the fuzzy test is the degree of loss of structural details of all text structures in the grayscale image, and the S γ In the fuzzy test grayscale image of the fuzzy degree, the degree of loss of structural details of all text structures is normalized to obtain the normalized parameters of the degree of loss of structural details of all text structures. The normalized parameter of the degree of loss of structural details of each text structure is S γ The text structure adjustment weight of this text structure in the fuzzy test grayscale image of the fuzziness level.

8. A text intelligent recognition system in complex power scenarios, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method for intelligent text recognition in complex power scenarios as described in any one of claims 1-7 are implemented.

9. A text intelligent recognition device in complex power scenarios, characterized in that: The device comprises: Image acquisition module: used to input the initial frame test image into the frequency domain conversion model to obtain the frequency domain image of the image to be analyzed; Image analysis and processing module: perform frequency domain mask processing of different sizes according to the obtained frequency domain image of the image to be analyzed, obtain test fuzzy grayscale images with different blur levels, compare the test fuzzy grayscale images with different blur levels with the initial frame test image, and obtain the degree of font detail loss of the test fuzzy grayscale images with different blur levels; According to the font detail loss degree of the test fuzzy grayscale image, clustering and connected domain analysis are performed at the same blur level to obtain different text structures at the same blur level, and the structural detail loss degree corresponding to different text structures at different blur levels; Text positioning module: used to obtain text structure adjustment weights for text recognition of electrical equipment in complex power scenarios with different fuzziness levels according to the degree of structural detail loss corresponding to the same text structure with different fuzziness levels, and to perform positioning frame correction and positioning of text with different fuzziness levels in the actual recognition process according to the text structure adjustment weights of different text recognition; The specific steps of adjusting the weights according to the text structures of different text recognitions to correct the positioning of the texts with different blur levels in the actual recognition process include the following: The machine learning algorithm is used to perform text recognition and text positioning frame positioning on the detection frame. When the structure recognition is performed during the positioning process, during the text structure connected domain detection process in the machine learning algorithm, when the detected connected domain is the same connected domain as the text structure under the current blur level, the possibility of the detected connected domain as the text structure is adjusted. The adjustment method is the possible judgment of the inherent machine learning algorithm that the current connected domain is a text structure plus the text structure adjustment weight of this text structure under the blur level; Finally, the adjusted probability is used to detect text recognition and text positioning in the frame.

Citation Information

Patent Citations

  • Label identification method for fertilizer production line

    CN115578732A

  • Gray and color image fusion method based on Laplacian pyramid and self-attention mechanism

    CN118071613A