Image content analysis method, device and equipment

By introducing local binarization processing based on strong and weak thresholds in image content analysis, the label region is preprocessed in a refined manner, which solves the problem of low recognition accuracy of optical character recognition in complex lighting environments in the existing technology and achieves higher recognition accuracy.

CN120877290APending Publication Date: 2025-10-31CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510901276.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In complex lighting conditions, existing optical character recognition technology struggles to handle unclear, small, or blurry device labels, resulting in low recognition accuracy.

Method used

By acquiring the image to be processed, identifying the label region, calculating the strong and weak thresholds of the neighborhood of each pixel, and judging the binarized image of each pixel based on the corresponding thresholds, text detection and recognition of text content are performed.

Benefits of technology

By performing fine-grained preprocessing on the label area through local binarization, the adaptability to label text wear, blurring, or shrinkage is enhanced, thereby improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877290A_ABST
    Figure CN120877290A_ABST
Patent Text Reader

Abstract

The invention discloses an image content analysis method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed image, and recognizing a tag region of the to-be-processed image; determining a strong threshold value and a weak threshold value corresponding to the neighborhood of each pixel point in the label area; based on the corresponding strong threshold value and the weak threshold value, the binarization value of each pixel point is judged, and a binarization image corresponding to the label area is obtained; and performing text detection and recognition on the binarized image to obtain text content in the to-be-processed image. Local binarization processing based on a strong threshold value and a weak threshold value is introduced, fine preprocessing is carried out on a label area before text detection and recognition, complex light interference can be effectively adapted, the adaptability to label character abrasion, blurring or reduction is enhanced, and therefore a clearer and distinguishable text image is obtained before recognition; the identification accuracy under various complex scenes is improved, and the problem that in the prior art, only model training optimization is relied on, and the limitation of image preprocessing is ignored is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing, and specifically relates to an image content analysis method, apparatus, device, and storage medium. Background Technology

[0002] In asset management, the characters that need to be recognized are all on the equipment labels. OCR (Optical Character Recognition) technology can be used to quickly and accurately recognize the content information on the equipment label codes and then automatically fill it into the corresponding information field in the system, reducing the workload of manual operation and lowering the error rate of manual entry.

[0003] In existing technologies, the YOLOv4 model can be used for OCR recognition. By acquiring a training image set for OCR recognition, YOLOv4 and text recognition algorithms are used to sequentially perform text localization and text recognition on the initial training dataset, thereby training an initial OCR recognition model. Furthermore, the recognition results and proofreading results are compared at set intervals, original image samples with recognition errors are collected, abnormal element statistics are performed, and an optimization dataset is constructed. The initial OCR recognition model is then optimized and trained, and finally, an optimized OCR recognition model is obtained.

[0004] However, in real-world applications, device tags are typically in complex lighting environments, and the light they receive is affected by indoor and outdoor light sources. They often have complex image backgrounds and ambient light. The methods described above rely solely on training and optimizing the model with new samples, which is a simplistic and brute-force approach that is illegible when the device tags are unclear, too small, or blurry. Summary of the Invention

[0005] The purpose of this application is to provide an image content analysis method, apparatus, device, and storage medium that can solve the problem that current optical character recognition technology is unable to cope with unclear, small, or blurry device labels.

[0006] In a first aspect, embodiments of this application provide an image content analysis method, the method comprising:

[0007] Acquire the image to be processed and identify the label region of the image to be processed;

[0008] Determine the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region;

[0009] Based on the corresponding strong threshold and weak threshold, the binarization value of each pixel is determined to obtain the binarized image corresponding to the label region;

[0010] Text detection and recognition are performed on the binarized image to obtain the text content in the image to be processed.

[0011] Optionally, identifying the label region of the image to be processed includes:

[0012] The image to be processed is input into a pre-trained label detection model to obtain label location information;

[0013] The width and height of the label are calculated based on the label location information to obtain the label area.

[0014] Optionally, determining the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region includes:

[0015] Determine the neighborhood of each pixel in the label region;

[0016] Calculate the gray-level weighted average and gray-level weighted variance of the pixels in the region;

[0017] Calculate the strong threshold corresponding to the neighborhood of each pixel based on the first hyperparameter, the gray-level weighted average value, and the gray-level weighted variance;

[0018] The weak threshold corresponding to the neighborhood of each pixel is calculated based on the second hyperparameter, the gray-level weighted average value, and the gray-level weighted variance.

[0019] Optionally, determining the binarized value of each pixel based on the corresponding strong threshold and weak threshold includes:

[0020] For each pixel, if the grayscale value of the pixel is lower than the strong threshold, then the binarization value of the pixel is determined to be 0.

[0021] If the grayscale value of the pixel is higher than the weak threshold, then the binarization value of the pixel is determined to be 255;

[0022] If the grayscale value of the pixel is between the strong threshold and the weak threshold, then check whether the pixel is connected to a confirmed pixel whose binarization value is 0. If connected, then the binarization value of the pixel is determined to be 0; otherwise, the binarization value of the pixel is determined to be 255.

[0023] Optionally, determining the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region includes:

[0024] Based on the first size and the second size, determine two candidate neighborhoods for each pixel in the label region;

[0025] Calculate the gray-level weighted average of the pixels in each candidate neighborhood;

[0026] The candidate neighborhood with the larger gray-scale weighted average is used as the neighborhood of each pixel in the label region, and the corresponding strong threshold and weak threshold are calculated.

[0027] Optionally, the step of performing text detection and recognition on the binarized image to obtain the text content in the image to be processed includes:

[0028] The binarized image is converted into a three-channel image;

[0029] Perform text detection on the three-channel image and output a text box;

[0030] The clarity of the text box is evaluated. If the clarity is below a threshold, the text box is super-resolution enhanced to obtain a high-resolution image.

[0031] Text recognition is performed on the high-resolution image to obtain the text content in the image to be processed.

[0032] Optionally, the step of performing super-resolution enhancement on the text box to obtain a high-resolution image includes:

[0033] The text super-resolution backbone network extracts low-resolution image features of the text box;

[0034] The low-resolution image features and the pre-acquired prior features are adaptively normalized to obtain high-resolution image features;

[0035] The high-resolution image features are reverse-mapped to obtain a high-resolution image.

[0036] Secondly, embodiments of this application provide an image content analysis apparatus, the apparatus comprising:

[0037] An acquisition module is used to acquire the image to be processed and identify the label region of the image to be processed;

[0038] The determination module is used to determine the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region;

[0039] The determination module is used to determine the binarization value of each pixel based on the corresponding strong threshold and weak threshold, so as to obtain the binarized image corresponding to the label region.

[0040] The recognition module is used to perform text detection and recognition on the binarized image to obtain the text content in the image to be processed.

[0041] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0042] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0043] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0044] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0045] As can be seen from the above, in this embodiment of the application, by introducing local binarization processing based on strong and weak thresholds, the label region is finely preprocessed before text detection and recognition. This can effectively adapt to complex lighting interference, enhance adaptability to label text wear, blurring, or shrinkage, thereby obtaining a clearer and more distinguishable text image before recognition, improving the recognition accuracy in various complex scenarios, and solving the problem of simply relying on model training optimization while ignoring the limitations of image preprocessing in the prior art. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating an image content analysis method according to an exemplary embodiment;

[0047] Figure 2 This is a flowchart illustrating a label detection process according to an exemplary embodiment;

[0048] Figure 3 This is a schematic diagram of an AISRNet structure according to an exemplary embodiment;

[0049] Figure 4 This is a flowchart illustrating an image content analysis method according to an exemplary embodiment;

[0050] Figure 5 This is a block diagram illustrating an image content analysis apparatus according to an exemplary embodiment;

[0051] Figure 6 This is a block diagram illustrating an electronic device according to an exemplary embodiment;

[0052] Figure 7 This is a schematic diagram of the hardware structure of an electronic device according to an exemplary embodiment. Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0054] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0055] First, let me explain the terms mentioned in this application:

[0056] Multimodal data:

[0057] Multimodal data comprises two or more different modalities. These modalities can be text, images, audio, video, etc. Multimodal data extracts information about the same descriptive object from different perspectives or domains, with each modality providing another way to describe that object. For example, a multimodal system might simultaneously process audio and image information from a video, as well as text captions. Combining these different modalities within the same dataset can provide more comprehensive and multi-dimensional information, enabling more accurate analysis and reasoning.

[0058] Image binarization:

[0059] Image binarization is an image processing technique that converts each pixel in an image into two values: black and white (binary). For common color images, the color image is usually first converted to a grayscale image, and then the grayscale image is binarized back to a black and white image. In short, image binarization is an important image processing technique that enables more efficient processing of image information.

[0060] Image super-resolution:

[0061] Super-resolution (SR) is an image processing technique designed to transform low-resolution (LR) images into high-resolution (HR) images using algorithms. High-resolution images have higher pixel density, more detail, and finer image quality. Super-resolution methods include traditional methods and deep learning methods. Deep learning methods generally outperform traditional methods in terms of performance.

[0062] Transformer:

[0063] The Transformer is a deep learning model that employs a self-attention mechanism. Compared to traditional recurrent neural networks and long short-term memory networks, the Transformer offers higher parallelism and computational efficiency. Based on an encoder-decoder structure, the Transformer primarily maps input sequences to continuous representation sequences, and then uses a decoder to generate the complete output sequence. The overall structure includes self-attention, point-wise processing, and fully connected layers.

[0064] In related technologies, YOLOv4 can be used for OCR recognition. By acquiring a training image set for OCR recognition, YOLOv4 and text recognition algorithms are used to sequentially perform text localization and text recognition on the initial training dataset, thereby training an initial OCR recognition model. Furthermore, the recognition results and proofreading results are compared at set intervals, original image samples with recognition errors are collected, abnormal element statistics are performed, and an optimization dataset is constructed. The initial OCR recognition model is then optimized and trained, and finally, an optimized OCR recognition model is obtained.

[0065] However, in real-world applications, device tags are typically located in complex lighting environments, where the light used for imaging is influenced by indoor and outdoor light sources. This often results in complex image backgrounds and ambient light. The methods described above, which rely solely on training and optimizing the model with new samples, are simplistic and ineffective in handling situations where device tags are unclear, too small, or blurry. Therefore, this application proposes an image content analysis method to address these issues.

[0066] The image content analysis method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0067] Figure 1 This is a flowchart illustrating an image content analysis method according to an exemplary embodiment, the image content analysis method including the following steps.

[0068] In step S11, the image to be processed is acquired, and the label region of the image to be processed is identified.

[0069] In this step, the device tag portion containing key information, i.e. the tag area, can be accurately located from the original image to be processed, which may contain a lot of irrelevant information.

[0070] First, it can receive input images to be processed, which may be photos of equipment taken in complex environments such as smart data centers or outdoor base stations. Next, using image processing or computer vision techniques (such as object detection algorithms), it searches for and determines the specific location of the device tag in the image, typically outputting it as a rectangular or bounding box, thus precisely limiting the subsequent processing scope to the tag area.

[0071] For example, a pre-trained YOLOv8m object detection model can be used to identify the label regions of the image to be processed. This model can well balance model accuracy and model complexity in the case of label detection.

[0072] For example, the flowchart of the object detection process can be as follows: Figure 2 As shown in the diagram. Before training the model, the images to be processed need to be acquired and preprocessed. The LabelImg tool is used to label the assets, and these labels are used as the ground truth for the training samples, resulting in a training set. In addition, a small number of unlabeled images containing only the devices are added to the training set as background images to reduce false positives.

[0073] When training the model, data augmentation is performed on the training samples. The data augmentation methods used include random brightness, random rotation, random translation, random scaling, random Gaussian blur, random cropping, and random erasure.

[0074] This effectively filters out other background elements in the image that may interfere with OCR recognition, such as the device itself, walls, and other devices, thus laying the foundation for subsequent fine processing and improving overall recognition efficiency and accuracy.

[0075] Furthermore, the label detection model can perform label detection on captured device photos, identify the category of each target in the device photo and the corresponding classification confidence, including targets that may be label regions. At the same time, it can also output the coordinates of the top left and bottom right corners of the label box in the original image to obtain the label position information.

[0076] In step S12, the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region are determined.

[0077] After successfully extracting the label region, issues such as uneven lighting, shadows, and blurred text edges may still exist in the image, resulting in poor OCR recognition performance when performed directly. Therefore, this step introduces an advanced local binarization preprocessing technique. The system analyzes each pixel within the label region, no longer using a globally uniform threshold, but instead examining a small neighborhood (e.g., 8 or 24 neighboring pixels) around the pixel. By calculating the statistical characteristics of the pixel grayscale values ​​within this neighborhood (such as mean, variance, etc.), two thresholds are dynamically calculated for the current pixel: a relatively high "strong threshold" and a relatively low "weak threshold".

[0078] These two thresholds reflect the brightness distribution characteristics of the local area. In other words, the judgment criteria can be adaptively adjusted according to the specific brightness of the label text and its surrounding background at different locations, thereby more accurately distinguishing between text pixels and background pixels. This is especially effective when dealing with situations where there are drastic changes in lighting or uneven contrast between text and background.

[0079] In step S13, based on the corresponding strong threshold and weak threshold, the binarization value of each pixel is determined to obtain the binarized image corresponding to the label region.

[0080] In this step, based on the strong and weak thresholds determined for each pixel, we can determine whether the pixel is displayed as white (representing text) or black (representing background) in the final binarized image. The core idea is: if a pixel's grayscale value is below the strong threshold, it is likely text and is set to white; if it is above the weak threshold, it is likely background and is set to black; pixels in between require more careful handling, possibly considering factors such as their connectivity with confirmed text pixels.

[0081] By employing this refined judgment based on dual thresholds, the system can generate a new image in which the text portion of the label is clearly highlighted (usually pure white), while the background is uniformly set to pure black. This optimized binarized image greatly simplifies the image information, eliminates most of the lighting and background interference, and enables subsequent text detection and recognition algorithms to more easily and accurately locate text lines, segment characters, and ultimately accurately identify the text content on the label.

[0082] In step S14, text detection and recognition are performed on the binarized image to obtain the text content in the image to be processed.

[0083] In the preceding steps, a high-quality binarized image was obtained through precise label region localization and refined binarization based on neighborhood analysis. In this image, the text portion of the label is clearly highlighted in white, while the background is black, and most interfering elements have been effectively removed.

[0084] In this step, text detection and recognition algorithms can be used to perform "text detection" on this binarized image. The algorithm scans the image to find connected regions composed of all white areas that match the shape characteristics of text lines or characters. Next, "text recognition" is performed. For each detected text region (or segmented individual character), the algorithm compares its image features with a built-in character model to determine which character or number each region represents.

[0085] In this way, starting from the initially acquired image to be processed, through a series of precise positioning and image enhancement preprocessing steps, the text information contained in the equipment tag is accurately "read," such as key asset data like equipment number, model, and installation date. This identified text content can then be automatically extracted for updating the asset management system, generating reports, or performing other business operations, thereby achieving automated and high-precision asset tag information collection.

[0086] In one implementation, step S11 involves identifying the label region of the image to be processed, including:

[0087] The image to be processed is input into a pre-trained label detection model to obtain label location information;

[0088] The width and height of the label are calculated based on the label location information to obtain the label area.

[0089] In this implementation, the acquired image to be processed is first input into a pre-trained label detection model. This model is typically trained on a large dataset of labeled telecommunications equipment images. During training, data augmentation techniques such as random brightness adjustment, rotation, and translation are used to effectively improve the recognition capability of label images under complex shooting angles and lighting conditions. It can quickly analyze image content and output the coordinates of the top-left and bottom-right corners of the label bounding box, thereby determining the specific location information of the label in the image. For example, the label detection model can use the YOLOv8m model.

[0090] After obtaining the label location information, the label area can be further optimized. Based on the coordinates of the label frame, the system automatically calculates the width and height of the label and appropriately enlarges it using a specific algorithm. The enlargement process uses the label center point as a reference, proportionally expanding the size of the label frame to ensure that all characters, patterns, and other content on the label are completely covered. This operation not only avoids missing label content due to shooting errors but also provides a sufficient and accurate processing area for subsequent steps such as image binarization and text detection, significantly improving the reliability and accuracy of the entire image content analysis process.

[0091] In one implementation, step S12 involves determining the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region, including:

[0092] Determine the neighborhood of each pixel in the label region;

[0093] In the field of computation, the gray-level weighted average and gray-level weighted variance of pixels are used.

[0094] Calculate the strong threshold corresponding to the neighborhood of each pixel based on the first hyperparameter, the gray-level weighted average value, and the gray-level weighted variance;

[0095] The weak threshold corresponding to the neighborhood of each pixel is calculated based on the second hyperparameter, the gray-level weighted average value, and the gray-level weighted variance.

[0096] In other words, in order to accurately determine the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region, the system first defines the local neighborhood range for each pixel. This neighborhood typically includes pixels within a certain range around the center pixel, and its specific size can be preset according to the image resolution and label features.

[0097] After defining the neighborhood, the system calculates the gray-level weighted average and gray-level weighted variance of all pixels within that neighborhood. The gray-level weighted average reflects the overall brightness level of the pixels within the neighborhood, while the gray-level weighted variance quantifies the dispersion of pixel brightness within the neighborhood, i.e., contrast or texture complexity.

[0098] Based on these two statistics, the system further introduces two hyperparameters—the first hyperparameter is used to adjust the strictness of the strong threshold, and the second hyperparameter is used to adjust the leniency of the weak threshold. The strong threshold is calculated by performing specific mathematical operations (such as linear combination or nonlinear mapping) on ​​the first hyperparameter with the gray-level weighted average and weighted variance; similarly, the weak threshold is calculated by performing operations on the second hyperparameter with the gray-level weighted average and weighted variance.

[0099] This dynamic calculation method ensures that the strong threshold and the weak threshold can be adaptively adjusted according to local image features, so that in subsequent binarization decisions, the character region and the background region can be effectively distinguished, especially in response to interference factors such as uneven illumination or stains.

[0100] For example, the calculation method of the weighted average value is as follows:

[0101] m (x,y) = ∑W (x,y) g (x,y)

[0102] Among them, (x, y) represents the coordinates of the center point in the cropped image. W (x,y) is the weight of the pixel in the neighborhood, and the sum of the weights of the pixels in the neighborhood is 1, and g ( x,y) is the pixel gray value of the center point. For each center point, select the neighborhood with the larger weighted average value as the neighborhood of this point.

[0103] After determining the neighborhood range, the calculation method of the weighted variance is as follows:

[0104]

[0105] The binarization threshold of the center point is:

[0106]

[0107] Among them, k and σ are two hyperparameters. This threshold takes into account both the weighted mean and the weighted variance of the gray level. In order to separate the relatively light character traces caused by wear or light reflection, the present invention uses two sets of hyperparameter configurations for the above calculation formula, corresponding to the strong threshold and the weak threshold respectively. Among them, for the strong threshold, the values of the two hyperparameters are: 0 < k < 0.5, σ = 1. The two hyperparameters of the weak threshold are: 0.5 < k < 1, 1 < σ < 2.

[0108] In one implementation, in step S13, based on the corresponding strong threshold and weak threshold, determining the binarization value of each pixel point includes:

[0109] For each pixel point, if the gray value of this pixel point is lower than the strong threshold, it is determined that the binarization value of this pixel point is 0;

[0110] If the gray value of this pixel point is higher than the weak threshold, it is determined that the binarization value of this pixel point is 255;

[0111] If the gray value of this pixel point is between the strong threshold and the weak threshold, check whether this pixel point is connected to the pixel points whose binarization values have been confirmed to be 0. If it is connected, it is determined that the binarization value of this pixel point is 0, otherwise it is determined that the binarization value of this pixel point is 255.

[0112] In the process of determining the binarization value of a pixel based on strong and weak thresholds, firstly, for each pixel in the label area, the system compares its gray value with the strong threshold: if the gray value is strictly lower than the strong threshold, it is directly determined to be a character pixel, and the binarization value is set to 0 (black). This rule can quickly capture clear character subjects, such as the text area with full ink on the label.

[0113] If a pixel's grayscale value is higher than the weak threshold, it is directly identified as a background pixel and its value is set to 255 (white). This filters out high-brightness background areas, such as the white background of a label or areas with strong light reflection. It is worth noting that the strong threshold and the weak threshold are in a relationship where the strong threshold is less than the weak threshold, forming a transition range between them. This is specifically designed to handle pixels with blurred grayscale values ​​caused by uneven lighting or label wear.

[0114] For pixels with grayscale values ​​between the strong and weak thresholds, the system initiates an eight-connected region search mechanism: centering on the identified character point (value 0), it checks its eight adjacent pixels in the top, bottom, left, right, and diagonal directions. If a pixel in the transition region is connected to the marked character point, it is identified as a character pixel (set to 0); otherwise, it is identified as background (set to 255). This process iterates through multiple rounds until all pixel states are stable, effectively connecting character strokes broken by wear or shadows. For example, it integrates faint character traces with clear strokes into a complete character region, avoiding the character fragmentation problem caused by traditional thresholding methods.

[0115] In one implementation, step S12 involves determining the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region, including:

[0116] Based on the first and second dimensions, determine two candidate neighborhoods for each pixel in the label region;

[0117] Calculate the gray-level weighted average of the pixels in each candidate neighborhood;

[0118] The candidate neighborhood with the larger gray-scale weighted average is used as the neighborhood of each pixel in the label region, and the corresponding strong threshold and weak threshold are calculated.

[0119] When determining the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region, this implementation improves the accuracy of threshold calculation through an adaptive neighborhood selection mechanism. First, the system defines two candidate neighborhoods of different sizes for each pixel: a large neighborhood and a small neighborhood, constructed according to a first size (e.g., D×D) and a second size (e.g., d×d, where D>1.5d), to cover local features at different scales. Then, a Gaussian kernel-weighted average is calculated for the pixel grayscale values ​​within the two candidate neighborhoods. This weighting method uses the distance from the pixel to the center as the basis for weight decay, ensuring that pixels closer to the center contribute more to the mean, thus more accurately reflecting the local grayscale distribution.

[0120] It's understandable that global thresholding binarization methods struggle to convert all characters to a minimum value of 0 due to the complex lighting conditions when photographing labels. Applying global thresholding to images with localized lighting may result in a situation where "no matter what threshold parameter is set, the overall image requirement cannot be met." Traditional local thresholding methods often use the mean of a neighborhood as the threshold for the center point of that region. Due to uneven lighting and label wear, some characters may be faintly visible, while other areas may have a darker white background. In these scenarios, the aforementioned local thresholding methods also fail to binarize the label image effectively.

[0121] Therefore, this invention improves upon the local thresholding method. By comparing the weighted average grayscale values ​​of two candidate neighborhoods, the system determines the neighborhood with the larger average value as the final calculated neighborhood for that pixel. This selection logic is based on the following considerations: a larger weighted average value usually corresponds to a more concentrated grayscale distribution, which can more accurately represent the typical features of the region where the pixel is located, especially suitable for adaptive differentiation of local bright and dark areas in uneven lighting scenes. After determining the neighborhood, the system then calculates a strong threshold and a weak threshold based on the weighted average grayscale value and variance of that neighborhood, combined with two different sets of hyperparameters. This provides a more accurate judgment benchmark for subsequent binarization processing, effectively improving the character and background segmentation accuracy of label images under complex lighting conditions.

[0122] In one implementation, step S14 involves performing text detection and recognition on the binarized image to obtain the text content in the image to be processed, including:

[0123] Convert a binary image into a three-channel image;

[0124] Perform text detection on a three-channel image and output text boxes;

[0125] The text box's sharpness is evaluated. If the sharpness is below a threshold, the text box is super-resolution enhanced to obtain a high-resolution image.

[0126] Text recognition is performed on high-resolution images to obtain the text content in the image to be processed.

[0127] In the process of text detection and recognition of binarized images, the system first converts the binarized black and white image into a three-channel RGB format. The three-channel structure is more compatible with subsequent deep learning-based text detection models, providing a standardized input format for pixel-level text probability prediction and avoiding model inference anomalies caused by channel mismatch.

[0128] After conversion, text detection models such as DBNet (Dilated Backbone Network) can be used to perform text detection on the three-channel image. This model is based on a segmentation algorithm, first outputting a probability map of each pixel belonging to text, then generating text contours after binarization, and finally calculating the minimum bounding rectangle of the contours and appropriately expanding it to form the final text box. Compared with traditional object detection, this segmentation-based detection method is more adaptable to curved text and irregular characters, and is especially suitable for tilted or deformed text that may appear in telecommunications labels.

[0129] After extracting the text box, the system evaluates its sharpness using the Laplace gradient variance algorithm: calculating the gradient distribution variance of the image within the text box; if this value is lower than a preset threshold, it indicates that the image is blurry (e.g., due to camera shake or insufficient resolution). At this point, the AISRNet (Audit Image Super Resolution Net) network is activated for enhancement—this network generates structural features of characters (based on discrete encodings of 6623 commonly used characters) through a structural prior module, combines the font style and position information extracted by the Transformer encoder, aligns the character regions in the super-resolution backbone network using ROIAlign (Region of Interest Alignment) technology, and generates a high-resolution image after AdaIN (Adaptive Instance Normalization) normalization, effectively restoring the stroke details of the blurry characters.

[0130] Finally, the refined text image is input into the text recognition model. The model first extracts character features through convolutional layers, then uses sequence modeling capabilities to parse the character sequence, and finally outputs the text content from the image to be processed. This text content is automatically filled into the corresponding fields in the asset management system, achieving accurate conversion from image to semantics. This process, through a closed-loop design of "format standardization - accurate detection - intelligent enhancement - semantic parsing," comprehensively improves the recognition reliability of telecommunications asset tags under complex lighting and wear scenarios.

[0131] In one implementation, super-resolution enhancement is performed on the text box to obtain a high-resolution image, including:

[0132] Text super-resolution backbone network extracts low-resolution image features of text boxes;

[0133] Adaptive normalization is performed on low-resolution image features and pre-acquired prior features to obtain high-resolution image features;

[0134] High-resolution images are obtained by inverse mapping of high-resolution image features.

[0135] In the process of super-resolution enhancement of text boxes to obtain high-resolution images, the text super-resolution backbone network first extracts features from the low-resolution text box image. It captures basic features such as character edges and contours through multiple convolutional operations, forming a low-resolution feature map. During this process, the network preserves local details of the text, providing foundational data for subsequent enhancement.

[0136] Next, the system fuses the low-resolution features extracted by the backbone network with the prior features generated in advance by the structural prior module. The prior features originate from the improved StyleGAN model (Style Generative Adversarial Network), which generates structural prior information of characters based on the discrete codes of 6623 commonly used characters (such as feature vectors of Chinese characters, letters, and numbers), including prior knowledge such as the stroke direction and corner curvature of the font.

[0137] This invention improves upon the original StyleGAN by removing the constants used as input and replacing them with discrete codes representing different characters. It also removes the hierarchical noise used in StyleGAN for randomly varying fine-grained features. The discrete code for each Chinese character is represented as follows: Where X is the encoding book storing all character features, where each encoding is a learnable feature vector with a feature dimension of 1x1x512. M is the number of discrete encodings in the encoding book, covering commonly used Chinese characters, English letters, and numbers, with a size of 6623. The definition of the structural prior module is as follows:

[0138] I x =B(x,F(z),Φ G )=B(x,w,Φ G )

[0139] In the formula, x is the discrete code of a single character, F represents the network that maps z to the latent space w, and Φ G For the model parameters of StyleGAN, I xThis represents the features of the fused high-resolution image.

[0140] Then, adaptive normalization can be performed on the two types of features: first, the instance mean and variance of the low-resolution features are calculated, and then the statistics of the prior features are embedded into them. This allows the low-resolution features to retain their original content while incorporating prior structural knowledge, thereby generating a high-resolution feature map containing clear structural information. This operation is particularly designed for the complex strokes of Chinese characters, avoiding the problem of detail loss in existing super-resolution algorithms when dealing with unknown fonts.

[0141] Furthermore, a reverse mapping mechanism is used to convert the high-resolution feature map into an image. Specifically, the character bounding box information output by the Transformer encoder is first used, and the enhanced features are aligned to their original positions using the ROIAlign operation. Then, the reverse ROIAlign operation maps the features of each character region to the corresponding positions of the complete text box, ultimately synthesizing a high-resolution image. This process ensures that the enhanced character details and the spatial positions of the original text box are accurately matched, effectively solving the positional misalignment problem of traditional super-resolution algorithms in multi-character scenes. This restores blurry device label text to a clear and recognizable high-resolution image, providing high-quality input for subsequent OCR recognition.

[0142] For example, the super-resolution network proposed in this invention is called AISRNet (Audit Image SuperResolution Net), which consists of three parts: a structure prior module, a Transformer encoder, and a backbone network. Its model structure is shown in the attached figure. Figure 3 As shown in the diagram.

[0143] In one implementation, such as Figure 4 The diagram shown is a flowchart of an image content analysis method provided in an embodiment of this application, which includes the following steps:

[0144] First, prepare an image content detection model based on DBNet and an image content recognition model based on CNN (Convolutional Neural Network) combined with LSTM (Long Short-Term Memory) or Transformer. In addition, two machine learning models are also needed.

[0145] When recognizing the character content in an asset management label, the image file is used for label detection using YoloV8m, and the detected target box is appropriately enlarged according to the width and height of the box.

[0146] Convert the cropped image to grayscale, calculate the weighted mean of two neighborhoods for each pixel, select the neighborhood with the larger mean, and calculate the weighted variance within the neighborhood.

[0147] For each pixel, a strong threshold and a weak threshold are calculated, meaning the image corresponds to two threshold maps: a strong threshold map and a weak threshold map. The binarization result for each pixel is determined step-by-step according to the strategy, and the image is converted into a 3-channel RGB image.

[0148] Image content detection is performed using DBNet to obtain the minimum bounding rectangle containing the text box based on an image segmentation algorithm. This rectangle is then appropriately expanded to form the final text box.

[0149] The expanded text box is cropped from the original image, and its sharpness is evaluated using the Laplace operator. If the sharpness is below a threshold, the text super-resolution network AISRNet is used to restore the image.

[0150] The specific process of image restoration is as follows: Low-resolution blurred images are processed by a text super-resolution backbone network to extract low-resolution image features. Then, the prior features generated by the structure prior module are embedded into the low-resolution characters. Next, the bounding boxes detected by the Transformer module are acquired and aligned to each low-resolution character using the ROIAlign operation. For each low-resolution character, AdaIN is used to normalize the prior distribution. Finally, the inverse RoIAlign is used to map the enhanced features back to their original positions, obtaining a clear text image. Image content recognition is then performed on the final text box content.

[0151] As can be seen from the above, in this embodiment of the application, by introducing local binarization processing based on strong and weak thresholds, the label region is finely preprocessed before text detection and recognition. This can effectively adapt to complex lighting interference, enhance adaptability to label text wear, blurring, or shrinkage, thereby obtaining a clearer and more distinguishable text image before recognition, improving the recognition accuracy in various complex scenarios, and solving the problem of simply relying on model training optimization while ignoring the limitations of image preprocessing in the prior art.

[0152] The image content analysis method provided in this application can be executed by an image content analysis device. This application uses an image content analysis device executing a data storage method as an example to illustrate the apparatus of the image content analysis method provided in this application.

[0153] Figure 5 This is a block diagram of an image content analysis apparatus according to an exemplary embodiment, comprising:

[0154] The acquisition module 301 is used to acquire the image to be processed and identify the label region of the image to be processed;

[0155] The determining module 302 is used to determine the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region;

[0156] The determination module 303 is used to determine the binarization value of each pixel based on the corresponding strong threshold and weak threshold, so as to obtain the binarized image corresponding to the label region.

[0157] The recognition module 304 is used to perform text detection and recognition on the binarized image to obtain the text content in the image to be processed.

[0158] As can be seen from the above, in this embodiment of the application, by introducing local binarization processing based on strong and weak thresholds, the label region is finely preprocessed before text detection and recognition. This can effectively adapt to complex lighting interference, enhance adaptability to label text wear, blurring, or shrinkage, thereby obtaining a clearer and more distinguishable text image before recognition, improving the recognition accuracy in various complex scenarios, and solving the problem of simply relying on model training optimization while ignoring the limitations of image preprocessing in the prior art.

[0159] The image content analysis method provided in this application can be executed by a terminal access terminal. This application uses the terminal access terminal executing the terminal access method as an example to illustrate the apparatus of the image content analysis method provided in this application.

[0160] The image content analysis device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0161] The image content analysis device provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0162] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the above-described image content analysis method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0163] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0164] Figure 7 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0165] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0166] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0167] As can be seen from the above, in this embodiment of the application, by introducing local binarization processing based on strong and weak thresholds, the label region is finely preprocessed before text detection and recognition. This can effectively adapt to complex lighting interference, enhance adaptability to label text wear, blurring, or shrinkage, thereby obtaining a clearer and more distinguishable text image before recognition, improving the recognition accuracy in various complex scenarios, and solving the problem of simply relying on model training optimization while ignoring the limitations of image preprocessing in the prior art.

[0168] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0169] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0170] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0171] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image content analysis method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0172] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0173] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image content analysis method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0174] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0175] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image content analysis method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0176] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0177] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0178] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image content analysis method, characterized in that, The method includes: Acquire the image to be processed and identify the label region of the image to be processed; Determine the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region; Based on the corresponding strong threshold and weak threshold, the binarization value of each pixel is determined to obtain the binarized image corresponding to the label region; Text detection and recognition are performed on the binarized image to obtain the text content in the image to be processed.

2. The image content analysis method according to claim 1, characterized in that, The identification of the label region of the image to be processed includes: The image to be processed is input into a pre-trained label detection model to obtain label location information; The width and height of the label are calculated based on the label location information to obtain the label area.

3. The image content analysis method according to claim 1, characterized in that, Determining the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region includes: Determine the neighborhood of each pixel in the label region; Calculate the gray-level weighted average and gray-level weighted variance of the pixels in the region; Calculate the strong threshold corresponding to the neighborhood of each pixel based on the first hyperparameter, the gray-level weighted average value, and the gray-level weighted variance; The weak threshold corresponding to the neighborhood of each pixel is calculated based on the second hyperparameter, the gray-level weighted average value, and the gray-level weighted variance.

4. The image content analysis method according to claim 1, characterized in that, The step of determining the binarized value of each pixel based on the corresponding strong threshold and weak threshold includes: For each pixel, if the grayscale value of the pixel is lower than the strong threshold, then the binarization value of the pixel is determined to be 0. If the grayscale value of the pixel is higher than the weak threshold, then the binarization value of the pixel is determined to be 255; If the grayscale value of the pixel is between the strong threshold and the weak threshold, then check whether the pixel is connected to a confirmed pixel whose binarization value is 0. If connected, then the binarization value of the pixel is determined to be 0; otherwise, the binarization value of the pixel is determined to be 255.

5. The image content analysis method according to claim 1, characterized in that, Determining the strong and weak thresholds corresponding to the neighborhood of each pixel in the label region includes: Based on the first size and the second size, determine two candidate neighborhoods for each pixel in the label region; Calculate the gray-level weighted average of the pixels in each candidate neighborhood; The candidate neighborhood with the larger gray-scale weighted average is used as the neighborhood of each pixel in the label region, and the corresponding strong threshold and weak threshold are calculated.

6. The image content analysis method according to claim 1, characterized in that, The step of performing text detection and recognition on the binarized image to obtain the text content in the image to be processed includes: The binarized image is converted into a three-channel image; Perform text detection on the three-channel image and output a text box; The clarity of the text box is evaluated. If the clarity is below a threshold, the text box is super-resolution enhanced to obtain a high-resolution image. Text recognition is performed on the high-resolution image to obtain the text content in the image to be processed.

7. The image content analysis method according to claim 6, characterized in that, The step of performing super-resolution enhancement on the text box to obtain a high-resolution image includes: The text super-resolution backbone network extracts low-resolution image features of the text box; The low-resolution image features and the pre-acquired prior features are adaptively normalized to obtain high-resolution image features; The high-resolution image features are reverse-mapped to obtain a high-resolution image.

8. An image content analysis device, characterized in that, The device includes: An acquisition module is used to acquire the image to be processed and identify the label region of the image to be processed; The determination module is used to determine the strong threshold and weak threshold corresponding to the neighborhood of each pixel in the label region; The determination module is used to determine the binarization value of each pixel based on the corresponding strong threshold and weak threshold, so as to obtain the binarized image corresponding to the label region. The recognition module is used to perform text detection and recognition on the binarized image to obtain the text content in the image to be processed.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the image content analysis method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image content analysis method as described in any one of claims 1-7.