Method and device for grouping pictures on large and small number sides based on computer vision
By using the word detection model and the cascade grouping method of normalized similarity calculation in the high-altitude foreign object detection of transmission lines, the problem of structured information errors and insufficient discrimination of feature similarity is solved, and high-precision picture grouping is realized to adapt to automatic grouping under complex working conditions.
Patent Information
- Application Number
- CN202510369110.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
In the detection of high-altitude foreign matter in transmission lines, there are problems such as picture grouping failure and insufficient discrimination of feature similarity caused by structured information errors. In particular, accurate grouping cannot be achieved in mixed watermark scenarios, which affects the accuracy of the positive sample algorithm.
Using a cascading grouping method based on computer vision, the watermark text is identified through the text detection model and combined with normalized similarity calculation, a complementary verification mechanism is formed to ensure that high-precision picture grouping is achieved under complex operating conditions.
The accuracy of image grouping is improved to 99.23%, the robustness and adaptability of the system are enhanced, and the automatic grouping needs are adapted to the mixed watermark scenarios.
Smart Images

Figure CN120296187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image recognition and image processing, and provides a method and device for grouping large and small side pictures based on computer vision. Background Art
[0002] In the field of high-altitude foreign object detection of transmission lines, the positive sample detection algorithm based on computer vision needs to pre-collect the scene base map of each shooting point as a reference, and each point corresponds to a single scene. Usually, there is only one camera for a single device corresponding to each point, and only one scene of data is collected. However, in practical applications, it is found that some devices have two cameras, resulting in the pictures collected containing both "large side" and "small side" scenes. Therefore, it is necessary to distinguish the pictures of the two scenes in this device. Normally, the corresponding side attributes can be configured in the structured information of the pictures of such devices, and the pictures can be directly distinguished according to the structured information of the pictures. However, in the actual process, there are cases where the side attributes configured in the structured information are incorrect, and it is difficult to correct such errors externally in the short term, resulting in the failure of the traditional scheme that relies on structured information grouping. This problem directly affects the accuracy of the positive sample algorithm, and there is an urgent need for a robust automatic grouping method.
[0003] Currently, for such picture grouping problems, the existing technologies mainly include the following two types of solutions:
[0004] 1. Structured information note grouping method: Group directly by reading the pre-device note information of the picture (such as "large side" or "small side" marked by the device source). This method has high efficiency and an accuracy of up to 100% when the information is accurate. However, in the actual scenario, due to the inconsistency between the note information and the picture content, the grouping error rate increases significantly, and it cannot meet the actual needs.
[0005] 2. Positive sample base map similar feature matching method: Match and group by extracting the feature similarity between the picture to be grouped and the reference base map. Although this method does not require an additional grouping module, it relies on the similarity calculation at the feature level and is difficult to distinguish the large and small side pictures with subtle scene differences, resulting in a low grouping accuracy rate.
[0006] The above methods all have obvious defects: the former fails due to incorrect annotation; the latter cannot guarantee the accuracy of scene grouping due to insufficient discrimination granularity of feature similarity. In addition, some pictures have "large side" or "small side" watermark information, but the existing technologies do not effectively utilize such explicit features, nor do they combine multi-modal information (such as text detection and image similarity) to achieve complementary discrimination.
[0007] In view of the above problems, the present invention proposes a cascaded grouping method based on computer vision, which solves the problem of picture grouping in the case of annotation errors by integrating object detection technology and normalized similarity calculation. This method not only uses watermark text detection to achieve explicit classification, but also compensates for missed detections through image similarity comparison, significantly improving the robustness and accuracy of grouping. Experiments show that the grouping accuracy rate of this solution can reach 99.23%, effectively overcoming the limitations of the existing technology. Summary of the Invention
[0008] In the scenario of detecting high-altitude foreign objects on transmission lines, the positive sample detection algorithm needs to establish a normal scenario reference based on the point map. However, for some devices, due to the complex actual working conditions (for example, there are two scenarios on the large-side and small-side at the same point), and the structured annotation information often conflicts with the image watermark content, the existing technical solutions have the following key technical problems:
[0009] Risk of misgrouping caused by dependence on structured information: Existing methods rely on structured information (such as remarks "large side / small side") from device sources or manual annotations for picture grouping. However, in actual scenarios, the annotation information often conflicts with the image watermark text (for example, a picture marked as "large side" actually shows the words "small side"). Directly relying on such incorrect information will lead to completely wrong grouping results.
[0010] Insufficient generalization of feature similarity discrimination: The solution that extracts the point map features through the positive sample algorithm for similarity matching can avoid the problem of dependence on structured information. However, its feature selection and calculation methods lack pertinence (for example, focusing on global features rather than key areas), resulting in a significant decrease in the correct rate of similarity discrimination under conditions where the scene details are less different or there are lighting changes (for example, misjudged as the same group due to background interference).
[0011] Lack of processing for mixed watermark scenarios: Existing technologies have not solved the complex situation where some pictures in the same device have clear watermarks while others do not (for example, the large-side pictures have watermarks while the small-side pictures are not marked). Methods that solely rely on structured information or feature similarity cannot achieve accurate grouping in such mixed scenarios.
[0012] Absence of a cascaded determination mechanism: Existing grouping methods are mostly single technical paths (for example, only relying on annotation or only relying on similarity), lacking a multi-level verification mechanism. As a result, when a certain link fails (such as missing watermark text detection or deviation in similarity calculation), the results cannot be corrected through complementary means, limiting the robustness of the grouping system.
[0013] The above technical problems directly affect the effectiveness of the positive sample algorithm in the transmission line detection system. There is an urgent need for a composite solution that combines semantic recognition and scene feature matching to achieve high-precision picture grouping determination under complex working conditions.
[0014] To achieve the above object, the present invention adopts the following technical solutions:
[0015] The present invention provides a method for grouping large and small size side pictures based on computer vision, comprising the following steps:
[0016] Establish a reference image library of the target device, where the reference image library contains a reference image group of reference large size side images and reference small size side images;
[0017] Obtain the detection image to be grouped and the corresponding reference image group;
[0018] Use an OCR model to identify the watermark text in the detection image;
[0019] When valid watermark text is detected, classify the detection image into the corresponding group according to the watermark content;
[0020] When no valid watermark text is detected: Normalize the detection image and the reference image group to a predetermined size; Calculate the first normalized similarity big_score between the detection image and the reference large size side image, and the second normalized similarity small_score between the detection image and the reference small size side image respectively;
[0021] Complete the scene grouping by comparing the first normalized similarity and the second normalized similarity.
[0022] In the above solution, the reference background image of the large size side scene corresponds to the shooting angle of the first camera of the device, and the reference background image of the small size side scene corresponds to the shooting angle of the second camera of the device.
[0023] In the above solution, the step of establishing the reference image library includes: The selection of reference images follows the following rules:
[0024] When both the images of the large size side scene and the small size side scene of the target device contain watermark text: Select the watermark image with the "large size side" label as the reference large size side image, and select the watermark image with the "small size side" label as the reference small size side image;
[0025] When the image of one side scene contains watermark text:
[0026] Select the image with the "large size side" watermark as the reference large size side image, and select the image of the small size side scene without watermark text as the reference small size side image; or select the image with the "small size side" watermark as the reference small size side image, and select the image of the large size side scene without watermark text as the reference large size side image;
[0027] When there is no watermark text in the scene images on both sides: Select the two images with the largest feature difference between the image of the large-sized side scene and the image of the small-sized side scene as the reference image group.
[0028] In the above solution, the calculation method of the first normalized similarity big_score is as follows:
[0029] Calculate the standard normalized similarity between test_imageM and the base big image base_big_imageM of the large-sized side:
[0030]
[0031] where, T ′ (x,y) and B ′ (x,y) respectively represent the pixel values after image centering;
[0032] The calculation method of the second normalized similarity small_score is as follows:
[0033] The standard normalized similarity between test_imageM and the base small image base_small_imageM of the small-sized side:
[0034]
[0035] where, T(x,y) and S(x,y) respectively represent the pixel values of test_imageM and the base big image base_big_imageM at the coordinate (x,y), and M is the image scaling size.
[0036] In the above solution, the calculation method of the normalized similarity is: complete scene grouping by comparing the first normalized similarity and the second normalized similarity. Specifically: compare the sizes of big_score and small_score. If big_score is greater than or equal to small_score, then the test image belongs to the image of the large-sized side scene; otherwise, the test image belongs to the image of the small-sized side scene.
[0037] The present invention provides a device for grouping large-sized and small-sized side pictures based on computer vision, including the following steps:
[0038] Establishment module: Establish a reference image library of the target device, and the reference image library includes a reference image group of reference large-sized side images and reference small-sized side images;
[0039] Acquisition module: Acquire the test image to be grouped and the corresponding reference image group;
[0040] Text detection module: Use a text detection model to identify the watermark text in the test image;
[0041] Classification module: When a valid watermark text is detected, classify the detected image into the corresponding group according to the watermark content;
[0042] Normalization module: When no valid watermark text is detected: Normalize the detected image and the reference image group to a predetermined size, and calculate the first normalized similarity big_score between the detected image and the reference large-side image and the second normalized similarity small_score between the detected image and the reference small-side image respectively;
[0043] Comparison module: Complete scene grouping by comparing the first normalized similarity and the second normalized similarity.
[0044] In the above device, the steps of establishing the reference image library include: The selection of reference images follows the following rules:
[0045] When the images of the large-side scene and the small-side scene of the target device both contain watermark texts: Select the watermark image with the "large side" logo as the reference large-side image, and select the watermark image with the "small side" logo as the reference small-side image;
[0046] When only one side scene image contains a watermark text:
[0047] Select the image with the "large side" watermark as the reference large-side image, and select the small-side scene image without watermark text as the reference small-side image; or select the image with the "small side" watermark as the reference small-side image, and select the large-side scene image without watermark text as the reference large-side image;
[0048] When neither of the two side scene images contains a watermark text: Select the two images with the largest feature difference between the large-side scene image and the small-side scene image as the reference image group.
[0049] In the above device, the calculation method of the first normalized similarity big_score is:
[0050] Calculate the standard normalized similarity between test_imageM and the large-side base image base_big_imageM:
[0051]
[0052] where, T ′ (x,y) and B ′ (x,y) respectively represent the pixel values after image centering;
[0053] The calculation method of the second normalized similarity small_score is:
[0054] Standard normalized similarity between test_imageM and the small-sized side bottom image base_small_imageM:
[0055]
[0056] Among them, T(x, y) and S(x, y) respectively represent the pixel values of test_imageM and the large-sized side bottom image base_big_imageM at the coordinate (x, y), and M is the image scaling size.
[0057] In the above device, the calculation method of the normalized similarity is as follows: scene grouping is completed by comparing the first normalized similarity and the second normalized similarity. Specifically, compare the sizes of big_score and small_score. If big_score is greater than or equal to small_score, the detected image belongs to the large-sized side scene picture; otherwise, the detected image belongs to the small-sized side scene picture.
[0058] Because the present invention adopts the above technical means, it has the following beneficial effects:
[0059] 1. In the present invention, the large-sized and small-sized side character detection models and the normalized similarity are cascaded and compensated for each other, mainly solving the problem that the large-sized and small-sized sides cannot be correctly distinguished due to incorrect note information. At the same time, it is more targeted at the similar features of the positive sample bottom image. The large-sized and small-sized side character detection models are used to solve the grouping of pictures with characters, and the similarity is used to solve the grouping of pictures without characters. Finally, the correct rate of the large-sized and small-sized side grouping using this solution reaches 99.23%.
[0060] 2. The cascaded compensation mechanism enhances the robustness of the system
[0061] The cascaded design of the front-stage watermark detection and the rear-stage similarity calculation forms a complementary verification mechanism. Even if there is a missed detection in the watermark detection stage (such as blurred or blocked text), the similarity calculation can still correct it further, ensuring the high fault tolerance of the overall grouping process and effectively reducing the risk of overall failure caused by a single technical defect.
[0062] 3. The latter stage discriminates by comparing the normalized similarity between the test image and the two large-sized and small-sized side bottom images after scaling the picture to a certain size, mainly targeting pictures without large-sized and small-sized side characters and detected images that cannot be discriminated in the previous stage.
[0063] 4. Flexible grouping suitable for mixed watermark scenarios
[0064] By constructing a dynamic reference image library (base_big / base_small), it is possible to flexibly configure the reference images according to the actual watermark situation of different devices (such as a mixed scenario with and without watermarks), without manual intervention for correction, achieving fully automatic grouping adaptation under complex watermark conditions. Brief Description of the Drawings
[0065] Figure 1 It is a schematic diagram of the present invention;
[0066] Figure 2 It is a schematic diagram of data annotation. Detailed Embodiments
[0067] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.
[0068] In addition, for a better illustration of the present invention, numerous specific details are given in the following detailed embodiments. Those skilled in the art will understand that the present invention can also be implemented without these specific details.
[0069] The present invention relates to the field of computer vision, and provides a method and device for grouping side pictures of large and small sizes based on computer vision. The main purpose is to solve the problem of mixed picture scenes caused by multi-camera shooting of power transmission equipment, and overcome the grouping failure problems of incorrect dependence on structured information, insufficient discrimination of feature similarity, and blank processing of mixed watermark scenes in the prior art.
[0070] Embodiment 1:
[0071] This embodiment provides a method for grouping side pictures of large and small sizes based on computer vision. Combining with Figure 1 the process shown, it includes the following steps:
[0072] Step 1: Establish a base map comparison library in advance for devices with "large side" and "small side":
[0073] A. Double-sided watermark scenario: If the pictures of the large side scene and the small side scene of the device both have the watermark words "large side" or "small side", then take a "large side" picture and scale it to a fixed size MxM (64x64 in this example) and save it as a picture named base_big.png; take a "small side" picture and scale it to a fixed length and width size MxM (64x64 in this example) and save it as a picture named base_small.png;
[0074] B. Single-sided watermark scenario:
[0075] If in the pictures of the two scenarios of the device, the picture of the large-side scenario has the watermark "large side", and the picture of the small-side scenario does not have the words "side", then take a picture with the watermark "large side" and scale it to a fixed size and save it as a picture named base_big.png, and take a picture without the words "side" and scale it to a fixed size and save it as a picture named base_small.png;
[0076] If in the pictures of the two scenarios of the device, the picture of the small-side scenario has the watermark "small side", and the picture of the large-side scenario does not have the words, then take a picture with the watermark "small side" and scale it to a fixed size and save it as a picture named base_small.png, and take a picture without the words "side" and scale it to a fixed size and save it as a picture named base_big.png;
[0077] C. Scenario without watermark: If in the pictures of the two scenarios of the device, all pictures do not have the words "side", take two pictures of different scenarios as base_big.png and base_small.png respectively. The words "side" refer to the watermark words "large side" or "small side";
[0078] Special note: The preferred size of the reference image is 64×64 pixels (M = 64). During actual implementation, other sizes (such as 32×32 or 128×128) can be adapted, and different-precision calculations can be achieved by adjusting the M value.
[0079] Step 2: Obtain the detection image test_image and the reference image group established for the corresponding device, that is, the base maps base_big.png and base_small.png; among them, base_big_image and base_small_image are loaded from the base map base_big.png of the large-side scenario and the base map base_small.png of the small-side scenario respectively, and the image size is MxM.
[0080] Step 3: Perform detection on the detection image for the words of the large and small sides. If the words "large side" or "small side" are not detected, execute Step 4; if the word "large side" is detected, then this detection image belongs to the large-side picture and jump to Step 8. If the word "small side" is detected, then this detection image belongs to the small-side picture and transfer to Step 8;
[0081] Step 4: Scale the detection image test_image to the size of MxM to obtain test_imageM;
[0082] Step 5: Calculate the standard normalized similarity big_score between test_imageM and the base map base_big_imageM of the large-side scenario:
[0083]
[0084] Among them, T(x, y) and B(x, y) respectively represent the pixel values of test_imageM and the base map base_big_imageM of the large-side scene at the coordinate (x, y). M is the image scaling size (64 in this example). T(x ′ ,y ′ ) and B(x ′ ,y ′ ) respectively represent the pixel values of test_imageM and the base map base_big_imageM of the large-side scene at the coordinate (x ′ ,y ′ );
[0085] Step 6: Calculate the standard normalized similarity small_score between test_imageM and the base map base_small_imageM of the small-side scene. Then
[0086]
[0087] Among them, T(x, y) and S(x, y) respectively represent the pixel values of test_imageM and the base map base_small_imageM of the large-side scene at the coordinate (x, y). M is the image scaling size (64 in this example); T(x ′ ,y ′ ) and S(x ′ ,y ′ ) respectively represent the pixel values of test_imageM and the base map base_small_imageM of the large-side scene at the coordinate (x ′ ,y ′ );
[0088] Step 7: Compare the magnitudes of big_score and small_score. If big_score is greater than or equal to small_score, then the test image belongs to the large-side scene picture; otherwise, the test image belongs to the small-side scene picture.
[0089] Step 8: Output the discrimination result.
[0090] Model training:
[0091] In this solution, a large-small side word detection model needs to be trained, and this model is used to detect the words "large side" and "small side" in the test image.
[0092] 1. Data collection: Collect image data with the words "large side" and "small side" on-site;
[0093] 2. Data annotation: Mark the "large size side" and "small size side" in the form of a rectangular box according to the object detection annotation method, such as Figure 2 shown;
[0094] 3. Model training: Use the annotated data to train the YOLO model as the detection model for the large and small size side words.
[0095] Example 2
[0096] The present invention provides a device that implements the above method for grouping large and small size side pictures based on computer vision.
Claims
1. A method for grouping side pictures of large and small sizes based on computer vision, characterized in that Including the following steps: Establish a reference image library for the target device, where the reference image library contains a reference image group of a reference large-side image and a reference small-side image; Obtain the detection images to be grouped and the corresponding reference image groups; Use a text detection model to identify the watermark text in the detection images; When valid watermark text is detected, classify the detection images into the corresponding groups according to the watermark content; When no valid watermark text is detected: Normalize the detection images and the reference image groups to a predetermined size, and calculate the first normalized similarity big_score between the detection image and the reference large-side image and the second normalized similarity small_score between the detection image and the reference small-side image respectively; Complete the scene grouping by comparing the first normalized similarity and the second normalized similarity.
2. The method according to claim 1, characterized in that, The step of establishing the reference image library includes: The selection of reference images follows the following rules: When the images of the large-side scene and the small-side scene of the target device both contain watermark text: Select the watermark image with the "large side" label as the reference large-side image, and select the watermark image with the "small side" label as the reference small-side image; When the image of a single-side scene contains watermark text: Select the image with the "large side" watermark as the reference large-side image, and select the small-side scene image without watermark text as the reference small-side image; or select the image with the "small side" watermark as the reference small-side image, and select the large-side scene image without watermark text as the reference large-side image; When the images of both side scenes do not contain watermark text: Select the two images with the largest feature difference between the large-side scene image and the small-side scene image as the reference image group.
3. The method according to claim 1, wherein The calculation method of the first normalized similarity big_score is: Calculate the standard normalized similarity between test_imageM and the large-side base image base_big_imageM; Among them, T ′ (x, y) and B ′ (x, y) respectively represent the pixel values after image centering; The calculation method of the second normalized similarity small_score is: The standard normalized similarity between test_imageM and the small-side base image base_small_imageM; Where T(x, y) and S(x, y) respectively represent the pixel values of test_imageM and the large-side base image base_big_imageM at the coordinate (x, y), and M is the image scaling size.
4. The method according to claim 1, wherein The calculation method of the normalized similarity is: Complete the scene grouping by comparing the first normalized similarity and the second normalized similarity. Specifically: Compare the sizes of big_score and small_score. If big_score is greater than or equal to small_score, then the detection image belongs to the large-side scene image, otherwise the detection image belongs to the small-side scene image.
5. A grouping device for side pictures of different sizes based on computer vision, characterized in that, Including the following steps: Establishment module: Establish a reference image library for the target device, where the reference image library contains a reference image group of a reference large-side image and a reference small-side image; Obtaining module: Obtain the detection images to be grouped and the corresponding reference image groups; Text detection module: Use a text detection model to identify the watermark text in the detected image; Classification module: When valid watermark text is detected, classify the detected image into the corresponding group according to the watermark content; Normalization module: When no valid watermark text is detected: Normalize the detected image and the reference image group to a predetermined size, and calculate the first normalized similarity big_score between the detected image and the reference large-side image and the second normalized similarity small_score between the detected image and the reference small-side image respectively; Comparison module: Complete scene grouping by comparing the first normalized similarity and the second normalized similarity.
6. The device according to claim 5, characterized in that, The steps of establishing the reference image library include: The selection of reference images follows the following rules: When the images of the large-side scene and the small-side scene of the target device both contain watermark text: Select the watermark image with the "large-side" logo as the reference large-side image, and select the watermark image with the "small-side" logo as the reference small-side image; When the image of one side scene contains watermark text: Select the image with the "large-side" watermark as the reference large-side image, and select the image of the small-side scene without watermark text as the reference small-side image; or select the image with the "small-side" watermark as the reference small-side image, and select the image of the large-side scene without watermark text as the reference large-side image; When the images of both side scenes do not contain watermark text: Select the two images with the largest feature difference between the image of the large-side scene and the image of the small-side scene as the reference image group.
7. The device according to claim 5, characterized in that, The calculation method of the first normalized similarity big_score is: Calculate the standard normalized similarity between test_imageM and the large-side base image base_big_imageM; Among them, T ′ (x, y) and B ′ (x, y) respectively represent the pixel values after image centering; The calculation method of the second normalized similarity small_score is: The standard normalized similarity between test_imageM and the small-side base image base_small_imageM; Where, T(x, y) and S(x, y) respectively represent the pixel values of test_imageM and the large-side base image base_big_imageM at the coordinate (x, y), and M is the image scaling size.
8. The device according to claim 5, characterized in that, The calculation method of the normalized similarity is: Complete scene grouping by comparing the first normalized similarity and the second normalized similarity. Specifically: Compare the sizes of big_score and small_score. If big_score is greater than or equal to small_score, then the detected image belongs to the image of the large-side scene; otherwise, the detected image belongs to the image of the small-side scene.