Target detection image frame visualization method and device

Through the vue framework, the unified package of fixed coordinates and multi-scene scaling, and the built-in configurable overlapping frame merging logic solves the problems of poor adaptability and insufficient overlapping frame processing of the target detection image visualization method in the prior art, and achieves the consistency of user experience and development efficiency improvement in multiple scenarios.

CN120355565APending Publication Date: 2025-07-22ROCK AI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224284.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the existing object detection image visualization method handles the needs of combining fixed coordinate systems and overlapping boxes, the adaptability is poor, the overlapping boxes are insufficiently processed, and the flexible configuration is lacking. The interaction logic between the thumbnail and the preview mode is logically scattered, resulting in high front-end development and maintenance costs and inconsistent user experience.

Method used

It provides a visualization method and device for object detection image frames. It uniformly packages fixed coordinates and multi-scene scaling through the vue framework, has built-in configurable overlapping frame merging logic, and uses a unified component design to manage coordinate conversion of thumbnails and preview modes to realize automatic merging and display.

Benefits of technology

Automatic mapping of fixed coordinates and synchronization of multi-scene scaling, automatic merging of overlapping boxes reduces front-end development work, improves user experience consistency and readability of labels, and reduces development and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355565A_ABST
    Figure CN120355565A_ABST
Patent Text Reader

Abstract

The invention provides a target detection image frame visualization method and device, and solves the problems in the prior art. The method specifically comprises the following steps: S1, obtaining an image, a scaling, a target detection frame coordinate, a fixed coordinate system with a size of A * B and a current display scene; s2, performing coordinate transformation according to the image, the image parameters, the scaling and the coordinates of the target detection frame to form transformed coordinates; s3, merging the target detection frames meeting the merging condition in the image to form a merged target detection frame set; and S4, according to the combined target detection frame set, displaying a thumbnail and / or a preview of the image, thereby improving the adaptability of a coordinate system, and enabling service configuration to be more flexible and convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and in particular, to a method and device for visualizing object detection image frames. Background Art

[0002] Visualization of fixed coordinate object detection image frames means that in an object detection task, the position and category information of the detected target objects in the image are displayed in a visual way. Usually, an object detection algorithm outputs the bounding boxes and category labels of the target objects, and the visualization of fixed coordinate object detection is to draw this information on the image for intuitive observation of the detection results. The common practice is usually:

[0003] (1) Pixel-level coordinates: Corresponding to the true resolution of the image, the front-end only needs to draw the box using absolute positioning or Canvas; Normalized coordinates: Represent the relative position within the range of 0 to 1, and the front-end needs to perform reverse calculation according to the width and height of the image; (2) Front-end scaling and preview: To adapt to different screens, the front-end will scale the image by a certain ratio and maintain the relative position and size of the detection box. Common human-computer interaction methods for object detection image visualization include clicking on a thumbnail to view the large image, magnifying glass, full-screen preview, etc. However, in some business scenarios, the model or the backend output is fixed in a virtual coordinate system (such as 1000×1000), while the actual resolution of the picture itself varies. At this time, the front-end requires additional conversion logic to correctly render the detection box onto the image. At the same time, the problem of multiple overlapping boxes also appears in some scenarios, posing challenges to the readability of the front-end interface and subsequent processing.

[0004] Existing object detection image visualization methods have the following advantages: (1) The visualization effect is relatively intuitive: Whether using pixel-level coordinates or normalized coordinates, most front-end drawing frameworks or components can quickly superimpose the detection results on the image, and the overall process is simple and easy to understand; (2) The interaction method is mature: Currently, there are a large number of mature solutions supporting interaction functions such as click to zoom in, full-screen preview, and hover hint, which can meet the daily needs of viewing detection results.

[0005] However, the disadvantages of the existing technology are also relatively obvious, specifically including: (1) Poor adaptability of the fixed coordinate system: When the detection system uniformly outputs coordinates of 1000×1000, the front end must manually implement coordinate conversion. If not handled properly, it will lead to misalignment or scaling imbalance of the detection box. Traditional pixel-level coordinate components or normalized coordinate components are not generally applicable to such fixed coordinate outputs; (2) Insufficient handling of overlapping boxes: In many scenarios, there will be multiple detection boxes of the same target overlapping each other. Conventional visualization components simply draw all the boxes, causing chaos in the interface and lacking the function of automatically merging or filtering overlapping boxes with the same target name. It is difficult for users to distinguish the target boundaries when viewing, and it is also not conducive to secondary analysis or marking; (3) Lack of flexible configuration: For complex business requirements, the front end needs to filter or merge according to multiple dimensions such as target name, overlapping area, and even confidence. Existing technologies often lack configurable strategies and switches, resulting in developers having to perform secondary encapsulation by themselves, increasing the maintenance cost; (4) Scattered visualization and interaction logic: Some solutions do not have unified coordinate management for the switching between thumbnails and enlarged previews, which is prone to marker offset or misalignment of positions, and it is also impossible to maintain a consistent user experience on various screen sizes.

[0006] Based on the above deficiencies, it can be seen that the existing technology has not yet formed a mature and easily extensible front-end component-based solution for handling the fixed coordinate system and the need for merging overlapping boxes, and there is an urgent need for innovation and improvement. Summary of the Invention

[0007] The present invention provides a method and device for visualizing target detection image frames to solve the above problems.

[0008] In the first aspect, the present invention provides a method for visualizing target detection image frames, specifically including the following steps:

[0009] Step S1, obtain an image, a scaling ratio, target detection box coordinates, a fixed coordinate system of size A×B, and the current display scene;

[0010] Step S2, perform coordinate conversion according to the image, image parameters, scaling ratio, and target detection box coordinates to form converted coordinates;

[0011] Step S3, merge the target detection boxes in the image that meet the merging conditions to form a set of target detection boxes after merging processing;

[0012] Step S4, draw and display the thumbnail and / or preview image of the image according to the set of target detection boxes after merging processing.

[0013] Wherein, both A and B are positive integers.

[0014] Preferably, in step S1, the display scene includes a thumbnail display scene and a preview display scene.

[0015] Preferably, in step S2, the target detection boxes with the same target name and overlapping in the image are merged to form a merged target detection box, which specifically includes the following steps:

[0016] Step S201: Obtain the actual width W iamge and actual height H image ;

[0017] Step S202: Calculate the image scaling ratio scale under the current display scene according to the current display scene;

[0018] Step S203: Convert the coordinates of the target detection box in the image to form the actual display coordinates of the target detection;

[0019] Step S204: Label the target detection box in the image container.

[0020] Among them, W iamge and H image are both positive integers.

[0021] Preferably, in step S203, when converting the coordinates of the target detection box in the image, the conversion formula is as follows:

[0022]

[0023] Among them, X model represents the original abscissa of the target detection box; X display represents the abscissa after the conversion of the target detection box coordinates; Y model represents the original ordinate of the target detection box; Y display represents the ordinate after the conversion of the target detection box coordinates; X model , X display , Y model , Y display are all positive integers.

[0024] Preferably, in step S204, in the image container, the target detection box is labeled by absolute positioning or a similar rendering mechanism.

[0025] Preferably, in step S3, the target detection boxes that meet the merging conditions in the image are merged to form a set of merged target detection boxes, which specifically includes the following steps:

[0026] Step S301: Determine whether the target detection boxes in the image meet the merging conditions;

[0027] Step S302: When two target detection boxes meet the merging conditions, merge the two target detection boxes to form a merged target detection box.

[0028] Step S303: Repeat Step S301 - Step S302 until all target detection boxes that meet the merging conditions in the image have been merged, and output the set of merged target detection boxes.

[0029] Preferably, in Step S301, the merging conditions include one or more of the following three conditions:

[0030] a) The two target detection boxes overlap and have the same target name;

[0031] b) The two target detection boxes overlap and the intersection - over - union ratio of the two target detection boxes is greater than the first threshold;

[0032] c) The two target detection boxes overlap and the confidence levels of the two target detection boxes are greater than the second threshold.

[0033] Preferably, in Step S302, when two target detection boxes that meet the merging conditions are merged, the merging formula is specifically as follows:

[0034] x1 = min{x f1 ,x s1}

[0035] y1 = min{y f1 ,y s1}

[0036] x2 = max{x f2 ,x s2}

[0037] y2 = max{y f2 ,y s2}

[0038] Where x1 < x2, y1 < y2; (x1, y1) represents the coordinates of the first vertex of the merged target detection box; (x2, y2) represents the coordinates of the second vertex of the merged target detection box (in this application, the "second vertex" represents the vertex opposite to the first vertex); (x f1 ,y f1 ) represents a rectangular box; represents the coordinates of the first vertex of the first target detection box to be merged; (x f2 ,y f2 ) represents the second vertex of the first target detection box to be merged; (x s1 ,y s1 ) represents the coordinates of the first vertex of the second target detection box to be merged; (xs2 , y s2 ) represents the second vertex of the second target detection box to be merged.

[0039] Preferably, in step S4, according to the set of target detection boxes after the merging process, a thumbnail and / or a preview image of the image is drawn and displayed, specifically including:

[0040] When displaying the thumbnail of the image, it specifically includes the following steps:

[0041] Step1. Obtain the maximum display width and height of the thumbnail, and calculate the thumbnail scaling ratio;

[0042] Step2. According to the image, the maximum display width and height of the thumbnail, and the thumbnail scaling ratio, perform coordinate transformation on all pixel point coordinates in the image to form the transformed thumbnail pixel point coordinates;

[0043] Step3. Draw and display the thumbnail according to the thumbnail pixel point coordinates;

[0044] When displaying the preview image of the image, it specifically includes the following steps:

[0045] Step4. Obtain the maximum display width and height of the preview image, and calculate the preview image scaling ratio;

[0046] Step5. According to the image, the maximum wire harness width and height of the preview image, and the preview image scaling ratio, perform coordinate transformation on all pixel point coordinates in the image to form the transformed preview image pixel point coordinates;

[0047] Step6. Draw and display the preview image according to the preview image pixel point coordinates.

[0048] Preferably, in Step3 and Step6, according to HTML / CSS, <canvas>Draw thumbnail or preview images in a way similar to that in WebGL.

[0049] Preferably, the steps S1 - S4 are encapsulated by the reactive function and ref function of the vue framework (in this application, "encapsulation" means bundling data and the methods for operating on this data together, and hiding internal implementation details as much as possible, only exposing necessary interfaces for external use) to form a vue component.

[0050] In a second aspect, the present invention also provides a target detection image box visualization device, which specifically includes the following modules:

[0051] A data acquisition module, used to acquire an image, a scaling ratio, target detection box coordinates, a fixed coordinate system of size A×B, and the current display scene;

[0052] A target detection box coordinate conversion module, used to perform coordinate conversion according to the image, image parameters, scaling ratio, and target detection box coordinates to form converted coordinates;

[0053] A target detection box merging module, used to merge target detection boxes that meet the merging conditions in the image to form a set of target detection boxes after merging processing;

[0054] An image display module, used to draw and display the thumbnail and / or preview image of the image according to the set of target detection boxes after merging processing.

[0055] Wherein, both A and B are positive integers.

[0056] Preferably, in the data acquisition module, the display scene includes a thumbnail display scene and a preview image display scene.

[0057] Preferably, the target detection box coordinate conversion module specifically includes the following sub - modules:

[0058] A first coordinate conversion module, used to obtain the actual width W iamge and actual height H image ;

[0059] A second coordinate conversion module, used to calculate the image scaling ratio scale in the current display scene according to the current display scene;

[0060] A third coordinate conversion module, used to convert the target detection box coordinates in the image to form the actual display coordinates of the target detection;

[0061] A fourth coordinate conversion module, used to label the target detection box in the image container.

[0062] Wherein, W iamge and H image are all positive integers.

[0063] Preferably, in the third coordinate conversion module, the coordinates of the target detection box in the image are converted, and the conversion formula is as follows:

[0064]

[0065] where X model represents the original abscissa of the target detection box; X display represents the abscissa of the target detection box after coordinate conversion; Y model represents the original ordinate of the target detection box; Y display represents the ordinate of the target detection box after coordinate conversion; X model and X display and Y model and Y display are all positive integers.

[0066] Preferably, in the fourth coordinate conversion module, in the image container, the target detection box is labeled by absolute positioning or a similar rendering mechanism.

[0067] Preferably, the target detection box merging module specifically includes the following sub-modules:

[0068] The first merging sub-module is used to determine whether the target detection boxes in the image meet the merging conditions;

[0069] The second merging sub-module is used to merge the two target detection boxes when the two target detection boxes meet the merging conditions to form a merged target detection box;

[0070] The third merging sub-module is used to repeatedly execute the first merging sub-module and the second merging sub-module until all the target detection boxes in the image that meet the merging conditions have been merged, and output the set of merged target detection boxes.

[0071] Preferably, in the first merging sub-module, the merging conditions include one or more of the following three conditions;

[0072] a) The two target detection boxes overlap and have the same target name;

[0073] b) The two target detection boxes overlap and the intersection over union of the two target detection boxes is greater than the first threshold;

[0074] c) The two target detection boxes overlap and the confidence levels of the two target detection boxes are greater than the second threshold.

[0075] Preferably, in the second merging sub-module, when two target detection boxes meeting the merging conditions are merged, the merging formula is specifically as follows:

[0076] x1 = min{x f1 ,x s1}

[0077] y1 = min{y f1 ,y s1}

[0078] x2 = max{x f2 ,x s2}

[0079] y2 = max{y f2 ,y s2}

[0080] Wherein, x1 < x2, y1 < y2; (x1, y1) represents the coordinates of the first vertex of the merged target detection box; (x2, y2) represents the coordinates of the second vertex of the merged target detection box (in this application, the "second vertex" represents the vertex opposite to the first vertex); (x f1 ,y f1 ) represents a rectangular box; represents the coordinates of the first vertex of the first target detection box to be merged; (x f2 ,y f2 ) represents the second vertex of the first target detection box to be merged; (x s1 ,y s1 ) represents the coordinates of the first vertex of the second target detection box to be merged; (x s2 ,y s2 ) represents the second vertex of the second target detection box to be merged.

[0081] Preferably, the image display module specifically includes the following sub-modules:

[0082] The first image display sub-module is used to display the thumbnail of the image;

[0083] The second image display sub-module is used to display the preview image of the image.

[0084] Preferably, the first image display sub-module specifically includes the following units:

[0085] The first unit is used to obtain the maximum display width and height of the thumbnail and calculate the thumbnail scaling ratio;

[0086] A second unit, configured to perform coordinate transformation on all pixel point coordinates in the image according to the image, the maximum display width and height of the thumbnail, and the thumbnail scaling ratio, to form transformed thumbnail pixel point coordinates;

[0087] A third unit, configured to draw and display the thumbnail according to the thumbnail pixel point coordinates.

[0088] Preferably, the second image display sub-module specifically includes the following units:

[0089] A fourth unit, configured to obtain the maximum display width and height of the preview image, and calculate the preview image scaling ratio;

[0090] A fifth unit, configured to perform coordinate transformation on all pixel point coordinates in the image according to the image, the maximum wire harness width and height of the preview image, and the preview image scaling ratio, to form transformed preview image pixel point coordinates;

[0091] A sixth unit, configured to draw and display the preview image according to the preview image pixel point coordinates.

[0092] Preferably, in the third unit and the sixth unit, according to HTML / CSS, <canvas>Draw thumbnail or preview images in a way similar to that in WebGL.

[0093] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for visualizing object detection image frames described in any one of the first aspects of the present application.

[0094] In a fourth aspect, the present invention also provides an electronic device, which includes: a memory storing a computer program; a processor communicatively connected to the memory, and when calling the computer program, executes a method for visualizing object detection image frames described in any one of the first aspects of the present application.

[0095] Compared with the prior art, the present invention has the following obvious prominent substantive features and remarkable advantages:

[0096] The present invention provides a method and device for visualizing object detection image frames, which solves the problems existing in the prior art. It has the following advantages:

[0097] (1) Automatic mapping of fixed coordinates and synchronization of multi-scene scaling

[0098] Combining "fixed coordinates", "actual image pixels", and "multi-scene scaling" in the vue framework for unified encapsulation and management, greatly simplifies the drawbacks of multiple manual calculations in the prior art. Without changing the backend coordinate output, it provides a configurable and integrated coordinate conversion process for multi-scene visualization on the front end.

[0099] (2) Automatically merge overlapping boxes with the same target name that can be configured

[0100] For business scenarios with frequent occurrence of duplicate / overlapping boxes (such as security monitoring, industrial quality inspection, autonomous driving, etc.), the present invention has built-in switchable overlapping merge logic at the component layer, which combines a large number of originally scattered small boxes into a larger box; at the same time, information such as names and colors is retained for visualization use, giving the front end the ability to perform secondary optimization on the detection boxes during the drawing stage, significantly improving the readability and reliability of the annotations.

[0101] (3) Unified management of thumbnail and preview modes

[0102] Compared with the traditional way of separating "thumbnail" and "large image preview" and performing local repeated calculations, the present invention adopts a unified component-based design. When entering the preview mode, the same set of coordinate conversion and overlapping merge logic is reused. By simply switching to different zoom ratios, perfect consistency of the annotation positions can be maintained, reducing the workload of secondary encapsulation by developers, reducing risks such as position offset and misalignment of the annotation boxes, and maintaining a consistent user experience at multiple front-end resolutions. Description of the Drawings

[0103] The accompanying drawings that form a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0104] Figure 1 It is a flowchart of a method for visualizing target detection image frames in a preferred embodiment of the present invention.

[0105] Figure 2 It is a flowchart of an example for visualizing target detection image frames in an industrial quality inspection system scenario in a preferred embodiment of the present invention.

[0106] Figure 3 It is a picture before merging target detection image frames in a preferred embodiment of the present invention.

[0107] Figure 4 It is a picture after merging target detection image frames in a preferred embodiment of the present invention. Detailed implementation manners

[0108] The present invention provides a method and device for visualizing target detection image frames. To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0109] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0110] Example 1:

[0111] As Figure 1 shown, a method for visualizing target detection image frames described in this embodiment specifically includes the following steps:

[0112] Step S1, obtain an image, a scaling ratio, target detection box coordinates, a fixed coordinate system of size A×B, and the current display scene; wherein, the display scene includes a thumbnail display scene and a preview display scene, and both A and B are positive integers.

[0113] Step S2: Perform coordinate transformation based on the image, image parameters, scaling ratio, and target detection box coordinates to form the transformed coordinates.

[0114] In the specific implementation of this embodiment, step S2 specifically includes the following steps:

[0115] Step S201: Obtain the actual width W iamge and actual height H image of the image; where W iamge , H image are both positive integers.

[0116] Step S202: Calculate the image scaling ratio scale under the current display scenario.

[0117] Step S203: Transform the target detection box coordinates in the image to form the actual display coordinates of the target detection; among them, when transforming the target detection box coordinates in the image, the transformation formula is as follows:

[0118]

[0119] where X model represents the original abscissa of the target detection box; X display represents the abscissa after the target detection box coordinate transformation; Y model represents the original ordinate of the target detection box; Y display represents the ordinate after the target detection box coordinate transformation; X model , X display , Y model , Y display are all positive integers.

[0120] Step S204: Mark the target detection box in the image container; among them, in the image container, mark the target detection box through absolute positioning or a similar rendering mechanism.

[0121] Step S3: Merge the target detection boxes in the image that meet the merging conditions to form a set of merged target detection boxes.

[0122] In the specific implementation of this embodiment, step S3 specifically includes the following steps:

[0123] Step S301: Determine whether the target detection boxes in the image meet the merging conditions.

[0124] Optionally, in step S301, the merging conditions include one or more of the following three conditions;

[0125] a) The two object detection frames overlap and have the same object name; b) The two object detection frames overlap and the intersection over union (IoU; IoU is an algorithm for calculating the overlapping ratio of different images) of the two object detection frames is greater than the first threshold; c) The two object detection frames overlap and the confidence levels of the two object detection frames are greater than the second threshold. The above not only includes merging the object detection frames with the same name, but also can merge or filter different objects with a high degree of overlap. Sometimes, the object names in the same area may be incorrect or there may be conflicts between similar categories (such as "car" and "truck"). In this scenario, the annotations of different categories with a high IoU can be merged into one.

[0126] Step S302, when the two object detection frames meet the merging conditions, merge the two object detection frames to form a merged object detection frame.

[0127] Among them, when merging the two object detection frames that meet the merging conditions, the merging formula is specifically as follows:

[0128] x1 = min{x f1 , x s1}

[0129] y1 = min{y f1 , y s1}

[0130] x2 = max{x f2 , x s2}

[0131] y2 = max{y f2 , y s2}

[0132] Among them, x1 < x2, y1 < y2; (x1, y1) represents the coordinates of the first vertex of the merged object detection frame; (x2, y2) represents the coordinates of the second vertex of the merged object detection frame (in this application, the "second vertex" represents the vertex opposite to the first vertex); (x f1 , y f1 ) represents a rectangular frame; represents the coordinates of the first vertex of the first object detection frame to be merged; (x f2 , y f2 ) represents the second vertex of the first object detection frame to be merged; (x s1 , y s1 ) represents the coordinates of the first vertex of the second object detection frame to be merged; (x s2 , y s2 ) represents the second vertex of the second object detection frame to be merged.

[0133] Step S303: Repeat Step S301 - Step S302 until all target detection boxes in the image that meet the merging conditions have been merged, and output the set of target detection boxes after the merging process.

[0134] Step S4: Draw and display the thumbnail and / or preview image of the image based on the set of target detection boxes after the merging process.

[0135] Among them, when displaying the thumbnail of the image, it specifically includes the following steps:

[0136] Step1: Obtain the maximum display width and height of the thumbnail, and calculate the thumbnail scaling ratio; Step2: According to the image, the maximum display width and height of the thumbnail, and the thumbnail scaling ratio, perform coordinate transformation on all pixel point coordinates in the image to form the transformed thumbnail pixel point coordinates; Step3: Draw and display the thumbnail based on the thumbnail pixel point coordinates.

[0137] Among them, when displaying the preview image of the image, it specifically includes the following steps:

[0138] Step4: Obtain the maximum display width and height of the preview image, and calculate the preview image scaling ratio; Step5: According to the image, the maximum wire harness width and height of the preview image, and the preview image scaling ratio, perform coordinate transformation on all pixel point coordinates in the image to form the transformed preview image pixel point coordinates; Step6: Draw and display the preview image based on the preview image pixel point coordinates.

[0139] In the specific implementation of this embodiment, in Step3 and Step6, based on HTML / CSS <canvas>Draw thumbnails or preview images in a way similar to that in WebGL. When a small number of preview images or thumbnails need to be drawn, use HTML / CSS for drawing. When preview images or thumbnails with a large number of object detection boxes need to be drawn (for example, in high-resolution remote sensing images, there are thousands or tens of thousands of object detection boxes, such as crop plots, building boundaries, etc.), use <canvas>Or perform batch rendering using WebGL, with high processing efficiency.

[0140] In addition, in this embodiment, the steps S1 - S4 are encapsulated by using the reactive function and ref function of the vue framework to form a vue component. It has the following advantages: (1) Information hiding: By restricting access to class members (properties or methods), it prevents external code from directly accessing or modifying the internal state of the object; (2) Improved security: Since external code cannot directly access the internal properties of the object, this reduces errors or vulnerabilities caused by misusing or abusing these properties; (3) Increased flexibility and maintainability: Developers can freely change the internal implementation of the class without changing the external interface, making the code easier to maintain and expand; (4) Defined clear interfaces: Encapsulation encourages defining a set of clear operation interfaces for each class, that is, public methods, which are the only way for the outside world to interact with the object. Such a design helps to reduce the coupling degree between system components and improve the modularity.

[0141] In the specific implementation of this embodiment, the merged target detection box can be used to select a specific target. For example, in a certain industrial production line, there are dozens of detected target categories, but the operator may only care about two types of defects, "cracks" and "dents". It is possible to only display the defect marks of the "crack" category, which is more intuitive visually, enabling the user to focus on the detection boxes of specific categories and reducing visual interference.

[0142] Example 2:

[0143] Such as Figure 2 shown, taking an industrial quality inspection system scenario as an example, a method for visualizing target detection images according to the present invention is used to detect the defects on the product surface and output the defect coordinates.

[0144] 1. Receive the picture imageData (the picture source includes: picture url and BASE64) (one picture includes one or more target detection boxes) and targets (detection results, that is, target detection boxes with annotations).

[0145] 2. Load the picture and obtain the picture size.

[0146] In an industrial quality inspection system, the resolution of a certain part photo is 3000×2000 pixels. Through the front - end script or browser API, the originalWidth = 3000 and originalHeight = 2000 of this photo are obtained. Bind all subsequent drawing operations to this real size without changing the backend output method.

[0147] 3. Determine whether overlapping merging needs to be performed by checking mergeEnabled (used as a configuration switch to control whether to execute). If it is required, perform the merge; otherwise, keep it unchanged.

[0148] After the system configuration sets mergeEnabled = true (mergeEnabled is used as a configuration switch to control whether to perform overlapping merging, "true" means execute, "false" means not execute), if there are two boxes (120, 200, 150, 230) and (125, 205, 145, 225) with the same "scratch" target name overlapping, they will be merged into a larger box. For example, (125, 205, 150, 230). Automatically judge and merge the target detection image boxes at the front end to avoid repeated drawing and improve readability.

[0149] 4. Calculate the scaling ratio scale of the display scene (including the preview display scene and the zoom display scene).

[0150] 5. For each target detection box in the picture, convert it using the coordinate transformation formula.

[0151] 6. In the image container or <canvas>Draw rectangles and text annotations on it.

[0152] 7. Repeat steps 5 - 6 until all the object detection boxes in the picture are processed.

[0153] 8. Display the thumbnail image and provide a preview function. After the user clicks on the thumbnail, enter the preview scene mode.

[0154] (1) Set the maximum width of the thumbnail to 600 pixels and calculate the thumbnail scaling ratio: scale = 600 / 3000 = 0.2; (2) Perform coordinate transformation on the merged boxes: for example, the position of the coordinate (120, 200, 150, 230) finally displayed on the thumbnail is (120 / 1000 × 3000 × 0.2, 200 / 1000 × 2000 × 0.2), etc. There is no need to manually write multiple formulas, and all mapping logics are uniformly processed within the component to reduce the error rate.

[0155] When the user clicks on an area with more defects, switch to the preview mode. The component automatically calculates the previewScale (preview range). For example, if the available space on the user's screen only allows a maximum area of 1500 × 1000, the scaling ratio in the preview mode may be 0.5 (based on the width). Perform coordinate transformation on all the object detection image boxes again and render them in the visible area of 1500 × 1000 to form a preview view with high - resolution details; the user can observe the exact position and size of the defects. The same coordinate transformation method is used for both thumbnail and preview, rather than independent processing; achieving complete consistency of coordinates and boxes.

[0156] 9. Redraw the object detection boxes on the preview image.

[0157] Finally, it can be seen in the preview image that a continuous area of scratches is only shown as a larger merged object detection image box on the front - end interface, reducing the mutual occlusion and the impact of duplicate boxes between the image boxes. The user can freely turn on or off the merging function to view more detailed or more overview information.

[0158] 10. After the display is finished, the user clicks to close the picture display.

[0159] Example 3:

[0160] Apply the object detection image box visualization method described in the present invention to the target search of an unmanned aerial vehicle. The unmanned aerial vehicle searches for specific targets and shows the marked target positions after taking pictures. Input a picture of a human scene and input "Detect the position of the girls in the picture", and output the object image box selection result, as Figure 3 shown. Figure 3 Two girls in the original picture of the character scene are respectively framed, and a target detection image frame with an overlapping relationship is generated. Through a target detection image frame visualization method described in the present invention, the target detection image frame shown in Figure 4 is generated to frame the two girls simultaneously.

[0161] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.< / canvas> < / canvas> < / canvas> < / canvas> < / canvas>

Claims

1. A method for visualizing object detection image frames, characterized in that, Specifically, it includes the following steps: Step S1, obtain an image, a scaling ratio, target detection box coordinates, a fixed coordinate system of size A×B, and the current display scene; Among them, the display scene includes a thumbnail display scene and a preview display scene; Step S2, perform coordinate conversion according to the image, image parameters, scaling ratio, and target detection box coordinates to form converted coordinates; Step S3, merge the target detection boxes in the image that meet the merging conditions to form a set of target detection boxes after merging processing; Step S4, draw and display the thumbnail and / or preview of the image according to the set of target detection boxes after merging processing; Among them, both A and B are positive integers.

2. The object detection image box visualization method according to claim 1, wherein In step S2, merge the target detection boxes in the image with the same target name and overlapping to form a merged target detection box. Specifically, it includes the following steps: Step S201: Obtain the actual width W and actual height H of the image according to the image iamge and actual height H image ; Step S202, calculate the image scaling ratio scale in the current display scene according to the current display scene; Step S203, convert the target detection box coordinates in the image to form the actual display coordinates of the target detection; Step S204, label the target detection boxes in the image container; Among them, W iamge , H image are all positive integers.

3. The method for visualizing a target detection image frame according to claim 2, characterized in that In step S203, when converting the target detection box coordinates in the image, the conversion formula is as follows: Among them, X model represents the original abscissa of the target detection box; X display represents the abscissa of the target detection box after coordinate transformation; Y model represents the original ordinate of the target detection box; Y display represents the ordinate of the target detection box after coordinate transformation; X model , X display , Y model , Y display are all positive integers.

4. A method for visualizing object detection image frames according to claim 2, characterized in that, In step S204, in the image container, label the target detection boxes through absolute positioning or a similar rendering mechanism.

5. A method for visualizing object detection image frames according to claim 1, characterized in that In step S3, merge the target detection boxes in the image that meet the merging conditions to form a set of target detection boxes after merging processing. Specifically, it includes the following steps: Step S301, determine whether the target detection boxes in the image meet the merging conditions; Step S302, when two target detection boxes meet the merging conditions, merge the two target detection boxes to form a merged target detection box; Step S303, repeat steps S301 - S302 until all the target detection boxes in the image that meet the merging conditions have been merged, and output the set of target detection boxes after merging processing.

6. A method for visualizing object detection image frames according to claim 5, characterized in that, In step S301, the merging conditions include one or more of the following three conditions; a) Two target detection boxes overlap and have the same target name; b) Two target detection boxes overlap and the intersection - over - union ratio of the two target detection boxes is greater than the first threshold; c) Two target detection boxes overlap and the confidence levels of the two target detection boxes are greater than the second threshold.

7. A method for visualizing object detection image frames according to claim 5, characterized in that In step S302, when two target detection boxes that meet the merging conditions are merged, the merging formula is specifically as follows: x1 = min{x f1 , x s1} y1 = min{y f1 , y s1} x2 = max{x f2 , x s2} y2 = max{y f2 , y s2} Among them, x1 < x2 and y1 < y2; (x1, y1) represents the coordinates of the first vertex of the merged target detection box; (x2, y2) represents the coordinates of the second vertex of the merged target detection box (in this application, the "second vertex" refers to the vertex opposite to the first vertex); (x f1 , y f1 ) represents a rectangular box; represents the coordinates of the first vertex of the first target detection box to be merged; (x f2 , y f2 ) represents the second vertex of the first target detection box to be merged; (x s1 , y s1 ) represents the coordinates of the first vertex of the second target detection box to be merged; (x s2 , y s2 ) represents the second vertex of the second target detection box to be merged.

8. A method for visualizing object detection image frames according to claim 1, characterized in that In step S4, draw and display the thumbnail and / or preview of the image according to the set of target detection boxes after merging processing. Specifically, it includes: When displaying the thumbnail of the image, it specifically includes the following steps: Step1, obtain the maximum display width and height of the thumbnail, and calculate the thumbnail scaling ratio; Step2, perform coordinate conversion on all pixel point coordinates in the image according to the image, the maximum display width and height of the thumbnail, and the thumbnail scaling ratio to form the converted thumbnail pixel point coordinates; Step 3. Draw and display the thumbnail according to the pixel coordinates of the thumbnail. When displaying the preview image of the image, it specifically includes the following steps: Step 4. Obtain the maximum display width and height of the preview image, and calculate the preview image scaling ratio. Step 5. Perform coordinate transformation on all pixel coordinates in the image according to the image, the maximum beam width and height of the preview image, and the preview image scaling ratio, to form the pixel coordinates of the transformed preview image. Step 6. Draw and display the preview image according to the pixel coordinates of the preview image.

9. The object detection image box visualization method according to claim 8, wherein In Step3 and Step6, according to HTML / CSS, <canvas>Draw the thumbnail or preview image in one of the ways in WebGL.< / canvas> 10. A method for visualizing object detection image frames according to claim 1, characterized in that, Encapsulate steps S1 - S4 through the reactive function and ref function of the vue framework to form a vue component.

11. An object detection image frame visualization device, characterized in that, Specifically, it includes the following modules: Data acquisition module, used to acquire the image, scaling ratio, target detection box coordinates, a fixed coordinate system of size A×B, and the current display scene; Among them, the display scene includes a thumbnail display scene and a preview image display scene; Target detection box coordinate transformation module, used to perform coordinate transformation according to the image, image parameters, scaling ratio, and target detection box coordinates to form the transformed coordinates; Target detection box merging module, used to merge the target detection boxes that meet the merging conditions in the image to form a set of target detection boxes after merging processing; Image display module, used to draw and display the thumbnail and / or preview image of the image according to the set of target detection boxes after merging processing; Among them, both A and B are positive integers.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a method for visualizing a target detection image box as described in any one of claims 1 - 10.

13. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for visualizing a target detection image box as described in any one of claims 1 - 10.