Image Detection Method, Image Detection Device, Electronic Device and Storage Medium
By splicing the local image of the previous frame image to the current frame image during the image acquisition process, combining the object detection results and position information, the continuity and accuracy of inter-frame detection are solved, and the completeness and reliability of object detection are achieved.
Patent Information
- Application Number
- CN202510388219.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-31
AI Technical Summary
During the image acquisition process, there are problems of continuity and repeated detection of object detection between continuous frames, which affects the integrity and accuracy of the detection results.
By acquiring the partial image of the previous frame image and splicing it on the top of the current frame image, combining the target detection results and position information, the target output frame is determined to avoid information loss caused by inter-frame switching.
It improves the continuity and accuracy of object detection, reduces the visual breaks between objects and ensures the integrity and reliability of object object detection.
Smart Images

Figure CN119887786B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image detection technologies, and particularly to an image detection method, an image detection device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the actual image acquisition process, in order to improve the acquisition efficiency, image acquisition is performed using a specific size, so as to ensure obtaining a certain amount of image information while being able to improve the acquisition speed, which is applicable to scenarios that require rapid acquisition of image data, such as industrial automation detection, high-speed video surveillance, etc.
[0003] However, when performing object detection and analysis on continuously acquired image frames, problems such as the continuity of objects between different frames and repeated detection will be faced, thereby affecting the detection results. Summary of the Invention
[0004] To overcome the problems existing in the related art, an exemplary embodiment of the present disclosure provides an image detection method, which includes: obtaining a current frame image; if there is a previous frame image, obtaining a first partial image extracted from the bottom of the previous frame image, where the first partial image is obtained based on a first object detection result corresponding to the previous frame image, the first object detection result includes a first detection box corresponding to the first partial image and first position information of the first detection box, the distance between the bottom edge position of the first detection box and the bottom edge position of the previous frame image is less than a first distance, and the first partial image includes the top edge of the first detection box; splicing the first partial image on the top of the current frame image to obtain a first target image; performing object detection on the first target image to obtain a second object detection result corresponding to the current frame image; determining and outputting a target output box corresponding to the current frame image based on the second object detection result and the first position information.
[0005] In some embodiments, determining and outputting a target output box corresponding to the current frame image based on the second object detection result and the first position information includes: respectively determining a plurality of detection boxes corresponding to the first target image and position information of each detection box according to the second object detection result; determining a target output box corresponding to the current frame image from the plurality of detection boxes based on the position information of each detection box and the first position information; outputting the target output box.
[0006] In some embodiments, determining a target output box corresponding to the current frame image from multiple detection boxes based on the position information of each detection box and the first position information includes: determining a first target position in a first target image, where the first target position is a position in a first partial image that is at a first distance from the top edge position of the current frame image; according to the position information of each detection box, the first target position, and the first position information, respectively determining the category of each detection box according to a classification strategy; determining a target output box corresponding to the current frame image based on the category of each detection box; where the classification strategy includes: a first classification strategy: regarding detection boxes with overlapping position information and the first target position as a first category, and the detection boxes of the first category include detection boxes overlapping with the first position information; a second classification strategy: regarding detection boxes with position information between the first target position and a second target position as a second category, where the second target position is a position in the current frame image that is at a first distance from the bottom edge position of the current frame image; a third classification strategy: regarding detection boxes with position information between the top edge position of the first partial image and the top edge position of the current frame image as a third category.
[0007] In some embodiments, determining a target output box corresponding to the current frame image based on the category of each detection box includes: regarding detection boxes that are the union of the first category and the second category but the difference set with the third category as the target output box corresponding to the current frame image; and / or determining a third detection box that is the intersection of the first category and the third category among the multiple detection boxes; according to the position information of the third detection box, determining the area of the region of the third detection box; based on the position information of a fourth detection box in the previous frame image, determining a reference detection box for duplicate removal detection in the previous frame image and the third position information of the reference detection box, where the fourth detection box is a detection box with overlapping position information and the position of the top edge of the first partial image in the previous frame image; according to the position information of the third detection box and the third position information, determining the intersection area between the third detection box and the reference detection box; and if the ratio between the intersection area and the area of the region is less than a specified threshold, then regarding the third detection box as the target output box corresponding to the current frame image.
[0008] In some embodiments, determining a target output box corresponding to the current frame image based on the category of each detection box further includes: if the distance between the bottom edge position of the detection box and the bottom edge position of the current frame image is less than the first distance, then regarding the detection box as a non-target output box corresponding to the current frame image; and / or if the ratio between the intersection area and the area of the region is greater than or equal to the specified threshold, then regarding the third detection box as a non-target output box corresponding to the current frame image.
[0009] In some embodiments, the method further includes: based on the second position information of each second detection box, respectively determining a second distance between the top edge position of each second detection box and the bottom edge position of the current frame image, where the second detection box is a detection box whose distance between the bottom edge position of the detection box and the bottom edge of the current frame image is less than the first distance; determining an image extraction height based on the maximum second distance; extracting a second partial image from the bottom of the current frame image based on a third target position corresponding to the image extraction height; saving the second partial image to be spliced on the top of the next frame image to determine a target output box corresponding to the next frame image.
[0010] In some embodiments, the classification strategy further includes a fourth classification strategy: regarding a detection box whose position information overlaps with the third target position as a fourth category; the method further includes: saving the position information corresponding to the detection box of the fourth category for determining a non-target output box corresponding to the next frame image.
[0011] In some embodiments, splicing the first partial image on the top of the current frame image to obtain a first target image includes: splicing the first partial image on the top of the current frame image to obtain a spliced image; determining the image height of the first partial image; determining a first splicing height of the spliced image according to the image height of the first partial image and the image height of the current frame image; determining a second splicing height of the previous frame image corresponding to the first target image; if the second splicing height is greater than the first splicing height, then according to the height difference between the second splicing height and the first splicing height, clearing the regional image corresponding to the height difference at the top of the previous frame image corresponding to the first target image to obtain an intermediate image; covering the spliced image on the intermediate image to obtain the first target image.
[0012] In some embodiments, the method further includes: if there is no previous frame image, performing target detection on the current frame image to obtain a third target detection result corresponding to the current frame image; determining and outputting a target output box corresponding to the current frame image based on the third target detection result.
[0013] In some embodiments, determining and outputting a target output box corresponding to the current frame image based on the third target detection result includes: respectively determining a plurality of detection boxes corresponding to the current frame image and the position information of each detection box according to the third target detection result; regarding the detection box belonging to the fifth category as the target output box corresponding to the current frame image according to the position information of each detection box; where the detection box of the fifth category is a detection box whose position information is between the top edge position of the current frame image and the second target position, and the second target position is a position whose distance from the bottom edge position of the current frame image in the current frame image is the first distance.
[0014] Second aspect, the present disclosure further provides an image detection device, the device includes: a first acquisition module, configured to acquire a current frame image; a second acquisition module, configured to acquire a first partial image extracted from the bottom of the previous frame image if there is a previous frame image, where the first partial image is obtained based on a first target detection result corresponding to the previous frame image, the first target detection result includes a first detection box corresponding to the first partial image and first position information of the first detection box, a distance between a bottom edge position of the first detection box and a bottom edge position of the previous frame image is less than a first distance, and the first partial image includes a top edge of the first detection box; a splicing module, configured to splice the first partial image on top of the current frame image to obtain a first target image; a first detection module, configured to perform target detection on the first target image to obtain a second target detection result corresponding to the current frame image; a first processing module, configured to determine and output a target output box corresponding to the current frame image based on the second target detection result and the first position information.
[0015] Third aspect, the present disclosure further provides an electronic device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the image detection method provided in any of the above aspects.
[0016] Fourth aspect, the present disclosure further provides a computer-readable storage medium, which stores the following program, and the program is used to execute the image detection method provided in any of the above aspects.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.
[0018] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: According to the image detection method provided by the present disclosure, in the case where there is a previous frame image, a first partial image extracted from the bottom of the previous frame image is obtained and spliced with the current frame image to obtain a first target image, which helps to reduce the visual break when the object moves between frames, and further improves the continuity and accuracy of target detection. Among them, the first partial image is obtained based on the first target detection result corresponding to the previous frame image, and the height between the first partial image and the image bottom edge of the previous frame image is determined according to the top edge position of the first detection box in the first partial image, so as to ensure the extraction reliability and flexibility of the first partial image, adapt to various image splicing scenarios, thereby helping to improve the accuracy of the spliced image and ensure the integrity of object target detection. Moreover, by performing target detection on the first target image and determining and outputting the target output box corresponding to the current frame image based on the obtained second target detection result and the first position information of the first detection box, the situation of information loss caused by frame switching can be effectively avoided, thereby effectively improving the reliability and accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] By describing the exemplary embodiments of the present disclosure in conjunction with the drawings, the present disclosure can be better understood. In the drawings:
[0020] Figure 1 It is a schematic flowchart of an image detection method shown in an exemplary embodiment of the present disclosure;
[0021] Figure 2 It is a schematic flowchart of another image detection method shown in an exemplary embodiment of the present disclosure;
[0022] Figure 3 It is a schematic diagram of the position distribution of a detection box shown in an exemplary embodiment of the present disclosure;
[0023] Figure 4 It is a schematic diagram of the position distribution of another detection box shown in an exemplary embodiment of the present disclosure;
[0024] Figure 5 It is a schematic diagram of the position distribution of yet another detection box shown in an exemplary embodiment of the present disclosure;
[0025] Figure 6 It is a schematic diagram of a target image shown in an exemplary embodiment of the present disclosure;
[0026] Figure 7 It is a schematic flowchart of yet another image detection method shown in an exemplary embodiment of the present disclosure;
[0027] Figure 8Another schematic diagram of the position distribution of the detection frame shown in an exemplary embodiment of the present disclosure;
[0028] Figure 9 Another schematic flowchart of an image detection method shown in an exemplary embodiment of the present disclosure;
[0029] Figure 10 A schematic diagram of a background image shown in an exemplary embodiment of the present disclosure;
[0030] Figure 11 Another schematic diagram of a target image shown in an exemplary embodiment of the present disclosure;
[0031] Figure 12 Another schematic flowchart of an image detection method shown in another exemplary embodiment of the present disclosure;
[0032] Figure 13 A schematic structural diagram of an image detection method device shown in another exemplary embodiment of the present disclosure;
[0033] Figure 14 A schematic structural diagram of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0034] The following will describe the detailed implementation manners of the present disclosure. It should be noted that in the specific description process of these implementation manners, for the sake of concise description, this specification cannot describe all features of the actual implementation manners in detail. It should be understood that in the actual implementation process of any implementation manner, just as in the process of any engineering project or design project, in order to achieve the specific goals of the developer and to meet system-related or business-related restrictions, various specific decisions are often made, and these will also change from one implementation manner to another. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present disclosure, some design, manufacturing, or production changes based on the technical content disclosed in the present disclosure are only conventional technical means and should not be understood as the content of the present disclosure being insufficient.
[0035] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings understood by those of ordinary skill in the technical field to which this disclosure pertains. The "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one. The terms such as "including" or "comprising" mean that the elements or objects appearing before "including" or "comprising" cover the elements or objects listed after "including" or "comprising" and their equivalent elements, and do not exclude other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, nor are they limited to direct or indirect connections.
[0036] In the field of industrial image detection, the traditional method of performing object detection frame by frame has many limitations in the process of continuous industrial shooting. For example, in practical applications, when an object is at the edge of the frame, the object is prone to be truncated, which in turn affects the integrity of the detection result.
[0037] In the actual image acquisition process, in order to improve the acquisition efficiency, image acquisition is performed using a specific size, so as to ensure a certain amount of image information while being able to improve the acquisition speed, which is applicable to scenarios that require rapid acquisition of image data, such as industrial automation detection, high-speed video surveillance, etc. For example, the size of the acquired image can be width*height, and height = 1 / 8 width.
[0038] However, when performing object detection and analysis on continuously acquired frame images, problems such as the continuity of the object between different frames and repeated detection will be faced, which in turn affects the detection result.
[0039] In the related art, to solve the above problems, a method of performing image stitching using a fixed cropping area is adopted to achieve cross-frame detection. However, since the sizes and moving speeds of objects are different, using this method for image detection is likely to affect the detection effect. For example, for an object with a large size or a slow moving speed, it may not be completely included in the cropping area, resulting in incomplete object information and thus detection failure. Or, for an object with a small size or a fast moving speed, the cropping area may be too large, resulting in unnecessary computational burden and false detection.
[0040] To solve the above problems, an exemplary embodiment of this disclosure provides an image detection method. As Figure 1 shown, the image detection method may include the following steps:
[0041] Step S110, obtain the current frame image.
[0042] The current frame image refers to the frame image that needs to be processed for object detection currently. This current frame image is one of a series of consecutive frame images. These multiple frame images can be obtained from local, cloud, or real-time acquisition. The image content corresponding to the current frame image can depend on the corresponding image acquisition scenario. For example, taking the ore sorting scenario as the image acquisition scenario, the image content corresponding to the current frame image can be the ore transported by the transmission mechanism of the ore sorter.
[0043] Step S120, if there is a previous frame image, obtain the first local image extracted from the bottom of the previous frame image.
[0044] Since in the actual process of image acquisition, there may be a situation where an object is truncated at the edge of the picture. Therefore, to determine whether there is an object truncated by the edge of the picture in the previous frame image, first judge whether there is a previous frame image. Here, the previous frame image refers to the frame image adjacent to the current frame image and that has undergone object detection processing.
[0045] If there is a previous frame image, obtain the first local image extracted from the bottom of the previous frame image. Among them, the first local image is obtained based on the first object detection result corresponding to the previous frame image. The first object detection result includes the first detection box corresponding to the first local image and the first position information of the first detection box. The distance between the bottom edge position of the first detection box and the bottom edge position of the previous frame image is less than the first distance, and the first local image includes the top edge of the first detection box.
[0046] That is, the first distance can be understood as the minimum distance for judging whether the information of the detection box is complete. According to the first object detection result, the position information of each detection box in the previous frame image can be determined. Compare the distance between the bottom edge position of each detection box and the bottom edge position of the previous frame image with the first distance respectively. If the distance between the bottom edge position of the detection box and the bottom edge position of the previous frame image is greater than or equal to the first distance, it indicates that the object corresponding to this detection box is not truncated by the edge of the picture during the actual image acquisition process, and the information of this detection box is complete. However, if the distance between the bottom edge position of the detection box and the bottom edge position of the previous frame image is less than the first distance, it indicates that the object corresponding to this detection box may be truncated by the edge of the picture during the actual image acquisition process, and the information of this detection box may be incomplete. Therefore, to ensure the integrity of the object's information, this detection box can be considered as the first detection box in the previous frame image that needs to perform cross-frame object detection.
[0047] To ensure the information integrity of the first detection frame, the height of the first local image is determined according to the position of the first detection frame to ensure the rationality of the extraction of the first local image. In some embodiments, if the number of first detection frames corresponding to the previous frame image is one, the height of the first local image can be greater than or equal to the distance between the top edge position of the first detection frame and the bottom edge position of the previous frame image, so as to avoid the waste of computing resources caused by excessive extraction of the local image. In other embodiments, if the number of first detection frames corresponding to the previous frame image is multiple, the height of the first local image can depend on the maximum distance between the top edge positions of the multiple first detection frames and the bottom edge position of the previous frame image, so as to avoid the situation that the first detection frame is incomplete due to under-extraction of the local image.
[0048] In some examples, the height of the first local image can be equal to the maximum distance between the top edge position of the first detection frame and the bottom edge position of the previous frame image, so that the first local image extracted from the bottom of the previous frame image can be quickly determined. For example, if the maximum distance between the top edge position of the first detection frame and the bottom edge position of the previous frame image is max_H, the first local image extracted from the bottom of the previous frame image is a local area with a height of max_H from the bottom edge of the previous frame image as the reference.
[0049] In other examples, to ensure the extraction integrity of the first detection frame, the tolerance pixel height of the corresponding detection frame label is determined according to the target detection algorithm used, and then the sum of the tolerance pixel height and the maximum distance between the top edge position of the first detection frame and the bottom edge position of the previous frame image is used as the height of the first local image. The target detection algorithm corresponds to the tolerance pixel height one by one. For example, if the tolerance pixel height is determined to be top_h and the maximum distance between the top edge position of the first detection frame and the bottom edge position of the previous frame image is max_H, the first local image extracted from the bottom of the previous frame image is a local area with a height of H = max_H + top_h from the bottom edge of the previous frame image as the reference. For example, taking the target detection algorithm as YOLOv5, the tolerance pixel height of its corresponding detection frame label is 2 - 3 pixels, and thus the tolerance pixel height top_h can be selected as 3.
[0050] In some implementation scenarios, taking the industrial image acquisition scenario as an example, the object can include but is not limited to ores, cotton, plastics, metals, etc., and can be specifically determined according to the actual image acquisition scenario.
[0051] Step S130, splice the first local image on the top of the current frame image to obtain the first target image.
[0052] Since the first partial image is the tail image of the previous frame image, and the distance between the bottom edge position of the first detection box in the first partial image and the bottom edge position of the previous frame image is less than the first distance, in order to determine whether the information of the first detection box is complete and whether the object image corresponding to the first detection box is truncated by the image edge, the first partial image is spliced on the top of the current frame image to obtain the first target image, so as to reduce the visual break when the same object moves between frames, which is beneficial to improving the continuity and accuracy of target detection.
[0053] Step S140: Perform target detection on the first target image to obtain the second target detection result corresponding to the current frame image.
[0054] Perform target detection on the first target image, and use the result of target detection on the first target image as the second target detection result corresponding to the current frame image, so that the distribution of object images in the current frame image can be determined according to the second target detection result subsequently, which is convenient for subsequent targeted tracking and recognition.
[0055] Step S150: Based on the second target detection result and the first position information, determine and output the target output box corresponding to the current frame image.
[0056] By combining the second target detection result with the first position information, the detection box information corresponding to the object image involved in target detection across image frames can be made more complete. Furthermore, it can be determined which detection boxes among the multiple detection boxes corresponding to the current frame image belong to the target detection boxes that can be output, so that the determined and output target output box corresponding to the current frame image is more reliable, effectively avoiding the loss of information caused by frame switching, and thus being more helpful to ensure the accuracy of target detection. That is, the target detection box can be understood as the detection box that can be output corresponding to the current frame image.
[0057] According to the image detection method provided by the present disclosure, in the case where there is a previous frame image, a first partial image extracted from the bottom of the previous frame image is obtained and spliced with the current frame image to obtain a first target image, which helps to reduce the visual break when the object moves between frames, and further improves the continuity and accuracy of target detection. Among them, the first partial image is obtained based on the first target detection result corresponding to the previous frame image, and the height between the first partial image and the image bottom edge of the previous frame image is determined according to the top edge position of the first detection box in the first partial image, so as to ensure the extraction reliability and flexibility of the first partial image, adapt to various image splicing scenarios, thereby helping to improve the accuracy of the spliced image and ensure the integrity of object target detection. Moreover, by performing target detection on the first target image and determining and outputting the target output box corresponding to the current frame image based on the obtained second target detection result and the first position information of the first detection box, the situation of information loss caused by frame switching can be effectively avoided, thereby effectively improving the reliability and accuracy of target detection.
[0058] In some embodiments, as Figure 2 shown, the above step S150 may include the following steps:
[0059] Step S151, according to the second target detection result, respectively determine a plurality of detection boxes corresponding to the first target image and the position information of each detection box.
[0060] According to the second target detection result, it can be determined whether there is an identified object image in the first target image. If there is an identified object image, then through the second target detection result, the detection box corresponding to each identified object image in the first target image and the position information of the detection box can be respectively determined.
[0061] Step S152, based on the position information of each detection box and the first position information, determine the target output box corresponding to the current frame image from the plurality of detection boxes.
[0062] According to the position information of each detection box, the distribution position of the corresponding object image in the target frame image can be determined. Furthermore, in combination with the first position information, targeted screening is performed on the plurality of detection boxes corresponding to the first target image to determine the target output box corresponding to the current frame image among the plurality of detection boxes, which can effectively reduce the occurrence of false detection or missed detection, thereby improving the accuracy of target detection.
[0063] Step S153, output the target output box.
[0064] Outputting the target output box helps relevant personnel to identify the object corresponding to the current frame image, so as to perform targeted processing on the object corresponding to the target output box subsequently and meet the requirements of relevant tasks. For example, performing processing such as tracking, identifying, or monitoring on the object corresponding to the target output box.
[0065] According to the above method, by combining the position information of multiple detection boxes corresponding to the first target image with the first position information of the previous frame, the target detection box corresponding to the current frame image is determined, which helps to improve the robustness and accuracy of target detection.
[0066] Moreover, in a dynamic scene, by determining the target detection box corresponding to the current frame image in this way, the object can also be effectively tracked and the detection error caused by scene changes can be reduced, thereby helping to improve the reliability and accuracy of target detection.
[0067] In some other embodiments, the above step S152 may include the following steps:
[0068] Step a1, determining the first target position in the first target image;
[0069] Step a2, respectively determining the category of each detection box according to the position information of each detection box, the first target position, and the first position information according to the classification strategy;
[0070] Step a3, determining the target output box corresponding to the current frame image based on the category of each detection box.
[0071] Specifically, to better distinguish whether the detection box is detected for the first time, first determine the first target position according to the first target image. Among them, the first target position can be understood as the position in the first local image that is at a first distance from the top edge position of the current frame image. For example, as Figure 3 shown, if the top edge position of the current frame image is start_y, the first target position is the position that is at a first distance h from start_y and belongs to the first local image.
[0072] Since the first distance is the minimum distance for determining whether the information of the detection box is complete. Therefore, according to the position information of the detection box, the first target position, and the first position information, the position types of each detection box can be divided, and then according to the category corresponding to each detection box, it can be judged whether each detection box is a repeatedly detected detection box and whether the information of the detection box is complete.
[0073] Among them, the classification strategies include: The first classification strategy: The detection boxes whose position information overlaps with the first target position are regarded as the first category, and the detection boxes of the first category include the detection boxes that overlap with the first position information; The second classification strategy: The detection boxes whose position information is between the first target position and the second target position are regarded as the second category, and the second target position is the position whose distance from the bottom edge position of the current frame image is the first distance; The third classification strategy: The detection boxes whose position information is between the top edge position of the first local image and the top edge position of the current frame image are regarded as the third category. For example, the detection boxes that meet the first category can be Figure 3 the detection boxes with filled colors in; The detection boxes that meet the second category can be Figure 4 the detection boxes with filled colors in; The detection boxes that meet the third category can be Figure 5 the detection boxes with filled colors in. Among them, Figures 3 - 5 in, the first target position is the position that is at a first distance h from start_y and belongs to the first local image, and the second target position is the position that is at a first distance h from the bottom edge position end_y of the current frame image and belongs to the current frame image. It should be noted that Figures 3 - 5 the positions of the detection boxes in are only for illustration, and the specific positions of each detection box depend on the actual second target detection result.
[0074] In some application scenarios, such as Figure 6 shown, for the convenience of description, it can be determined that the top edge position of the current frame image is start_y, the bottom edge position of the current frame image is end_y, the top edge position is higher than the bottom edge position, but the value of end_y is greater than the value of start_y, and the first distance is h. When dividing the detection boxes according to the first classification strategy, the detection boxes whose top edge position is higher than or equal to the first target position, the bottom edge position is lower than or equal to the first target position and the bottom edge position is less than or equal to end_y can be divided into the first category. When dividing the detection boxes according to the second classification strategy, the detection boxes whose top edge position is lower than or equal to the first target position and the bottom edge position is higher than or equal to the second target position can be divided into the second category. When dividing the detection boxes according to the third classification strategy, the detection boxes whose bottom edge position is lower than or equal to start_y can be divided into the third category.
[0075] From the results of dividing according to each classification strategy, it can be seen that
[0076] There are detection boxes in the detection boxes of the first category that overlap with the first position information of the first detection box. Therefore, the detection boxes of this category may be obtained based on cross-frame detection or may be obtained by detecting the first local image;
[0077] The detection boxes of the second category are obtained by performing object detection on the current frame image. Such detection boxes are not obtained by repeatedly detecting the same object image and have complete information;
[0078] The detection boxes of the third category are obtained by performing object detection on the previous frame image. There may be a situation of repeatedly detecting the same object image, but the information of such detection boxes may be incomplete.
[0079] After clarifying the categories corresponding to each detection box, targeted screening can be performed according to the categories of each detection box, so as to determine the target output box corresponding to the current frame image, ensuring the reliability and accuracy of the determination of the target output box.
[0080] In some examples, the above step a2 can determine the target output box in any one of the following methods or a combination of them:
[0081] Method 1: The detection boxes that are the union of the first category and the second category but the difference set from the third category are used as the target output boxes corresponding to the current frame image. Since the detection boxes of the third category are obtained by performing object detection on the previous frame image, there may be detection boxes in the detection boxes of the first category or the second category that overlap with the position information of the detection boxes of the third category. Therefore, to avoid the repeated output of detection boxes of the same object image, the detection boxes in the first category that overlap with the position information of the third category and the detection boxes in the second category that overlap with the position information of the third category are excluded, and the detection boxes that are the union of the first category and the second category but the difference set from the third category are used as the target output boxes corresponding to the current frame image, ensuring the reliability and accuracy of the determination of the target output box.
[0082] Method 2: Determine the third detection boxes that are the intersection of the first category and the third category among multiple detection boxes; according to the position information of the third detection boxes, determine the area of the region of the third detection boxes; based on the position information of the fourth detection boxes in the previous frame image, determine the reference detection boxes for duplicate detection removal in the previous frame image and the third position information of the reference detection boxes; according to the position information of the third detection boxes and the third position information, determine the intersection area between the third detection boxes and the reference detection boxes; and if the ratio between the intersection area and the region area is less than the specified threshold, then use the third detection boxes as the target output boxes corresponding to the current frame image.
[0083] Specifically, during the actual execution of the object detection algorithm for object detection processing, due to detection accuracy issues, different detection frames may be generated when performing multiple object detections on the same object image. For example, there are differences in the size or position of the detection frames. The detection frames of the first category are the detection frames whose position information overlaps with the first target position, and the detection frames of the third category are the detection frames whose position information is between the top edge position of the first partial image and the top edge position of the current frame image. If the same detection frame belongs to both the first category and the third category, it can be considered that the object image corresponding to this detection frame partially or entirely belongs to the previous frame image. Therefore, to avoid repeatedly outputting the detection frames corresponding to the same object, first, the detection frames that belong to both the first category and the third category among the multiple detection frames are used as the third detection frames, and the area of the region of the third detection frame is determined according to the position information of the third detection frame.
[0084] Determine the position information of the fourth detection frame in the previous frame image. Among them, the fourth detection frame can be understood as the detection frame whose position information overlaps with the position of the top edge of the first partial image in the previous frame image, and the position information of each detection frame in the previous frame image is determined according to the first object detection result. The object image corresponding to the fourth detection frame can be considered as the object image that has been recognized in the previous frame image. Since the first partial image is extracted from the previous frame image, part or all of the fourth detection frame is included in the first partial image. Therefore, the fourth detection frame whose position information overlaps with the position of the first partial image in the previous frame image is used as the reference detection frame for duplicate detection removal, and the third position information for determining this reference detection frame is obtained.
[0085] To determine whether the object image corresponding to the third detection frame and the object image corresponding to the reference detection frame are the same object image, it can be judged by the Intersection over Union (IOU). That is, according to the position information of the third detection frame and the third position information of the reference detection frame, the intersection area between the third detection frame and the reference detection frame is determined. If the ratio of the intersection area to the area of the region is less than the specified threshold, it indicates that the area of the image region where the third detection frame and the reference detection frame intersect accounts for a relatively small proportion in the area of the region of the third detection frame. Furthermore, it can be determined that the object image corresponding to the third detection frame and the object image corresponding to the reference detection frame are not the same object image. Therefore, it can be determined that the object image corresponding to the third detection frame has not been detected repeatedly. Thus, the third detection frame can be used as the target output frame corresponding to the current frame image. Among them, the specified threshold can be understood as the minimum ratio at which two object images are considered to be the same object image within the allowable error range.
[0086] In some other examples, step a2 above can also determine the non-target output frames in any one of the following ways or a combination thereof:
[0087] Method 1: If the distance between the bottom edge position of the detection box and the bottom edge position of the current frame image is less than the first distance, the detection box is used as the non-target output box corresponding to the current frame image. Since the first distance is the minimum distance for determining whether the information of the detection box is complete, if the distance between the bottom edge position of the detection box and the bottom edge position of the current frame image is less than the first distance, it indicates that the information of the detection box may be incomplete. Therefore, to ensure the reliability of object detection, the detection box is used as the non-target output box corresponding to the current frame image to avoid misidentification and thus ensure the accuracy of the target output box.
[0088] Method 2: If the ratio of the intersection area to the area of the region is greater than or equal to the specified threshold, the third detection box is used as the non-target output box corresponding to the current frame image. If the ratio of the intersection area to the area of the region is greater than or equal to the specified threshold, it indicates that the area of the image region where the third detection box intersects with the reference detection box accounts for a relatively large proportion in the area of the third detection box. It can be considered that the object image corresponding to the third detection box and the object image corresponding to the reference detection box are the same object image. Since the object image corresponding to the reference detection box is an object image that has been recognized and marked, to avoid the situation where the detection boxes of the same object image are repeatedly output, the third detection box can be used as the non-target output box corresponding to the current frame image to avoid affecting the output result of the target output box corresponding to the current frame image, thereby helping to reduce the need for manual intervention to reduce duplicates, improve the reliability of object detection, and facilitate reducing the occurrence of false alarms triggered by repeated detections.
[0089] As Figure 7 shown, the image detection method may include the following steps:
[0090] Step S110, obtain the current frame image.
[0091] Step S120, if there is a previous frame image, obtain the first partial image extracted from the bottom of the previous frame image.
[0092] Step S130, splice the first partial image on the top of the current frame image to obtain the first target image.
[0093] Step S140, perform object detection on the first target image to obtain the second target detection result corresponding to the current frame image.
[0094] Step S150, based on the second target detection result and the first position information, determine and output the target output box corresponding to the current frame image.
[0095] Step S160: Based on the second position information of each second detection box, respectively determine the second distance between the top edge position of each second detection box and the bottom edge position of the current frame image.
[0096] Among them, the second detection box is a detection box whose distance between the bottom edge position of the detection box and the bottom edge of the current frame image is less than the first distance. The number of second detection boxes is at least one. To ensure the integrity of the information of the object corresponding to the second detection box, the second distance between the top edge position of each second detection box and the bottom edge position of the current frame image is respectively determined according to the second position information of each second detection box, so as to determine the maximum second distance therefrom.
[0097] Step S170: Based on the maximum second distance, determine the image extraction height.
[0098] According to the maximum second distance, it can be ensured that the top edge positions of all second detection boxes can be included. Therefore, based on this maximum second distance, the image extraction height is determined. Among them, the image extraction height refers to the position at a distance of the image extraction height upward along the top edge direction of the current frame image with the bottom edge position of the current frame image as the reference.
[0099] In some examples, the maximum second distance can be directly used as the maximum second distance, which helps to improve the determination efficiency of the image extraction height.
[0100] In other examples, due to different tolerance pixel sizes when different object detection algorithms mark detection boxes. The tolerance pixel size includes the tolerance pixel width and the tolerance pixel height. Therefore, in the case of determining the maximum second distance, the tolerance pixel height corresponding to the current object detection algorithm can be combined to determine the image extraction height. For example, the image extraction height can be the sum of the maximum second distance and the tolerance pixel height corresponding to the current object detection algorithm. In some application scenarios, the tolerance pixel height can be 3 pixels or 5 pixels, which can be specifically determined according to the actual object detection algorithm and requirements.
[0101] Step S180: Extract the second local image from the bottom of the current frame image based on the third target position corresponding to the image extraction height.
[0102] According to the image extraction height, determine the third target position in the current frame image that is at a distance of the image extraction height from the bottom edge position of the current frame image, and intercept the current frame image according to this third target position, so as to obtain the second local image extracted from the bottom of the current frame image.
[0103] Step S190: Save the second local image to splice it on the top of the next frame image and determine the target output box corresponding to the next frame image.
[0104] Save the second partial image so that the second partial image can be stitched on top of the next frame image later, and perform cross-frame detection using the position information of the second detection box corresponding to the second partial image to determine the target output box corresponding to the next frame image.
[0105] According to the image detection method provided by the present disclosure, based on the distance between the bottom edge position of each detection box in each frame image and the bottom edge position of the corresponding frame image, the height of the partial image to be extracted can be dynamically adjusted, so that the image detection method can be adapted to application scenarios for target detection of various objects. Furthermore, it can not only improve the reliability of image stitching, but also help ensure the accuracy of cross-frame target detection, thereby contributing to enhancing the reliability and versatility of target detection and meeting the needs of various target detections.
[0106] In some other embodiments, the classification strategy further includes a fourth classification strategy: regarding the detection box whose position information overlaps with the third target position as the fourth category. For example, the detection box that conforms to the fourth category can be Figure 8 the detection box with a filling color in Figure 8 . It should be noted that Figure 6 the positions of the detection boxes in
[0107] are only for illustration, and the specific positions of the detection boxes depend on the actual second target detection results. Combining Figure 9 , the image extraction height between the third target position and the current frame image is H. Then, when dividing the detection boxes according to the fourth classification strategy, the detection boxes whose position information overlaps with H can be classified into the fourth category. Subsequently, through the detection boxes of these four categories, it can be assisted to determine whether there is a situation of repeated detection of the same object image in the next frame image, so as to ensure the reliability and accuracy of target detection in the next frame image. Therefore, the image detection method may further include: saving the position information corresponding to the detection boxes of the fourth category for determining the non-target output box corresponding to the next frame image, so as to ensure the orderly progress of the task of target detection for multiple frame images.
[0108] Step S131: Stitch the first partial image on top of the current frame image to obtain a stitched image.
[0109] Step S132: Determine the image height of the first partial image.
[0110] Step S133: Determine the first stitching height of the stitched image according to the image height of the first partial image and the image height of the current frame image.
[0111] Since the image height of the first partial image is determined based on the distance between the top edge position of the first detection box and the bottom edge position of the previous frame image, and the height of the first partial image corresponding to each frame image is different. Therefore, to determine the overall image height of the spliced image, the first splicing height of the spliced image is determined according to the image height of the first partial image and the image height of the current frame image.
[0112] Step S134, determine the second splicing height of the first target image corresponding to the previous frame image.
[0113] The first target image corresponding to the previous frame image is an image formed by splicing the partial image extracted from the bottom of the previous frame image of the previous frame image and the previous frame image. Since the partial image that can be extracted corresponding to each frame image is determined based on the corresponding object detection result, the height of the partial image corresponding to each frame image may be different. To improve the accuracy of object detection for the current frame image, the second splicing height of the first target image corresponding to the previous frame image is determined.
[0114] Step S135, if the second splicing height is greater than the first splicing height, then according to the height difference between the second splicing height and the first splicing height, clear the regional image corresponding to the top of the first target image corresponding to the previous frame image and the height difference, and obtain an intermediate image.
[0115] If the second splicing height is greater than the first splicing height, it means that if the spliced image is directly placed on the first target image corresponding to the previous frame image for object detection, the regional image of the higher part will affect the accuracy of the second object detection result of object detection for the spliced image. For example, the situation of misrecognition occurs. Therefore, according to the height difference between the second splicing height and the first splicing height, clear the regional image corresponding to the top of the first target image corresponding to the previous frame image and the height difference to eliminate the interference of this regional image on the image detection method for the spliced image, so as to obtain an intermediate image.
[0116] In some examples, if the second splicing height is less than or equal to the first splicing height, the spliced image can be directly covered on the first target image corresponding to the previous frame image.
[0117] Step S136, cover the spliced image on the intermediate image to obtain the first target image.
[0118] Covering the spliced image on the intermediate image can ensure that the image sizes for object detection for each frame image are the same. Furthermore, when object detection is performed based on the obtained first target image subsequently, the efficiency of object detection can be effectively improved.
[0119] By obtaining the first target image in the above manner, the interference of the first target image corresponding to the previous frame image on the target detection of the current frame image can be avoided, which helps to ensure the accuracy and efficiency of target detection.
[0120] In some optional application scenarios, to improve the image acquisition efficiency and ensure that the sizes of the target images (including the first target image and the second target image) corresponding to each frame image are the same when performing target detection on each frame image, a background image as shown in Figure 10 is created. The background color of this background image is white, which helps with subsequent image processing and visualization and is convenient for distinguishing the detection frames and the targets. The background image can be a square image with the same height and width. For example, the height and width of the background image can be the same as the width of the image frame, which is convenient for pasting images and cross-frame detection during subsequent target detection.
[0121] In some other optional application scenarios, combined with the background image shown in Figure 10 , the first target image obtained after splicing can be as shown in Figure 11 .
[0122] The present disclosure also provides another image detection method. As shown in Figure 12 , this image detection method may include the following steps:
[0123] Step S110, obtain the current frame image.
[0124] Step S120, if there is a previous frame image, obtain the first partial image extracted from the bottom of the previous frame image.
[0125] Step S130, splice the first partial image on the top of the current frame image to obtain the first target image.
[0126] Step S140, perform target detection on the first target image to obtain the second target detection result corresponding to the current frame image.
[0127] Step S150, based on the second target detection result and the first position information, determine and output the target output box corresponding to the current frame image.
[0128] Step S1100, if there is no previous frame image, perform target detection on the current frame image to obtain the third target detection result corresponding to the current frame image.
[0129] If there is no previous frame image, it indicates that the current frame image is the first frame image. To ensure the integrity of the information of the detection box of the object image at the top of the first frame image, target detection is performed on the current frame image to identify the object images included in the current frame image and mark the corresponding detection boxes for the identified object images, thereby obtaining the third target detection result corresponding to the current frame image.
[0130] Step S1110: Based on the third target detection result, determine and output the target output box corresponding to the current frame image.
[0131] According to the third target detection result, the detection boxes corresponding to the recognized object images in the current frame image can be determined, and based on the positions of the detection boxes, the target output box corresponding to the current frame image can be determined and output to ensure the accuracy of target detection.
[0132] According to the image detection method provided by the present disclosure, the utilization of detection information between consecutive frames can be fully considered, thereby effectively ensuring the information integrity of an object during cross-frame detection. Furthermore, it can not only ensure the continuity of target detection but also contribute to the accuracy of target detection. Therefore, applying the image detection method provided by the present disclosure in a dynamic scenario can completely capture the features of each object, avoid information loss caused by frame switching, and make the target detection process more reliable.
[0133] In some examples, the above step S1110 may include the following steps:
[0134] Step b1: According to the third target detection result, respectively determine multiple detection boxes corresponding to the current frame image and the position information of each detection box;
[0135] Step b2: According to the position information of each detection box, use the detection boxes belonging to the fifth category as the target output box corresponding to the current frame image.
[0136] Since the current frame image is the first frame image, all the multiple detection boxes obtained by performing target detection on this frame image are detected for the first time. Since the distance between the second target position and the bottom edge position of the current frame image is the first distance, to ensure the information integrity and accuracy of the detection boxes, the detection boxes belonging to the fifth category are screened from the multiple detection boxes, and the detection boxes of the fifth category are used as the target output box corresponding to the current frame image to ensure the reliability of determining the target output box. Among them, the detection boxes of the fifth category are the detection boxes whose position information is between the top edge position of the current frame image and the second target position, and the second target position is the position in the current frame image where the distance from the bottom edge position of the current frame image is the first distance.
[0137] In some application scenarios, if the value corresponding to the first distance is smaller than the image height value of the actual object image, the detection boxes of the third category can be considered as misrecognized noises. Furthermore, to improve the accuracy of the target output box, the detection boxes that are the union of the first category and the second category but the difference set with the third category can be used as the target output box corresponding to the current frame image to improve the reliability of determining the target output box.
[0138] In some optionally applicable scenarios, the process of performing target detection by the image detection method provided by the present disclosure may be as follows:
[0139] Create a background image with the same width as the width of the frame image. Wherein, the width of the frame image is greater than the height of the frame image. For example, if the height of the frame image is a and the width of the frame image is 8a, then the size of the background image is 8a 8a.
[0140] For the first frame image: Obtain the current frame image and overlay the current frame image on the background image with the bottom of the current frame image flush with the bottom of the background image to obtain a second target image. Based on the currently used target detection algorithm, perform image scaling processing on the second target image so that the size of the processed second target image can meet the requirements of the corresponding processing model of the target detection algorithm, thereby helping to improve the accuracy of target detection.
[0141] Perform target detection on the second target image to obtain the third target detection result corresponding to the current frame image, and based on the third target detection result, respectively determine a plurality of detection frames corresponding to the second target image and the position information of each detection frame.
[0142] Based on the position information of each detection frame, use the detection frame belonging to the fifth category as the target output frame corresponding to the current frame image and output it. In some examples, since the image height of the second target image is greater than the image height of the current frame image, therefore, based on the position information of each detection frame, the detection frame belonging to the third category can also be used as the target output frame and output to ensure the output integrity of the target output frame.
[0143] Based on the third target detection result, determine a second detection frame whose distance between the bottom edge position of the detection frame and the image bottom edge of the current frame image is less than the first distance. Based on the second position information of each second detection frame, respectively determine the second distance between the top edge position of each second detection frame and the bottom edge position of the current frame image. Based on the maximum second distance, determine the image extraction height; extract a second local image from the bottom of the current frame image based on the third target position corresponding to the image extraction height; save the second local image to be spliced on the top of the next frame image to determine the target output frame corresponding to the next frame image.
[0144] For the second frame image: The difference from the processing of the first frame image is that the second partial image needs to be spliced on the top of the second frame image, and the splicing result is overlaid on the corresponding first target image of the first frame image, so as to obtain the first target image corresponding to the second frame image. Moreover, when determining the target output box corresponding to the second frame image, it is necessary to jointly determine it by combining the position information of the second detection box in the second partial image (wherein, the first frame image can be regarded as the previous frame image of the second frame image, the second partial image is the first partial image extracted from the bottom of the previous frame image, the second detection box can be regarded as determined based on the first target detection result of the previous frame image, and the position information of the second detection box can be regarded as the first position information of the first detection box) and the target detection result of the first target image corresponding to the second frame image, so as to ensure the information integrity of the object during cross-frame detection.
[0145] For the middle frame image and the last frame image: The difference from the processing of the second frame image is that it is necessary to compare the first splicing height of the splicing image obtained by splicing the current frame image and the first partial image of the previous frame image with the second splicing height of the corresponding first target image of the previous frame image. If the second splicing height is greater than the first splicing height, according to the height difference between the second splicing height and the first splicing height, the regional image corresponding to the height difference at the top of the corresponding first target image of the previous frame image is cleared to obtain the middle image; and the splicing image is overlaid on the middle image to obtain the first target image corresponding to the current frame image, so as to avoid the occurrence of false detection and ensure the reliability of target detection for continuous frame images. If the second splicing height is less than or equal to the first splicing height, the splicing image can be directly overlaid on the corresponding first target image of the previous frame image.
[0146] In some other optional application scenarios, the image detection method provided by the present disclosure can be applied to ore sorting equipment. For example, during the process of the conveying mechanism conveying ores, the image detection method provided by the present disclosure can be used to perform target detection on the images collected of the currently conveyed ores, so as to determine the positions of the ores according to the target output boxes corresponding to each frame of image obtained, which is helpful to improve the accuracy of ore position determination, and thus the sorting efficiency can be improved during subsequent ore sorting, ensuring the performance of the ore sorting equipment.
[0147] Based on the same inventive concept, the present disclosure also provides an image detection device. As Figure 13 shown, the image detection device 200 includes:
[0148] A first acquisition module 210, configured to acquire a current frame image;
[0149] A second acquisition module 220, configured to, if there is a previous frame image, acquire a first partial image extracted from the bottom of the previous frame image, where the first partial image is obtained based on a first object detection result corresponding to the previous frame image, the first object detection result includes a first detection box corresponding to the first partial image and first position information of the first detection box, the distance between the bottom edge position of the first detection box and the bottom edge position of the previous frame image is less than a first distance, and the first partial image includes the top edge of the first detection box;
[0150] A splicing module 230, configured to splice the first partial image at the top of the current frame image to obtain a first target image;
[0151] A first detection module 240, configured to perform object detection on the first target image to obtain a second object detection result corresponding to the current frame image;
[0152] A first processing module 250, configured to determine and output a target output box corresponding to the current frame image based on the second object detection result and the first position information.
[0153] In some embodiments, the first processing module 250 includes: a first determination unit, configured to respectively determine a plurality of detection boxes corresponding to the first target image and position information of each detection box according to the second object detection result; a first screening unit, configured to determine a target output box corresponding to the current frame image from the plurality of detection boxes based on the position information of each detection box and the first position information; and an output unit, configured to output the target output box.
[0154] In some embodiments, the first screening unit includes: a position determination unit, configured to determine a first target position in the first target image, where the first target position is a position in the first partial image that is at a first distance from the top edge position of the current frame image; a classification unit, configured to respectively determine the category of each detection box according to the position information of each detection box, the first target position, and the first position information according to a classification strategy; and a screening subunit, configured to determine a target output box corresponding to the current frame image based on the category of each detection box; where the classification strategy includes: a first classification strategy: regarding the detection box whose position information overlaps with the first target position as a first category, and the detection boxes of the first category include the detection boxes that overlap with the first position information; a second classification strategy: regarding the detection box whose position information is between the first target position and a second target position as a second category, where the second target position is a position in the current frame image that is at a first distance from the bottom edge position of the current frame image; a third classification strategy: regarding the detection box whose position information is between the top edge position of the first partial image and the top edge position of the current frame image as a third category.
[0155] In some embodiments, the screening subunit includes: a first execution unit configured to use a detection box that is the union of the first category and the second category but the difference set from the third category as the target output box corresponding to the current frame image; and / or a second execution unit configured to determine a third detection box that is the intersection of the first category and the third category among a plurality of detection boxes; a third execution unit configured to determine the area of the region of the third detection box according to the position information of the third detection box; a fourth execution unit configured to determine the intersection area between the third detection box and a reference detection box according to the position information of the third detection box and the third position information; a fifth execution unit configured to determine, based on the position information of a fourth detection box in the previous frame image, the reference detection box for duplicate removal detection and the third position information of the reference detection box in the previous frame image, where the fourth detection box is a detection box whose position information overlaps with the position of the top edge of the first partial image in the previous frame image; and a sixth execution unit configured to use the third detection box as the target output box corresponding to the current frame image if the ratio between the intersection area and the area of the region is less than a specified threshold.
[0156] In some embodiments, the screening subunit further includes: a seventh execution unit configured to use a detection box as a non-target output box corresponding to the current frame image if the distance between the bottom edge position of the detection box and the bottom edge position of the current frame image is less than a first distance; and / or an eighth execution unit configured to use the third detection box as a non-target output box corresponding to the current frame image if the ratio between the intersection area and the area of the region is greater than or equal to the specified threshold.
[0157] In some embodiments, the image detection device 200 further includes: a second processing module configured to respectively determine a second distance between the top edge position of each second detection box and the bottom edge position of the current frame image based on the second position information of each second detection box, where the second detection box is a detection box whose distance between the bottom edge position and the image bottom edge of the current frame image is less than the first distance; a third processing module configured to determine an image extraction height based on the maximum second distance; an extraction module configured to extract a second partial image from the bottom of the current frame image based on a third target position corresponding to the image extraction height; and a saving module configured to save the second partial image to splice it on the top of the next frame image to determine a target output box corresponding to the next frame image.
[0158] In some embodiments, the classification strategy further includes a fourth classification strategy: using a detection box whose position information overlaps with the third target position as the fourth category; the image detection device 200 further includes: a saving module configured to save the position information corresponding to the detection box of the fourth category for determining a non-target output box corresponding to the next frame image.
[0159] In some embodiments, the splicing module includes: a first processing unit configured to splice a first partial image on top of the current frame image to obtain a spliced image; a second determination unit configured to determine the image height of the first partial image; a third determination unit configured to determine a first splicing height of the spliced image according to the image height of the first partial image and the image height of the current frame image; a fourth determination unit configured to determine a second splicing height of a first target image corresponding to the previous frame image; a second processing unit configured to, if the second splicing height is greater than the first splicing height, clear a region image corresponding to the height difference at the top of the first target image corresponding to the previous frame image according to the height difference between the second splicing height and the first splicing height to obtain an intermediate image; and a third processing unit configured to cover the spliced image on the intermediate image to obtain a first target image.
[0160] In some embodiments, the image detection device 200 further includes: a second detection module configured to perform target detection on the current frame image to obtain a third target detection result corresponding to the current frame image if there is no previous frame image; and a fourth processing module configured to determine and output a target output box corresponding to the current frame image based on the third target detection result.
[0161] In some embodiments, the fourth processing module includes: a fifth determination unit configured to respectively determine a plurality of detection boxes corresponding to the current frame image and position information of each detection box according to the third target detection result; and a second screening unit configured to use the detection boxes belonging to the fifth category as the target output box corresponding to the current frame image according to the position information of each detection box; wherein, the detection boxes of the fifth category are the detection boxes whose position information is between the top edge position of the current frame image and a second target position, and the second target position is a position whose distance from the bottom edge position of the current frame image in the current frame image is a first distance.
[0162] Regarding the image detection device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0163] Based on the same inventive concept, as Figure 14As shown, an embodiment of the present disclosure provides an electronic device 300. The electronic device includes: one or more processors 310, a memory 320, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple electronic devices can be connected, and each device provides part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 14 In the figure, a processor 310 is taken as an example.
[0164] The processor 310 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 310 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0165] Among them, the memory 320 stores instructions executable by at least one processor 310, so that at least one processor 310 executes the image detection method shown in the above embodiments.
[0166] The memory 320 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 320 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 320 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0167] The memory 320 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 320 can also include a combination of the above types of memories.
[0168] The electronic device further includes an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected via a bus or other means. Figure 13 Taking the connection via the bus as an example.
[0169] The input device 330 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 340 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0170] Based on the same inventive concept, the present disclosure also provides a computer-readable storage medium storing the following program, and the program is used to execute the image detection method of any of the foregoing embodiments.
[0171] The present disclosure uses specific terms to describe the embodiments of the present disclosure. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present disclosure can be combined appropriately.
[0172] In the context of the present disclosure, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list, and the method or device may also include other steps or elements.
[0173] Similarly, it should be noted that, in order to simplify the expression of the present disclosure and thus help the understanding of one or more embodiments of the application, in the foregoing description of the embodiments of the present disclosure, sometimes multiple features are merged into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the object of the present disclosure are more than the features required to be protected. In fact, the features of the embodiment are less than all the features of the single embodiment disclosed above.
[0174] The basic concepts have been described above. Obviously, for those skilled in the art, the above disclosure is only an example and does not constitute a limitation to the present disclosure. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to the present disclosure. Such modifications, improvements, and corrections are suggested in the present disclosure, so such modifications, improvements, and corrections still fall within the spirit and scope of the embodiments of the present disclosure.
Claims
1. An image detection method, characterized in that: The method comprises: Get the current frame image; If there is a previous frame image, obtain a first partial image extracted from the bottom of the previous frame image, wherein the first partial image is obtained based on a first target detection result corresponding to the previous frame image, the first target detection result includes a first detection frame corresponding to the first partial image and first position information of the first detection frame, the distance between the bottom edge position of the first detection frame and the bottom edge position of the previous frame image is less than a first distance, and the first partial image includes a top edge of the first detection frame; Stitching the first partial image on top of the current frame image to obtain a first target image, including: stitching the first partial image on top of the current frame image to obtain a stitched image; determining the image height of the first partial image; determining the first stitching height of the stitched image according to the image height of the first partial image and the image height of the current frame image; determining the second stitching height of the previous frame image corresponding to the first target image; if the second stitching height is greater than the first stitching height, clearing the top of the previous frame image corresponding to the first target image and the area image corresponding to the height difference according to the height difference between the second stitching height and the first stitching height to obtain an intermediate image; overlaying the stitched image on the intermediate image to obtain the first target image; Performing target detection on the first target image to obtain a second target detection result corresponding to the current frame image; Based on the second target detection result and the first position information, a target output frame corresponding to the current frame image is determined and output.
2. The image detection method according to claim 1, characterized in that: The determining and outputting a target output frame corresponding to the current frame image based on the second target detection result and the first position information includes: Determine, according to the second target detection result, a plurality of detection frames corresponding to the first target image and position information of each detection frame; Based on the position information of each of the detection frames and the first position information, determining a target output frame corresponding to the current frame image from the plurality of detection frames; The target output frame is output.
3. The image detection method according to claim 2, characterized in that: The step of determining a target output frame corresponding to the current frame image from a plurality of the detection frames based on the position information of each of the detection frames and the first position information includes: Determine a first target position in the first target image, where the first target position is a position in the first partial image that is the first distance away from a top edge position of the current frame image; Determining the category of each detection frame according to the position information of each detection frame, the first target position, and the first position information according to a classification strategy; Based on the category of each of the detection frames, determining a target output frame corresponding to the current frame image; The classification strategy includes: A first classification strategy: taking the detection frame whose position information overlaps with the first target position as a first category, wherein the detection frame of the first category includes the detection frame overlapping with the first position information; A second classification strategy: taking the detection frame whose position information is between the first target position and the second target position as a second category, wherein the second target position is a position in the current frame image whose distance from the bottom edge position of the current frame image is the first distance; The third classification strategy: the detection box whose position information is between the top edge position of the first local image and the top edge position of the current frame image is taken as the third category.
4. The image detection method according to claim 3, characterized in that: The step of determining a target output frame corresponding to the current frame image based on the category of each detection frame includes: Using the detection frame that is a union of the first category and the second category but a difference from the third category as a target output frame corresponding to the current frame image; and / or Determine a third detection frame among the plurality of detection frames that intersects the first category and the third category; Determining the area of the third detection frame according to the position information of the third detection frame; Based on the position information of the fourth detection frame in the previous frame image, determine a reference detection frame for deduplication detection in the previous frame image and third position information of the reference detection frame, the fourth detection frame being a detection frame whose position information overlaps with the position of the top edge of the first partial image in the previous frame image; Determining an intersection area between the third detection frame and the reference detection frame according to the position information of the third detection frame and the third position information; and If the ratio of the intersection area to the region area is less than a specified threshold, the third detection frame is used as a target output frame corresponding to the current frame image.
5. The image detection method according to claim 4, characterized in that: The step of determining a target output frame corresponding to the current frame image based on the category of each of the detection frames further includes: If the distance between the bottom edge position of the detection frame and the bottom edge position of the current frame image is less than the first distance, the detection frame is used as a non-target output frame corresponding to the current frame image; and / or If the ratio of the intersection area to the region area is greater than or equal to the specified threshold, the third detection frame is used as a non-target output frame corresponding to the current frame image.
6. The image detection method according to claim 5, characterized in that: The method further comprises: Based on the second position information of each second detection frame, respectively determine a second distance between the top edge position of each second detection frame and the bottom edge position of the current frame image, the second detection frame being the detection frame whose distance between the bottom edge position of the detection frame and the image bottom edge of the current frame image is less than the first distance; Determining an image extraction height based on the largest of the second distances; Extracting a second partial image from the bottom of the current frame image based on a third target position corresponding to the image extraction height; The second partial image is saved to be spliced on top of a next frame image, and a target output frame corresponding to the next frame image is determined.
7. The image detection method according to claim 6, characterized in that: The classification strategy also includes a fourth classification strategy: taking the detection frame whose position information overlaps with the third target position as a fourth category; The method further comprises: The position information corresponding to the detection frame of the fourth category is saved to determine the non-target output frame corresponding to the next frame image.
8. The image detection method according to claim 1, characterized in that: The method further comprises: If the previous frame image does not exist, performing target detection on the current frame image to obtain a third target detection result corresponding to the current frame image; Based on the third target detection result, a target output frame corresponding to the current frame image is determined and output.
9. The image detection method according to claim 8, characterized in that: The determining and outputting a target output frame corresponding to the current frame image based on the third target detection result includes: According to the third target detection result, respectively determine a plurality of detection frames corresponding to the current frame image and position information of each detection frame; According to the position information of each detection frame, taking the detection frame belonging to the fifth category as the target output frame corresponding to the current frame image; Among them, the detection box of the fifth category is a detection box whose position information is between the top edge position of the current frame image and the second target position, and the second target position is a position in the current frame image whose distance from the bottom edge position of the current frame image is the first distance.
10. An image detection device, characterized in that: The device comprises: A first acquisition module, used to acquire a current frame image; a second acquisition module, configured to acquire, if there is a previous frame image, a first partial image extracted from the bottom of the previous frame image, wherein the first partial image is obtained based on a first target detection result corresponding to the previous frame image, the first target detection result includes a first detection frame corresponding to the first partial image and first position information of the first detection frame, the distance between the bottom edge position of the first detection frame and the bottom edge position of the previous frame image is less than a first distance, and the first partial image includes a top edge of the first detection frame; A stitching module, used for stitching the first partial image on the top of the current frame image to obtain a first target image, including: stitching the first partial image on the top of the current frame image to obtain a stitched image; determining the image height of the first partial image; determining the first stitching height of the stitched image according to the image height of the first partial image and the image height of the current frame image; determining the second stitching height of the previous frame image corresponding to the first target image; if the second stitching height is greater than the first stitching height, then according to the height difference between the second stitching height and the first stitching height, clearing the top of the previous frame image corresponding to the first target image and the regional image corresponding to the height difference to obtain an intermediate image; overlaying the stitched image on the intermediate image to obtain the first target image; A first detection module, configured to perform target detection on the first target image to obtain a second target detection result corresponding to the current frame image; The first processing module is used to determine and output a target output frame corresponding to the current frame image based on the second target detection result and the first position information.
11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image detection method according to any one of claims 1 to 9 by executing the computer instructions. 12 . A computer-readable storage medium storing the following program, wherein the program is used to execute the image detection method according to claim 1 .
Citation Information
Patent Citations
Target tracking method and device
CN119722738A