Image detection method and device

By generating target images of different scales and fusing the detection results, the problem of insufficient precision in multi-scale target detection is solved, and the accuracy and efficiency of image detection are improved.

CN115063660BActive Publication Date: 2025-09-19BEIJING JINGDONG QIANSHITECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210851050.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-09-19
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

Existing image detection methods have problems with insufficient detection accuracy when processing multi-scale targets, especially near and far targets, and increase detection resource usage, affecting terminal operation efficiency.

Method used

By acquiring target images of different scales of multiple target objects, an image to be detected of a preset size is generated, and the image detection model is used for detection. The detection results of different scales are fused to generate the final detection result of the original image.

Benefits of technology

The accuracy of multi-scale target detection is improved, the detection area is increased, and the detection effect of the overall target object is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063660B_ABST
    Figure CN115063660B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide an image detection method and device. The image detection method includes: first, in response to obtaining an original image including multiple target objects, obtaining target images of different scales corresponding to the multiple target objects based on the original image, then generating an image to be detected of a preset size based on the original image and the target images of different scales, performing image detection on the image to be detected based on an image detection model, obtaining a detection result corresponding to the image to be detected, and finally generating an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image, which can obtain target images of different scales corresponding to different target objects based on the original image, and splicing the target images of different scales into a single-frame multi-scale image, so that each different target object can complete multi-scale detection within different scale ranges, thereby improving the accuracy of overall target object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the fields of computer technology and Internet technology, specifically to the fields of artificial intelligence technology and image processing technology, and more particularly to image detection methods and devices. Background Art

[0002] With the continuous development and widespread application of deep neural networks, artificial intelligence has rapidly advanced in the field of computer vision, such as image detection. During image detection, if multiple targets are detected within the detection area, and due to imaging principles, such as near targets being larger and distant targets being smaller, the image detection task, while maintaining high accuracy, will require simultaneous detection of both near and distant targets. This large scale difference will affect the final detection results.

[0003] Existing detection methods present significant challenges in multi-scale detection or small object detection. Even with FPN (feature pyramid networks) detection techniques, or through methods such as adjusting the preset detection boxes (anchors), increasing model depth, and increasing the model input scale, deviations in the detector's scale range persist. Furthermore, continuously increasing the preset detection boxes (anchors), scaling the model, or increasing the number of input samples (shape) increases resource usage, impacting terminal operational efficiency. Summary of the Invention

[0004] Embodiments of the present disclosure provide an image detection method, apparatus, electronic device, and computer-readable medium.

[0005] In a first aspect, an embodiment of the present disclosure provides an image detection method, the method comprising: in response to acquiring an original image including multiple target objects, acquiring target images of different scales corresponding to the multiple target objects based on the original image; generating an image to be detected of a preset size based on the original image and the target images of different scales; performing image detection on the image to be detected based on an image detection model, and acquiring a detection result corresponding to the image to be detected; and generating an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0006] In some embodiments, in response to acquiring an original image including multiple target objects, acquiring target images of different scales corresponding to the multiple target objects based on the original image includes: in response to acquiring the original image including multiple target objects, determining multiple expected scale intervals corresponding to the multiple target objects; calculating optimal image scales corresponding to the multiple target objects based on the multiple expected scale intervals; and acquiring target images of different scales corresponding to the multiple target objects based on the optimal image scale and the original image.

[0007] In some embodiments, based on multiple expected scale intervals, the optimal image scales corresponding to multiple target objects are calculated, including: determining the priorities of the multiple target objects based on the application scenario corresponding to the original image; determining the target object with the highest priority from the multiple target objects; and calculating the optimal image scales corresponding to the multiple target objects based on the expected scale interval of the target object with the highest priority.

[0008] In some embodiments, the method further includes: in response to determining that the object corresponding to the optimal image scale is a partial target object, calculating the expanded image scale corresponding to the remaining target objects based on the expected scale range of the remaining target objects; and, based on the optimal image scale and the original image, obtaining target images of different scales corresponding to the multiple target objects, including: obtaining target images of different scales corresponding to the multiple target objects based on the original image, the optimal image scale and the expanded image scale.

[0009] In some embodiments, the detection result corresponding to the image to be detected includes sub-detection results of multiple target images; and, based on the detection result corresponding to the image to be detected and the original image, an image detection result corresponding to the original image is generated, including: fusing the sub-detection results of multiple target images into the original image to obtain a fusion result of multiple target objects in the original image; based on the fusion result, an image detection result corresponding to the original image is generated.

[0010] In some embodiments, the image detection model is acquired based on the following steps: obtaining a training sample set, wherein the training sample set includes a sample image and an image annotation result corresponding to the sample image, wherein the sample image includes multiple images of different scales; using a machine learning method, taking the sample image as input and the image annotation result corresponding to the input image as the expected output, training the initial deep neural network to obtain an image detection model.

[0011] In some embodiments, the method further includes: obtaining multiple image detection results within a preset time period; confirming the analysis results between the multiple image detection results, and judging whether the analysis results meet the preset conditions; in response to determining that the analysis results do not meet the preset conditions, adjusting the expected scale interval corresponding to the target object.

[0012] In a second aspect, an embodiment of the present disclosure provides an image detection device, which includes: an acquisition module, configured to, in response to acquiring an original image including multiple target objects, acquire target images of different scales corresponding to the multiple target objects based on the original image; a first generation module, configured to generate an image to be detected of a preset size based on the original image and the target images of different scales; a detection module, configured to perform image detection on the image to be detected based on an image detection model, and obtain a detection result corresponding to the image to be detected; and a second generation module, configured to generate an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0013] In some embodiments, the acquisition module includes: a determination unit, configured to determine multiple expected scale intervals corresponding to the multiple target objects in response to acquiring an original image including multiple target objects; a calculation unit, configured to calculate the optimal image scale corresponding to the multiple target objects based on the multiple expected scale intervals; and an acquisition unit, configured to acquire target images of different scales corresponding to the multiple target objects based on the optimal image scale and the original image.

[0014] In some embodiments, the computing unit is further configured to: determine the priorities of multiple target objects based on the application scenario corresponding to the original image; determine the target object with the highest priority from the multiple target objects; and calculate the optimal image scale corresponding to the multiple target objects based on the expected scale range of the target object with the highest priority.

[0015] In some embodiments, the computing unit is further configured to: in response to determining that the object corresponding to the optimal image scale is a partial target object, calculate the expanded image scale corresponding to the remaining target objects based on the expected scale range of the remaining target objects; and the acquiring unit is further configured to: acquire target images of different scales corresponding to multiple target objects based on the original image, the optimal image scale, and the expanded image scale.

[0016] In some embodiments, the detection result corresponding to the image to be detected includes sub-detection results of multiple target images; and the second generation module is further configured to: fuse the sub-detection results of multiple target images into the original image to obtain the fusion result of multiple target objects in the original image; based on the fusion result, generate the image detection result corresponding to the original image.

[0017] In some embodiments, the image detection model is acquired based on the following steps: obtaining a training sample set, wherein the training sample set includes a sample image and an image annotation result corresponding to the sample image, wherein the sample image includes multiple images of different scales; using a machine learning method, taking the sample image as input and the image annotation result corresponding to the input image as the expected output, training the initial deep neural network to obtain an image detection model.

[0018] In some embodiments, the device also includes: a judgment module and an adjustment module; the acquisition module is further configured to: acquire multiple image detection results within a preset time period; the judgment module is further configured to: confirm the analysis results between the multiple image detection results, and judge whether the analysis results meet the preset conditions; the adjustment module is further configured to: in response to determining that the analysis results do not meet the preset conditions, adjust the expected scale interval corresponding to the target object.

[0019] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by one or more processors, the one or more processors implement the image detection method described in any embodiment of the first aspect.

[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the image detection method as described in any embodiment of the first aspect.

[0021] The image detection method provided by the embodiment of the present disclosure is as follows: the execution subject first responds to obtaining an original image including multiple target objects, obtains target images of different scales corresponding to the multiple target objects based on the original image, then generates an image to be detected of a preset size based on the original image and the target images of different scales, performs image detection on the image to be detected based on an image detection model, obtains a detection result corresponding to the image to be detected, and finally generates an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image. The method can obtain target images of different scales corresponding to different target objects based on the original image, and splice the target images of different scales into a single-frame multi-scale image, so that each different target object can complete multi-scale detection within different scale ranges, thereby adding more scale detection areas and fusing the detection results under multiple scales as the final image detection result, so that more target objects are within a scale range with a better image detection model, thereby improving the accuracy of overall target object detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings:

[0023] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0024] Figure 2 is a flow chart of an embodiment of an image detection method according to the present disclosure;

[0025] Figure 3 is a schematic diagram of an application scenario of the image detection method according to the present disclosure;

[0026] Figure 4 is a flow chart of an embodiment of acquiring target images of different scales according to the present disclosure;

[0027] Figure 5 is a flowchart of another embodiment of acquiring target images of different scales according to the present disclosure;

[0028] Figure 6 is a flow chart of one embodiment of obtaining an image detection model according to the present disclosure;

[0029] Figure 7 is a flow chart of an embodiment of adjusting the expected scale interval according to the present disclosure;

[0030] Figure 8 is a structural schematic diagram of an embodiment of an image detection device according to the present disclosure;

[0031] Figure 9 It is a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0032] The present disclosure is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant disclosure and are not intended to limit the disclosure. It should also be noted that, for ease of description, only portions relevant to the relevant disclosure are shown in the accompanying drawings.

[0033] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0034] Figure 1 An exemplary system architecture 100 is shown to which the image detection method and apparatus according to embodiments of the present disclosure may be applied.

[0035] like Figure 1As shown, system architecture 100 may include terminal devices 104, 105, 106, a network 107, and servers 101, 102, 103. Network 107 is used to provide a medium for communication links between terminal devices 104, 105, 106 and servers 101, 102, 103. Network 107 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0036] Users can interact with servers 101, 102, and 103 belonging to the same server cluster via terminal devices 104, 105, and 106 via network 107 to receive or send information. Various applications can be installed on terminal devices 104, 105, and 106, such as item display applications, data analysis applications, and search applications.

[0037] Terminal devices 104, 105, and 106 can be either hardware or software. If the terminal device is hardware, it can be any electronic device with a display screen that supports communication with the server, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. If the terminal device is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules, or as a single software program or software module. This is not specifically limited here.

[0038] The terminal devices 104, 105, and 106 can obtain an original image including multiple target objects, obtain target images of different scales corresponding to the multiple target objects based on the original image, and then generate an image to be detected of a preset size based on the original image and the target images of different scales, and perform image detection on the image to be detected based on the image detection model to obtain the detection result corresponding to the image to be detected, and finally generate an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0039] Servers 101, 102, and 103 may be servers that provide various services, such as backend servers that receive requests sent by terminal devices that establish communication connections therewith. The backend servers may receive and analyze the requests sent by the terminal devices and generate processing results.

[0040] Servers 101, 102, and 103 can obtain an original image including multiple target objects, obtain target images of different scales corresponding to the multiple target objects based on the original image, and then generate an image to be detected of a preset size based on the original image and the target images of different scales, and perform image detection on the image to be detected based on the image detection model to obtain a detection result corresponding to the image to be detected, and finally generate an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0041] It should be noted that a server can be either hardware or software. When the server is hardware, it can be various electronic devices that provide various services to terminal devices. When the server is software, it can be implemented as multiple software programs or software modules that provide various services to terminal devices, or it can be implemented as a single software program or software module that provides various services to terminal devices. This is not specifically limited here.

[0042] It should be noted that the image detection method provided in the embodiments of the present disclosure can be executed by the terminal devices 104, 105, 106 or the servers 101, 102, 103. Accordingly, the image detection apparatus is provided in the terminal devices 104, 105, 106 or the servers 101, 102, 103.

[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0044] Continue to refer Figure 2 , shows a process 200 of an embodiment of an image detection method according to the present disclosure. The image detection method includes the following steps:

[0045] Step 210 : In response to obtaining an original image including a plurality of target objects, obtaining target images of different scales corresponding to the plurality of target objects based on the original image.

[0046] In this step, the execution subject (e.g. Figure 1 The terminal devices 104, 105, 106 or servers 101, 102, 103 in the system can obtain an original image that requires image detection. The original image can be an image obtained by photographing multiple target objects, including multiple target objects to be detected at different scales. Since different target objects can be accurately detected at a certain image scale, and different target objects can correspond to different image scales, after the above-mentioned execution entity obtains the original image, it can crop the original image according to the image scales corresponding to different target objects before performing image detection on the target object, thereby obtaining multiple target images of different scales, each of which can include at least one target object.

[0047] As an example, if the original image includes three target objects, namely A, B, and C, the above-mentioned execution entity can obtain the image scale corresponding to A as A1, the image scale corresponding to B as B1, and the image scale corresponding to C as C1. The above-mentioned execution entity can crop the original image according to the image scale A1 to obtain target image 1 including target object A and with an image scale of A1. It can also crop the original image according to the image scale B1 to obtain target image 2 including target object B and with an image scale of B1. It can also crop the original image according to the image scale C1 to obtain target image 3 including target object C and with an image scale of C1. In this way, the above-mentioned execution entity can crop the original image according to multiple target objects to obtain target images of different scales.

[0048] Step 220 : Generate an image to be detected of a preset size based on the original image and target images of different scales.

[0049] In this step, after obtaining target images of different scales, the execution entity may obtain an image detection model for detecting the image and obtain input image information for the image detection model. The input image information may include the image size of the input image and may also include the image style of the input image, which may represent the layout format of multiple images in the input image. The execution entity may crop, splice, and layout the original image and target images of different scales based on the input image information of the image detection model to generate an image to be detected of a preset size. The preset size may be the same as the image size of the input image of the image detection model.

[0050] As an example, after the above-mentioned execution entity obtains the input image information of the image detection model, it determines the image style of the input image and the image size of the input image. If 1 target image is obtained, the image style of the input image is arranged in vertical columns, and the image size of the input image is N*N, then the image sizes of the original image and the target image can be cropped to N / 2*N respectively, and the cropped original image and target image are arranged and spliced ​​in the vertical column direction in turn to obtain an image to be detected of a preset size; if 2 target images are obtained, the image sizes of the original image and the target image can be cropped to N / 3*N respectively, and the cropped original image and the 2 target images are arranged and spliced ​​in the vertical column direction in turn to obtain an image to be detected of a preset size; or, if 3 target images are obtained, the image style of the input image is arranged in a grid pattern, and the image size of the input image is N*N, then the image sizes of the original image and the 3 target images can be cropped to N / 2*N / 2 respectively, and the cropped original image and the 3 target images are arranged and spliced ​​in a grid pattern to obtain an image to be detected of a preset size.

[0051] Step 230: Perform image detection on the image to be detected based on the image detection model to obtain a detection result corresponding to the image to be detected.

[0052] In this step, after the execution subject obtains the image to be detected, it can input the image to be detected into the image detection model, perform image detection on the image to be detected through the image detection model, and output the detection result of the image to be detected.

[0053] The image detection model can perform image detection on the original image and target image included in the image to be detected respectively, and output the results of each target object in the original image and the results of each target object in the target image, thereby obtaining the detection results corresponding to the image to be detected.

[0054] Step 240 : Based on the detection result corresponding to the image to be detected and the original image, generate an image detection result corresponding to the original image.

[0055] In this step, after obtaining the detection results corresponding to the image to be detected, the execution entity may fuse the results of the original image and the target image in the image to be detected. For each target object, the execution entity obtains the results in the original image and the target image respectively, and then analyzes and fuses the results in the original image and the target image to obtain the corresponding detection result. Thus, through the analysis and information fusion of each target object, the execution entity obtains the image detection result corresponding to the original image. This image detection result may include the detection results and detection scores of multiple target objects.

[0056] As an optional implementation method, the detection result corresponding to the image to be detected may include sub-detection results of multiple target images, and the above-mentioned step 240, based on the detection result corresponding to the image to be detected and the original image, generates the image detection result corresponding to the original image, which can be implemented based on the following steps: fusing the sub-detection results of multiple target images into the original image to obtain the fusion result of multiple target objects in the original image; based on the fusion result, generating the image detection result corresponding to the original image.

[0057] Specifically, after the execution subject obtains the detection result of the image to be detected, it can determine the sub-detection results corresponding to the multiple target images respectively. Since the scales of different target images are inconsistent, the multiple target images need to be normalized and normalized to the scale of the original image. Then the execution subject fuses the normalized sub-detection results of each target image into the original image to obtain a fusion result of the sub-detection results of multiple target objects in the multiple target images in the original image, that is, the fusion result can include the sub-detection results of each target object in different target images and the sub-detection results in the original image. After the execution subject obtains the fusion result of multiple target objects in the original image, it performs non-maximum suppression processing on the sub-detection result of each target object in the fusion result to obtain the final detection result of each target object, thereby generating the image detection result corresponding to the original image.

[0058] In this implementation, the image detection result corresponding to the original image is obtained by fusing and processing the sub-detection results of multiple target images in the image to be detected, thereby ensuring the accuracy of the detection result of each target object and improving the accuracy of the image detection result.

[0059] Continue to see Figure 3 , Figure 3 This is a schematic diagram of an application scenario of the image detection method according to this embodiment. Figure 3 In an application scenario, a user captures an original image 301 through a terminal, and the terminal obtains the original image 301, which includes multiple target objects. Based on the original image 301 and the multiple target objects, the terminal obtains target images 302 of different scales corresponding to the multiple target objects, and then splices the captured original image 301 and the target images 302 of different scales into an image to be detected 303 of a preset size. The terminal can input the image to be detected 303 into an image detection model for image detection, and obtain a detection result corresponding to the image to be detected 303. The detection result includes the detection results of each target object in the original image 301 and the detection results of each target object in the target image 302. Finally, based on the detection results corresponding to the image to be detected 303 and the original image 301, an image detection result corresponding to the original image 301 is generated.

[0060] The image detection method provided by the embodiment of the present disclosure is as follows: the execution subject first responds to obtaining an original image including multiple target objects, obtains target images of different scales corresponding to the multiple target objects based on the original image, then generates an image to be detected of a preset size based on the original image and the target images of different scales, performs image detection on the image to be detected based on an image detection model, obtains a detection result corresponding to the image to be detected, and finally generates an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image. The method can obtain target images of different scales corresponding to different target objects based on the original image, and splice the target images of different scales into a single-frame multi-scale image, so that each different target object can complete multi-scale detection within different scale ranges, thereby adding more scale detection areas and fusing the detection results under multiple scales as the final image detection result, so that more target objects are within a scale range with a better image detection model, thereby improving the accuracy of overall target object detection.

[0061] refer to Figure 4 , Figure 4 A flowchart 400 illustrating an embodiment of acquiring target images at different scales may include the following steps:

[0062] Step 410 : In response to acquiring an original image including a plurality of target objects, determining a plurality of expected scale intervals corresponding to the plurality of target objects.

[0063] In this step, after obtaining an original image including multiple target objects, the execution entity may perform image processing on the original image to obtain category information for the multiple target objects. This category information may represent the type of the target object, such as a traffic light, a pedestrian, etc. Based on the category information for each target object, the execution entity may search for the corresponding expected scale interval in a locally stored table of correspondences between category information and expected scale intervals. Each category information may correspond to an expected scale interval, which may represent a scale interval that enables a good detection effect for the target object in the detection model. That is, within this expected scale interval, the target object can achieve a good detection effect in the detection model. Thus, the execution entity obtains the expected scale interval corresponding to each target object, thereby obtaining multiple expected scale intervals corresponding to the multiple target objects.

[0064] The expected scale interval corresponding to each target object may be the expected scale intervals corresponding to different target objects statistically calculated during the training of the image detection model.

[0065] Step 420 : Calculate optimal image scales corresponding to multiple target objects based on the multiple expected scale intervals.

[0066] In this step, after obtaining multiple expected scale intervals corresponding to multiple target objects, the execution entity calculates the intersection scale between the multiple expected scale intervals. Specifically, the execution entity may perform the intersection calculation for each expected scale interval to determine the intersection scale between the multiple expected scale intervals. The execution entity then uses the calculated intersection scale as the optimal image scale for the multiple target objects. In other words, this optimal image scale ensures that all target objects can be detected effectively in the image detection model.

[0067] Alternatively, the execution entity performs intersection calculation on multiple expected scale intervals. If some of the expected scale intervals have intersection scales, and the remaining expected scale intervals do not have intersection with the intersection scale, the calculated intersection scale is used as the optimal image scale corresponding to the multiple target objects. That is, the optimal image scale can enable all of the target objects corresponding to the intersection to have better detection effects in the image detection model, and the optimal image scale is calculated based on the largest number of expected scale intervals.

[0068] As an example, if the original image includes three target objects, namely A, B, and C, the above-mentioned execution entity can obtain the expected scale interval corresponding to A as [30-50], the expected scale interval corresponding to B as [35-60], and the expected scale interval corresponding to C as [20-45]. The above-mentioned execution entity can first calculate the intersection between the expected scale intervals corresponding to A and B to obtain a first intersection scale [35-50]. Then, the above-mentioned execution entity can calculate the intersection between the first intersection scale and the expected scale interval corresponding to C to obtain an intersection scale [35-45] between multiple expected scale intervals. The optimal image scale corresponding to A, B, and C is determined to be [35-45]. The optimal image scale of [35-45] can enable all target objects to correspond to better detection effects in the image detection model.

[0069] As an example, if the original image includes three target objects, namely A, B, and C, the above-mentioned execution entity can obtain the expected scale interval corresponding to A as [30-50], the expected scale interval corresponding to B as [35-60], and the expected scale interval corresponding to C as [10-25]. The above-mentioned execution entity can first calculate the intersection between the expected scale intervals corresponding to A and B to obtain the first intersection scale [35-50]. Then, the above-mentioned execution entity can calculate the intersection between the first intersection scale and the expected scale interval corresponding to C. By calculation, it is determined that there is no intersection between the first intersection scale and the expected scale interval corresponding to C. Then, the optimal image scale corresponding to A, B, and C is determined to be [35-50]. The optimal image scale of [35-50] can enable A and B to have a better corresponding detection effect in the image detection model and include as many target objects as possible.

[0070] As an optional implementation, the above step 420, calculating the optimal image scales corresponding to the multiple target objects based on the multiple expected scale intervals, may include the following steps:

[0071] Step 1: Determine the priorities of multiple target objects based on the application scenario corresponding to the original image.

[0072] Specifically, after the above-mentioned execution entity obtains the original image, it can perform scene analysis on the original image to determine the application scene corresponding to the original image. The application scene can represent the current focus scene of the original image, etc. For example, if the original image is taken on a sidewalk, the application scene corresponding to the original image can be a pedestrian driving scene.

[0073] After determining the application scenario corresponding to the original image, the execution entity determines the priority of different target objects in the original image based on the application scenario. The priority represents the level of attention the target object receives in the application scenario. For example, if the application scenario corresponding to the original image is a pedestrian driving scene and the original image includes a sidewalk signal light and a motorway signal light, the execution entity may determine that the sidewalk signal light has a higher priority than the motorway signal light based on the pedestrian driving scene.

[0074] Step 2: Determine the target object with the highest priority from multiple target objects.

[0075] Specifically, after the execution subject obtains the priorities of multiple target objects, it can determine the target object with the highest priority from the multiple target objects.

[0076] Step 3: Based on the expected scale interval of the target object with the highest priority, calculate the optimal image scale corresponding to multiple target objects.

[0077] Specifically, after the execution entity determines the target object with the highest priority, it can calculate the intersection between multiple expected scale intervals based on the expected scale interval of the target object. That is, based on the expected scale interval of the target object with the highest priority, the intersection between multiple expected scale intervals is determined, so that as many target objects as possible correspond to the intersection.

[0078] As an example, if the original image includes three target objects, namely A, B, and C, the above-mentioned execution entity can obtain the expected scale interval corresponding to A as [30-50], the expected scale interval corresponding to B as [35-60], and the expected scale interval corresponding to C as [10-25]. If A has the highest priority, the above-mentioned execution entity can first calculate the intersection between the expected scale intervals corresponding to A and B based on the expected scale interval of A to obtain the first intersection scale [35-50]. Then, the above-mentioned execution entity can calculate the intersection between the first intersection scale and the expected scale interval corresponding to C. By calculation, it is determined that there is no intersection between the first intersection scale and the expected scale interval corresponding to C. Then, the optimal image scale corresponding to A, B, and C is determined to be [35-50]. The optimal image scale of [35-50] can enable A and B to have better detection effects in the image detection model and include as many target objects as possible.

[0079] In this implementation, by performing calculation based on the target object with the highest priority, the optimal image scale can correspond to more target objects and target objects with higher priorities, thereby improving the accuracy of the optimal image scale.

[0080] Step 430 : acquiring target images of different scales corresponding to a plurality of target objects based on the optimal image scale and the original image.

[0081] In this step, after the execution subject determines the optimal image scales corresponding to the multiple target objects, the original image can be cropped according to the optimal image scales to obtain target images of different scales corresponding to the multiple target objects.

[0082] In this implementation, by calculating multiple expected scale intervals, the optimal image scale corresponding to multiple target objects is determined to obtain target images of different scales corresponding to multiple target objects, thereby improving the targeting of target objects in the target image and enabling more target objects to be within a better detection range, thereby improving the accuracy of overall target object detection.

[0083] refer to Figure 5 , Figure 5 A flowchart 500 illustrating another embodiment of acquiring target images of different scales may include the following steps:

[0084] Step 510 : In response to acquiring an original image including a plurality of target objects, determining a plurality of expected scale intervals corresponding to the plurality of target objects.

[0085] Step 510 of this embodiment can be performed in accordance with Figure 4 Step 410 in the illustrated embodiment is performed in a similar manner and will not be described in detail here.

[0086] Step 520 : Calculate optimal image scales corresponding to multiple target objects based on the multiple expected scale intervals.

[0087] Step 520 of this embodiment can be performed in accordance with Figure 4 Step 420 in the illustrated embodiment is performed in a similar manner and will not be described in detail here.

[0088] Step 530 : In response to determining that the object corresponding to the optimal image scale is a partial target object, based on the expected scale interval of the remaining target objects, calculate the expanded image scale corresponding to the remaining target objects.

[0089] In this step, after determining the optimal image scale, the execution entity may analyze the target object corresponding to the optimal image scale to determine whether the target object is the entire target object in the original image. If it is determined that the object corresponding to the optimal image scale is only a partial target object, the execution entity may determine the target objects that are not within the optimal image scale based on the optimal image scale.

[0090] Furthermore, the execution entity may analyze the target objects corresponding to the optimal image scale to determine the target objects corresponding to the optimal image scale, and compare the optimal image scale with the expected scale interval of each target object to determine whether the optimal image scale is a boundary value of the expected scale interval for a target object. If the optimal image scale is determined to be a boundary value of the expected scale interval for a target object, the target object is determined. The execution entity may treat target objects that are not within the optimal image scale and target objects whose expected scale intervals have a boundary value of the optimal image scale as remaining target objects and exclude them from being objects corresponding to the optimal image scale.

[0091] After the execution entity determines the remaining target objects, it performs calculations based on the expected scale intervals of the remaining target objects to determine the expanded image scales corresponding to the remaining target objects. The expanded image scale can represent the image scale corresponding to the remaining target objects, and can be the intersection scale of the expected scale intervals of multiple remaining target objects, or can be the corresponding image scale of each remaining target object.

[0092] Specifically, after the execution entity determines the remaining target objects, if there are multiple remaining target objects, the expected scale intervals corresponding to each remaining target object can be intersected to determine the intersection scale between the expected scale intervals corresponding to the multiple remaining target objects. The execution entity then uses the calculated intersection scale as the expanded image scale corresponding to the remaining target objects. That is, the expanded image scale corresponding to the remaining target objects can enable all remaining target objects to achieve better detection results in the image detection model. If there is only one remaining target object, the expanded image scale corresponding to the remaining target object can be directly determined based on the expected scale interval of the remaining target object.

[0093] Step 540 : Acquire target images of different scales corresponding to a plurality of target objects based on the original image, the optimal image scale, and the expanded image scale.

[0094] In this step, after the above-mentioned execution entity determines the optimal image scale and the expanded image scale corresponding to the multiple target objects, the original image can be cropped according to the optimal image scale, and the original image can also be cropped according to the expanded image scale to obtain target images of different scales corresponding to the multiple target objects.

[0095] In this implementation, by calculating multiple expected scale intervals, the optimal image scale and expanded image scale corresponding to multiple target objects are determined to obtain target images of different scales corresponding to multiple target objects, so that all target objects in the original image can have corresponding image scales and target images of different scales, which improves the targeting of target objects in the target image and can make all target objects within a better detection range, thereby improving the accuracy of overall target object detection.

[0096] refer to Figure 6 , Figure 6 A flowchart 600 illustrating an embodiment of obtaining an image detection model may include the following steps:

[0097] Step 610: Acquire a training sample set, wherein the training sample set includes sample images and image annotation results corresponding to the sample images.

[0098] In this step, the execution entity may obtain sample images for training the model, where the sample images may be images spliced ​​together from multiple images of different scales, and annotate each sample image to obtain image annotation results for each sample image.

[0099] In step 620 , a machine learning method is used to train an initial deep neural network by taking the sample image as input and the image annotation result corresponding to the input image as the desired output to obtain an image detection model.

[0100] In this step, after the execution subject obtains the training sample set, it obtains the initial deep neural network. The execution subject can use machine learning methods to train the initial deep neural network based on the training sample set to obtain an image detection model.

[0101] Specifically, the execution subject may take the sample image as input, and obtain corresponding prediction information after processing by an initial deep neural network. The initial deep neural network may be any of various existing neural networks.

[0102] If the prediction information does not meet the constraints, the network parameters of the initial deep neural network are adjusted and the sample image is re-input to continue training. If the prediction information meets the constraints, model training is completed, and an image detection model is obtained. The constraint condition can be that the difference between the prediction information and the image annotation result meets a preset threshold. The preset threshold can be pre-set based on experience and is not specifically limited in this disclosure.

[0103] In this implementation, an initial deep neural network is trained with the acquired sample images to obtain an image detection model, which can improve the accuracy and extensiveness of the image detection model in detecting images.

[0104] refer to Figure 7 , Figure 7 A flowchart showing an embodiment of adjusting the expected scale interval may include the following steps:

[0105] Step 710: Acquire multiple image detection results within a preset time period.

[0106] In this step, the above-mentioned execution entity can read the image detection results output by the image detection model, including multiple image detection results within a preset time period.

[0107] Step 720 , confirming the analysis results among the multiple image detection results, and determining whether the analysis results meet the preset conditions.

[0108] In this step, after obtaining multiple image detection results, the execution entity may perform data analysis on the multiple image detection results, evaluate the image detection results, and determine whether the analysis results meet a preset condition. The preset condition may indicate that the image detection results are within a stable range, such as a preset interval of detection scores. Specifically, the execution entity may calculate the detection scores of the multiple image detection results and determine whether the calculated detection scores are within a stable range.

[0109] Step 730 : In response to determining that the analysis result does not meet the preset condition, adjusting the expected scale interval corresponding to the target object.

[0110] In this step, the execution subject determines that the analysis result does not meet the preset conditions, and then determines that there is a problem with the expected scale interval of the target object, triggering the adjustment logic of the initial expected scale interval of the target object, and adjusting the expected scale interval corresponding to the target object, so that the target object with a larger scale influence is adjusted to a better scale interval.

[0111] In this implementation, by analyzing the image detection results within a preset time period and adjusting the expected scale interval corresponding to the target object when the analysis results do not meet the preset conditions, the target objects with better tolerance can maintain a more tolerant scale interval, and the target objects with greater scale influence can be adjusted to a better scale interval. Finally, the overall multi-scale detection is completed by combining the multi-scale detection method.

[0112] Further references Figure 8 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image detection device. Figure 2 The method embodiment shown corresponds to the embodiment shown.

[0113] like Figure 8 As shown, the image detection device 800 of this embodiment may include: an acquisition module 810 , a first generation module 820 , a detection module 830 and a second generation module 840 .

[0114] The acquisition module 810 is configured to, in response to acquiring an original image including multiple target objects, acquire target images of different scales corresponding to the multiple target objects based on the original image;

[0115] The first generating module 820 is configured to generate an image to be detected of a preset size based on the original image and target images of different scales;

[0116] The detection module 830 is configured to perform image detection on the image to be detected based on the image detection model and obtain a detection result corresponding to the image to be detected;

[0117] The second generating module 840 is configured to generate an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0118] In some optional implementations of this embodiment, the acquisition module 810 includes: a determination unit, configured to determine multiple expected scale intervals corresponding to the multiple target objects in response to acquiring an original image including multiple target objects; a calculation unit, configured to calculate optimal image scales corresponding to the multiple target objects based on the multiple expected scale intervals; and an acquisition unit, configured to acquire target images of different scales corresponding to the multiple target objects based on the optimal image scale and the original image.

[0119] In some optional implementations of this embodiment, the computing unit is further configured to: determine the priorities of multiple target objects based on the application scenario corresponding to the original image; determine the target object with the highest priority from the multiple target objects; and calculate the optimal image scale corresponding to the multiple target objects based on the expected scale range of the target object with the highest priority.

[0120] In some optional implementations of this embodiment, the computing unit is further configured to: in response to determining that the object corresponding to the optimal image scale is a partial target object, calculate the expanded image scale corresponding to the remaining target objects based on the expected scale range of the remaining target objects; and the acquiring unit is further configured to: acquire target images of different scales corresponding to multiple target objects based on the original image, the optimal image scale, and the expanded image scale.

[0121] In some optional implementations of this embodiment, the detection result corresponding to the image to be detected includes sub-detection results of multiple target images; and the second generation module 840 is further configured to: fuse the sub-detection results of multiple target images into the original image to obtain the fusion result of multiple target objects in the original image; based on the fusion result, generate the image detection result corresponding to the original image.

[0122] In some optional implementations of the present embodiment, the image detection model is obtained based on the following steps: obtaining a training sample set, wherein the training sample set includes a sample image and an image annotation result corresponding to the sample image, wherein the sample image includes multiple images of different scales; using a machine learning method, taking the sample image as input and the image annotation result corresponding to the input image as the expected output, training the initial deep neural network to obtain an image detection model.

[0123] In some optional implementations of this embodiment, the device also includes: a judgment module and an adjustment module; the acquisition module is further configured to: obtain multiple image detection results within a preset time period; the judgment module is further configured to: confirm the analysis results between the multiple image detection results, and judge whether the analysis results meet the preset conditions; the adjustment module is further configured to: in response to determining that the analysis results do not meet the preset conditions, adjust the expected scale interval corresponding to the target object.

[0124] The image detection device provided by the above-mentioned embodiment of the present disclosure, the above-mentioned execution body first responds to obtaining an original image including multiple target objects, obtains target images of different scales corresponding to the multiple target objects based on the original image, and then generates an image to be detected of a preset size based on the original image and the target images of different scales, and performs image detection on the image to be detected based on the image detection model to obtain the detection result corresponding to the image to be detected, and finally generates an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image. It can obtain target images of different scales corresponding to different target objects according to the original image, and splice the target images of different scales into a single-frame multi-scale image, so that each different target object can complete multi-scale detection within different scale ranges, thereby adding more scale detection areas and fusing the detection results under multiple scales as the final image detection result, so that more target objects are within a better scale range of the image detection model, thereby improving the accuracy of overall target object detection.

[0125] Those skilled in the art will appreciate that the above-mentioned device also includes some other well-known structures, such as a processor, a memory, etc. In order to unnecessarily obscure the embodiments of the present disclosure, these well-known structures are not described in detail. Figure 8 Not shown.

[0126] Reference below Figure 9 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as smart screens, laptop computers, PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0127] like Figure 9 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0128] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 9 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0129] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed. It should be noted that the computer-readable medium of the embodiment of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wire, optical cable, RF (radio frequency), etc., or any suitable combination thereof.

[0130] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0131] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0132] The units involved in the embodiments described in this application may be implemented in software or hardware. The units described may also be provided in a processor. For example, it may be described as follows: a processor includes an acquisition module, a first generation module, a detection module, and a second generation module, wherein the names of these modules do not, in some cases, constitute limitations on the modules themselves.

[0133] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device; or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: in response to acquiring an original image including multiple target objects, acquires target images of different scales corresponding to the multiple target objects based on the original image; generates an image to be detected of a preset size based on the original image and the target images of different scales; performs image detection on the image to be detected based on an image detection model, and acquires a detection result corresponding to the image to be detected; and generates an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

[0134] The above description is merely a preferred embodiment of the present disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by mutually replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An image detection method, comprising: In response to acquiring an original image including a plurality of target objects, determining an expected scale interval corresponding to each target object in the plurality of target objects; Calculating optimal image scales corresponding to the multiple target objects based on the expected scale intervals corresponding to the target objects; Based on the optimal image scale and the original image, acquiring target images of different scales corresponding to the multiple target objects; Generating an image to be detected of a preset size based on the original image and the target images of different scales; Performing image detection on the image to be detected based on the image detection model to obtain a detection result corresponding to the image to be detected; Based on the detection result corresponding to the image to be detected and the original image, an image detection result corresponding to the original image is generated.

2. The method according to claim 1, wherein The calculating the optimal image scales corresponding to the plurality of target objects based on the expected scale intervals corresponding to the target objects includes: Determining priorities of the multiple target objects based on an application scenario corresponding to the original image; determining a target object with the highest priority from the multiple target objects; Based on the expected scale interval of the target object with the highest priority, optimal image scales corresponding to the multiple target objects are calculated.

3. The method according to claim 1, further comprising: In response to determining that the object corresponding to the optimal image scale is a partial target object, calculating an expanded image scale corresponding to the remaining target object based on an expected scale interval of the remaining target object; as well as, The acquiring, based on the optimal image scale and the original image, target images of different scales corresponding to the multiple target objects includes: Based on the original image, the optimal image scale, and the expanded image scale, target images of different scales corresponding to the multiple target objects are acquired.

4. The method according to claim 1, wherein The detection result corresponding to the image to be detected includes sub-detection results of multiple target images; as well as, The generating, based on the detection result corresponding to the image to be detected and the original image, an image detection result corresponding to the original image, includes: fusing the sub-detection results of the multiple target images into the original image to obtain a fusion result of the multiple target objects in the original image; Based on the fusion result, an image detection result corresponding to the original image is generated.

5. The method according to any one of claims 1 to 4, wherein: The image detection model is obtained based on the following steps: Acquire a training sample set, wherein the training sample set includes sample images and image annotation results corresponding to the sample images, wherein the sample images include multiple images of different scales; By using a machine learning method, the sample image is taken as input, and the image annotation result corresponding to the input image is taken as the expected output, and an initial deep neural network is trained to obtain the image detection model.

6. The method according to claim 1, wherein The method further comprises: Obtain multiple image detection results within a preset time period; Determine the analysis result among the multiple image detection results, and judge whether the analysis result meets the preset conditions; In response to determining that the analysis result does not meet the preset condition, the expected scale interval corresponding to the target object is adjusted.

7. An image detection device, comprising: The acquisition module includes: a determination unit configured to, in response to acquiring an original image including a plurality of target objects, determine an expected scale interval corresponding to each of the plurality of target objects; a calculation unit configured to calculate an optimal image scale corresponding to the plurality of target objects based on the expected scale interval corresponding to each of the target objects; and an acquisition unit configured to acquire target images of different scales corresponding to the plurality of target objects based on the optimal image scale and the original image; A first generating module is configured to generate an image to be detected of a preset size based on the original image and the target images of different scales; A detection module is configured to perform image detection on the image to be detected based on an image detection model, and obtain a detection result corresponding to the image to be detected; The second generating module is configured to generate an image detection result corresponding to the original image based on the detection result corresponding to the image to be detected and the original image.

8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method for detecting defects in multi-scale images and computing device utilizing method

    US20220198228A1

  • Method and apparatus for detecting object in fisheye image, and storage medium

    WO2022000862A1