Target detection method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202211497056.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-25
AI Technical Summary
然而,该模型读取图像文件,采用该模型进行目标检测的耗时高,导致单张图像的处理耗时较高,目标检测的效率较低
[0106]本公开实施例提供的目标检测方法,可以获取第一待检测图像,之后,对所述第一待检测图像进行多尺度变换,得到所述第一待检测图像的第一图像集合,然后,创建子进程,以使所述子进程基于所述第一图像集合中的图像确定所述第一待检测图像中的目标对象的第一检测结果,最后,基于所述第一检测结果,确定所述第一待检测图像中的所述目标对象的最终检测结果。由此方法,可以通过创建的子进程基于多尺度变换得到的第一图像集合中的图像,来确定第一待检测图像中的目标对象的第一检测结果,进而获得第一待检测图像中的目标对象的最终检测结果,由此,提高了目标检测的效率。
Smart Images

Figure CN115908336B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a target detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] Object detection refers to processing a given image using a specific strategy to determine whether it contains a target object (such as a face). If so, the detection results, such as the target's location and confidence level, are returned.
[0003] One existing object detection method uses the MTCNN (Multi-task convolutional neural network) model to process a single image to obtain the final detection result. However, this model is time-consuming to read image files and perform object detection, resulting in high processing time for a single image and low object detection efficiency. Summary of the Invention
[0004] In view of this, in order to solve some or all of the above-mentioned technical problems, the present disclosure provides a target detection method, apparatus, electronic device and storage medium.
[0005] In a first aspect, embodiments of this disclosure provide a target detection method, the method comprising:
[0006] Obtain the first image to be detected;
[0007] Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected;
[0008] A subprocess is created to determine a first detection result of a target object in the first image to be detected based on images in the first image set.
[0009] Based on the first detection result, the final detection result of the target object in the first image to be detected is determined.
[0010] In one possible implementation, the multi-scale transformation of the first image to be detected includes:
[0011] Obtain a set of scale information with a preset base number;
[0012] Based on the scale information in the scale information set, scale transformation is performed on the first image to be detected; and
[0013] The creation of the subprocess includes:
[0014] Create the preset number of subprocesses;
[0015] The images in the first image set are allocated to the subprocesses of the preset number of subprocesses.
[0016] In one possible implementation, acquiring the first image to be detected includes:
[0017] Acquire the target video;
[0018] Determine the size information of the target video;
[0019] If the size represented by the size information is greater than or equal to a preset threshold, the size of the target video is reduced by a first percentage to obtain a size-adjusted video;
[0020] If the size represented by the size information is less than the preset threshold, the size of the target video is reduced by a second percentage to obtain a size-adjusted video, wherein the second percentage is greater than the first percentage;
[0021] The image is extracted from the resized video to obtain the first image to be detected.
[0022] In one possible implementation, the scale information set is determined in the following manner:
[0023] Determine the size information of the first image to be detected;
[0024] If the size represented by the size information is greater than or equal to a preset threshold, the scale information set is determined to be a first preset scale information set;
[0025] If the size represented by the size information is less than the preset threshold, the scale information set is determined to be a second preset scale information set;
[0026] Wherein, the second target scale in the second scale sequence is greater than the first target scale in the first scale sequence; the second scale sequence is a sequence obtained by sorting the scales in the second preset scale information set according to the first preset order; the first scale sequence is a sequence obtained by sorting the scales in the first preset scale information set according to the first preset order; the index of the second target scale in the second scale sequence is the same as the index of the first target scale in the first scale sequence.
[0027] In one possible implementation, the first image to be detected is extracted from the target video; and
[0028] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0029] Determine whether the preset number of first detection results have been obtained;
[0030] If the preset number of first detection results are obtained, the final detection result of the target object in the first image to be detected is determined based on the preset number of first detection results.
[0031] In one possible implementation, after determining whether the preset number of first detection results have been obtained, the method further includes:
[0032] Extract the second image to be detected from the target video;
[0033] Based on the scale information in the scale information set, the second image to be detected is scaled and transformed to obtain the second image set.
[0034] The images in the second image set are allocated to the subprocesses in the preset number of subprocesses, so that the subprocesses determine the second detection result of the target object in the second image to be detected based on the images in the second image set;
[0035] Based on the preset number of second detection results, the final detection result of the target object in the second image to be detected is determined.
[0036] In one possible implementation, allocating images from the second image set to the subprocesses among the preset number of subprocesses includes:
[0037] Determine the target size set for the first image set allocated to the preset number of subprocesses;
[0038] For each of the preset number of subprocesses, based on the size of the image allocated to that subprocess from the target size set, the size of the image in the second image set to be allocated to that subprocess is determined;
[0039] Based on the dimensions of the images in the determined second image set, allocate images from the second image set to the subprocesses of the predetermined number of subprocesses.
[0040] In one possible implementation, determining the size of the image in the second image set allocated to each of the preset number of subprocesses, based on the size of the image allocated to that subprocess determined from the target size set, includes:
[0041] The dimensions in the target size set are sorted according to a second preset order to obtain a first size sequence;
[0042] Determine the subprocess sequence corresponding to the first size sequence, wherein the first size in the first size sequence corresponds one-to-one with the subprocess in the subprocess sequence, and the image of the first size has been assigned to the corresponding subprocess;
[0043] The inverted sequence of the first size sequence is determined as the second size sequence;
[0044] For each subprocess in the subprocess sequence, the second size corresponding to that subprocess in the second size sequence is determined as the size of the image in the second image set to be allocated to that subprocess; wherein the index of the second size corresponding to that subprocess in the second size sequence is the same as the index of that subprocess in the subprocess sequence.
[0045] In one possible implementation, acquiring the first image to be detected includes one of the following:
[0046] Retrieve the first image to be detected from memory;
[0047] Obtain the first image to be detected from the video memory.
[0048] In one possible implementation, the subprocess uses the P-Net model to determine a first detection result of the target object in the first image to be detected based on images in the first image set; and
[0049] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0050] Using the R-Net model, based on the first detection result, a third detection result of the target object in the first image to be detected is determined;
[0051] Using the O-Net model, based on the third detection result, the final detection result of the target object in the first image to be detected is determined.
[0052] In one possible implementation, the P-Net model, the R-Net model, and the O-Net model are models optimized by TensorRT.
[0053] Secondly, embodiments of this disclosure provide a target detection device, the device comprising:
[0054] The acquisition unit is used to acquire the first image to be detected;
[0055] The first transformation unit is used to perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected.
[0056] A creation unit is used to create a subprocess, so that the subprocess determines a first detection result of a target object in the first image to be detected based on the images in the first image set;
[0057] The first determining unit is configured to determine the final detection result of the target object in the first image to be detected based on the first detection result.
[0058] In one possible implementation, the multi-scale transformation of the first image to be detected includes:
[0059] Obtain a set of scale information with a preset base number;
[0060] Based on the scale information in the scale information set, scale transformation is performed on the first image to be detected; and
[0061] The creation of the subprocess includes:
[0062] Create the preset number of subprocesses;
[0063] The images in the first image set are allocated to the subprocesses of the preset number of subprocesses.
[0064] In one possible implementation, acquiring the first image to be detected includes:
[0065] Acquire the target video;
[0066] Determine the size information of the target video;
[0067] If the size represented by the size information is greater than or equal to a preset threshold, the size of the target video is reduced by a first percentage to obtain a size-adjusted video;
[0068] If the size represented by the size information is less than the preset threshold, the size of the target video is reduced by a second percentage to obtain a size-adjusted video, wherein the second percentage is greater than the first percentage;
[0069] The image is extracted from the resized video to obtain the first image to be detected.
[0070] In one possible implementation, the scale information set is determined in the following manner:
[0071] Determine the size information of the first image to be detected;
[0072] If the size represented by the size information is greater than or equal to a preset threshold, the scale information set is determined to be a first preset scale information set;
[0073] If the size represented by the size information is less than the preset threshold, the scale information set is determined to be a second preset scale information set;
[0074] Wherein, the second target scale in the second scale sequence is greater than the first target scale in the first scale sequence; the second scale sequence is a sequence obtained by sorting the scales in the second preset scale information set according to the first preset order; the first scale sequence is a sequence obtained by sorting the scales in the first preset scale information set according to the first preset order; the index of the second target scale in the second scale sequence is the same as the index of the first target scale in the first scale sequence.
[0075] In one possible implementation, the first image to be detected is extracted from the target video; and
[0076] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0077] Determine whether the preset number of first detection results have been obtained;
[0078] If the preset number of first detection results are obtained, the final detection result of the target object in the first image to be detected is determined based on the preset number of first detection results.
[0079] In one possible implementation, after determining whether the preset number of first detection results have been obtained, the device further includes:
[0080] An extraction unit is used to extract a second image to be detected from the target video;
[0081] The second transformation unit is used to perform scale transformation on the second image to be detected based on the scale information in the scale information set, so as to obtain the second image set.
[0082] An allocation unit is configured to allocate images from the second image set to the subprocesses among the preset number of subprocesses, so that the subprocesses determine a second detection result of the target object in the second image to be detected based on the images in the second image set;
[0083] The second determining unit is used to determine the final detection result of the target object in the second image to be detected based on the preset number of second detection results.
[0084] In one possible implementation, allocating images from the second image set to the subprocesses among the preset number of subprocesses includes:
[0085] Determine the target size set for the first image set allocated to the preset number of subprocesses;
[0086] For each of the preset number of subprocesses, based on the size of the image allocated to that subprocess from the target size set, the size of the image in the second image set to be allocated to that subprocess is determined;
[0087] Based on the dimensions of the images in the determined second image set, allocate images from the second image set to the subprocesses of the predetermined number of subprocesses.
[0088] In one possible implementation, determining the size of the image in the second image set allocated to each of the preset number of subprocesses, based on the size of the image allocated to that subprocess determined from the target size set, includes:
[0089] The dimensions in the target size set are sorted according to a second preset order to obtain a first size sequence;
[0090] Determine the subprocess sequence corresponding to the first size sequence, wherein the first size in the first size sequence corresponds one-to-one with the subprocess in the subprocess sequence, and the image of the first size has been assigned to the corresponding subprocess;
[0091] The inverted sequence of the first size sequence is determined as the second size sequence;
[0092] For each subprocess in the subprocess sequence, the second size corresponding to that subprocess in the second size sequence is determined as the size of the image in the second image set to be allocated to that subprocess; wherein the index of the second size corresponding to that subprocess in the second size sequence is the same as the index of that subprocess in the subprocess sequence.
[0093] In one possible implementation, acquiring the first image to be detected includes one of the following:
[0094] Retrieve the first image to be detected from memory;
[0095] Obtain the first image to be detected from the video memory.
[0096] In one possible implementation, the subprocess uses the P-Net model to determine a first detection result of the target object in the first image to be detected based on images in the first image set; and
[0097] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0098] Using the R-Net model, based on the first detection result, a third detection result of the target object in the first image to be detected is determined;
[0099] Using the O-Net model, based on the third detection result, the final detection result of the target object in the first image to be detected is determined.
[0100] In one possible implementation, the P-Net model, the R-Net model, and the O-Net model are models optimized by TensorRT.
[0101] Thirdly, embodiments of this disclosure provide an electronic device, including:
[0102] Memory, used to store computer programs;
[0103] A processor is configured to execute a computer program stored in the aforementioned memory, and when the computer program is executed, to implement the method of any embodiment of the target detection method of the first aspect of this disclosure.
[0104] Fourthly, embodiments of this disclosure provide a computer-readable storage medium that, when executed by a processor, implements the method of any embodiment of the target detection method of the first aspect described above.
[0105] Fifthly, embodiments of this disclosure provide a computer program including computer-readable code that, when executed on a device, causes a processor in the device to execute instructions for implementing the steps of the method as described in any embodiment of the target detection method of the first aspect above.
[0106] The target detection method provided in this disclosure can acquire a first image to be detected, then perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected, then create a subprocess to determine a first detection result of a target object in the first image to be detected based on the images in the first image set, and finally determine the final detection result of the target object in the first image to be detected based on the first detection result. This method can determine the first detection result of a target object in the first image to be detected by using a subprocess based on the images in the first image set obtained by multi-scale transformation, thereby obtaining the final detection result of the target object in the first image to be detected, thus improving the efficiency of target detection. Attached Figure Description
[0107] Figure 1 A schematic flowchart of a target detection method provided in an embodiment of this disclosure;
[0108] Figure 2 A flowchart illustrating another target detection method provided in this embodiment of the present disclosure;
[0109] Figure 3 To execute the target using a main process and child processes Figure 2 A flowchart illustrating the target detection method;
[0110] Figure 4 A schematic flowchart of another target detection method provided in this disclosure embodiment;
[0111] Figure 5 To execute the target using a main process and child processes Figure 4 A flowchart illustrating the target detection method;
[0112] Figure 6 In response to Figure 4 A schematic diagram of an application scenario;
[0113] Figure 7A This is a schematic diagram of the processing procedure for the P-Net model;
[0114] Figure 7B This is a schematic diagram of the processing procedure of the R-Net model;
[0115] Figure 7C This is a schematic diagram of the processing procedure for the O-Net model;
[0116] Figures 8A-8C A schematic flowchart of another target detection method provided in an embodiment of this disclosure;
[0117] Figure 8D A schematic diagram illustrating a method for calculating the intersection-union ratio (IU) according to an embodiment of this disclosure;
[0118] Figure 9 This is a schematic diagram of the structure of a target detection device provided in an embodiment of the present disclosure;
[0119] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0120] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0121] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of this disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they indicate the logical order between them.
[0122] It should also be understood that in this embodiment, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0123] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0124] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0125] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0126] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0127] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0128] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0129] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. To facilitate understanding of the embodiments of this disclosure, the disclosure will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0130] Figure 1This is a flowchart illustrating a target detection method provided in an embodiment of this disclosure. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0131] In some cases, this method can be executed via the main process on the aforementioned electronic device.
[0132] like Figure 1 As shown, the method specifically includes:
[0133] 101. Obtain the first image to be detected.
[0134] In this embodiment, the first image to be detected can be an image on which object detection (e.g., face detection) is performed. This first image to be detected can be any frame from a video (e.g., the target video described later), or it can be an image obtained by direct capture. The first image to be detected may or may not include a target object (e.g., a face).
[0135] The aforementioned target object can be an image obtained by photographing a target (such as a human face).
[0136] 102. Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected.
[0137] In this embodiment, multiple different scales can be used to scale the first image to be detected, thereby obtaining multiple images of different sizes.
[0138] The first set of images mentioned above can be multiple images of different sizes.
[0139] 103. Create a subprocess so that the subprocess determines a first detection result of the target object in the first image to be detected based on the images in the first image set.
[0140] In this embodiment, the subprocess can determine the first detection result of the target object in the first image to be detected based on all or part of the images in the first image set.
[0141] The aforementioned first detection result can be the detection result of a target object in the first image to be detected. The first detection result can characterize whether the first image to be detected includes a target object. If the first image to be detected includes a target object, the first detection result may further include the location information of the target object in the first image to be detected, and the confidence level of the location information of the target object in the first image to be detected.
[0142] Here, the subprocess can use face detection algorithms such as LFFD (A Light and Fast Face Detector for EdgeDevices) to determine the first detection result of the target object in the first image to be detected based on the images in the first image set.
[0143] 104. Based on the first detection result, determine the final detection result of the target object in the first image to be detected.
[0144] In this embodiment, the aforementioned final detection result may be the final detection result obtained by performing face detection on the target object in the first image to be detected. The final detection result can characterize whether the first image to be detected includes a target object. If the first image to be detected includes a target object, the final detection result may further include the location information of the target object in the first image to be detected, and the confidence level of the location information of the target object in the first image to be detected.
[0145] As an example, the confidence level of the first detection result that is greater than or equal to a preset threshold, and the location information corresponding to the confidence level that is greater than or equal to the preset threshold, can be determined as the final detection result.
[0146] In some cases, steps 101 and 104 can be executed via the main process, step 102 can be executed via a child process, and the step of determining the first detection result of the target object in the first image to be detected based on the images in the first image set can be executed.
[0147] In some optional implementations of this embodiment, step 101 can be performed in the following manner to obtain the first image to be detected:
[0148] Retrieve the first image to be detected from memory.
[0149] It is understandable that, among the above optional implementation methods, obtaining the first image to be detected from memory can improve the speed of obtaining the first image to be detected, thereby improving the speed of target detection.
[0150] In some optional implementations of this embodiment, step 101 can also be performed in the following manner to obtain the first image to be detected:
[0151] Obtain the first image to be detected from the video memory.
[0152] It is understandable that, in the above optional implementation methods, by obtaining the first image to be detected from the video memory, the copying time of the first image to be detected from the video memory to the main memory and then back to the video memory can be reduced, thereby improving the acquisition speed of the first image to be detected and thus improving the target detection speed.
[0153] The target detection method provided in this disclosure can acquire a first image to be detected, then perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected, then create a subprocess to determine a first detection result of a target object in the first image to be detected based on the images in the first image set, and finally determine the final detection result of the target object in the first image to be detected based on the first detection result. This method can determine the first detection result of a target object in the first image to be detected by using a created subprocess based on the images in the first image set obtained by multi-scale transformation, thereby obtaining the final detection result of the target object in the first image to be detected, thus improving the efficiency of target detection.
[0154] Figure 2 This is a schematic flowchart illustrating another target detection method provided in an embodiment of this disclosure. Figure 2 As shown, the method specifically includes:
[0155] 201. Obtain the first image to be detected.
[0156] In this embodiment, step 201 and Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.
[0157] 202. Obtain a set of scale information with a preset base number.
[0158] In this embodiment, the cardinality of the scale information set, that is, the number of scale information items contained in the scale information set, can be a preset number. For example, the preset number could be 8, 10, etc. For example, the scale information set could be "0.3000, 0.2127, 0.1508, 0.1069, 0.0758, 0.0538, 0.0381, 0.0270". Each scale information item in the scale information set can represent the scaling ratio of the image.
[0159] 203. Based on the scale information in the scale information set, scale transformation is performed on the first image to be detected to obtain the first image set of the first image to be detected.
[0160] In this embodiment, the scale represented by each scale information in the scale information set is used to perform a scale transformation on the first image to be detected, thereby obtaining an image. Therefore, the cardinality of the first image set (i.e., the number of images contained in the first image set) can be the aforementioned preset number.
[0161] 204. Create the preset number of subprocesses.
[0162] 205. Assign images from the first image set to the subprocesses in the preset number of subprocesses, so that the subprocesses determine the first detection result of the target object in the first image to be detected based on the images in the first image set.
[0163] In this embodiment, images from the first image set can be randomly assigned to the subprocesses among the preset number of subprocesses. Alternatively, images from the first image set can be assigned to the subprocesses among the preset number of subprocesses according to a preset allocation strategy.
[0164] Here, a single subprocess can determine a first detection result of the target object in the first image to be detected based on a single image in the first image set. Thus, a preset number of first detection results can be obtained.
[0165] 206. Based on the first detection result, determine the final detection result of the target object in the first image to be detected.
[0166] In this embodiment, step 206 and Figure 1 Step 104 in the corresponding embodiment is basically the same, and will not be repeated here.
[0167] In some optional implementations of this embodiment, step 101 can be performed in the following manner to obtain the first image to be detected:
[0168] First, obtain the target video.
[0169] Next, the size information of the target video is determined.
[0170] The aforementioned size information can characterize the length, width, and other dimensions of an image in the target video. In some cases, this size information can characterize the length of the shorter side of the image's length and width in the target video.
[0171] Then, if the size represented by the size information is greater than or equal to a preset threshold (e.g., 500 pixels), the size of the target video is reduced by a first percentage (e.g., 30%) to obtain a size-adjusted video.
[0172] Subsequently, if the size indicated by the size information is less than the preset threshold, the size of the target video is reduced by a second percentage (e.g., 15%) to obtain a resized video. The second percentage is greater than the first percentage.
[0173] Finally, images are extracted from the resized video to obtain the first image to be detected.
[0174] It is understandable that among the above optional implementation methods, different sizes can be adjusted for videos of different sizes, which can improve the speed and accuracy of target detection.
[0175] In some optional implementations of this embodiment, the scale information set can be determined in the following manner:
[0176] First, the size information of the first image to be detected is determined. This size information can represent the length, width, and other dimensions of the first image to be detected. In some cases, this size information can represent the length of the shorter side of the length and width of the first image to be detected.
[0177] Subsequently, if the size represented by the size information is greater than or equal to a preset threshold, the scale information set is determined to be the first preset scale information set.
[0178] As an example, the preset threshold mentioned above could be 500 pixels.
[0179] Then, if the size represented by the size information is less than the preset threshold, the scale information set is determined to be a second preset scale information set.
[0180] In this sequence, the second target scale in the second scale sequence is larger than the first target scale in the first scale sequence. The second scale sequence is a sequence obtained by sorting the scales in the second preset scale information set according to a first preset order (e.g., ascending or descending). The first scale sequence is a sequence obtained by sorting the scales in the first preset scale information set according to the first preset order. The index of the second target scale in the second scale sequence is the same as the index of the first target scale in the first scale sequence. An index can represent the position of an element in a sequence. For example, the index of the second target scale in the second scale sequence can represent the position of the second target scale in the second scale sequence. The index of the first target scale in the first scale sequence can represent the position of the first target scale in the first scale sequence.
[0181] In some cases, the second target scale can be a preset multiple of the first target scale. For example, the second target scale can be twice the first target scale.
[0182] As an example, if the first preset order is descending, the second scale sequence can be "0.3000, 0.2127, 0.1508, 0.1069, 0.0758, 0.0538, 0.0381, 0.0270", and the first scale sequence can be "0.1500, 0.1064, 0.0754, 0.0534, 0.0379, 0.0269, 0.0191, 0.1351".
[0183] Therefore, if the second target scale is "0.3000", the first target scale can be "0.1500".
[0184] It is understood that in the above optional implementation, the scale information set of the first image to be detected is determined by the size information of the first image to be detected. In this way, after performing multi-scale transformation on the larger image to be detected, a first image set with a large size difference from the image to be detected can be obtained. After performing multi-scale transformation on the smaller image to be detected, a first image set with a small size difference from the image to be detected can be obtained, thereby improving the uniformity obtained after multi-scale transformation.
[0185] In some optional implementations of this embodiment, the first image to be detected is extracted from the target video. Based on this, step 206 can be performed as follows to determine the final detection result of the target object in the first image to be detected based on the first detection result:
[0186] The first step is to determine whether the preset number of first detection results have been obtained.
[0187] The second step is to determine the final detection result of the target object in the first image to be detected based on the preset number of first detection results after obtaining the preset number of first detection results.
[0188] As an example, various methods can be used to determine the final detection result of the target object in the first image to be detected based on the preset number of first detection results. Please refer to the corresponding descriptions recorded in this disclosure for details, which will not be elaborated here.
[0189] It is understood that in the above optional implementation, the final detection result of the target object in the first image to be detected is determined only when the preset number of first detection results are obtained. In this way, if the number of first detection results obtained is less than the preset number, the main process can perform other operations (such as performing other steps included in this method), thereby further improving the speed of target detection.
[0190] In some application scenarios of the above optional implementation methods, after performing the first step, the following steps can also be performed:
[0191] Step 1: Extract the second image to be detected from the target video.
[0192] The second image to be detected can be any frame extracted from the target video. For example, the second image to be detected can be an image extracted from the target video again after the first image to be detected has been extracted, according to a preset frame extraction strategy.
[0193] Step 2: Based on the scale information in the scale information set, scale transformation is performed on the second image to be detected to obtain the second image set.
[0194] The single image in the second image set can be an image obtained by scaling the second image to be detected using a single scale from the scale information set.
[0195] Here, by using the scale represented by each scale information in the scale information set, a scale transformation is performed on the second image to be detected to obtain an image. Therefore, the cardinality of the second image set (i.e., the number of images contained in the second image set) can be the aforementioned preset number.
[0196] Step 3: Assign images from the second image set to the subprocesses in the preset number of subprocesses, so that the subprocesses can determine the second detection result of the target object in the second image to be detected based on the images in the second image set.
[0197] Here, images from the second image set can be randomly assigned to the subprocesses among the preset number of subprocesses. Alternatively, images from the second image set can be assigned to the subprocesses among the preset number of subprocesses according to a preset allocation strategy.
[0198] Here, a single subprocess can determine a second detection result of the target object in the second image to be detected based on a single image in the second image set. Thus, a preset number of second detection results can be obtained.
[0199] As an example, for each of the preset number of subprocesses, the size of the image in the second image set to be allocated to that subprocess can be determined based on the duration of its most recent image processing. For instance, if the duration of the most recent image processing for each of the preset number of subprocesses is relatively long, a smaller image can be allocated to that subprocess from the second image set.
[0200] Step four: Based on the preset number of second detection results, determine the final detection result of the target object in the second image to be detected.
[0201] The final detection result of the target object in the second image to be detected can be the final detection result obtained by performing face detection on the target object in the second image to be detected. The final detection result can characterize whether the second image to be detected includes the target object. If the second image to be detected includes the target object, the final detection result may further include the location information of the target object in the second image to be detected, and the confidence level of the location information of the target object in the second image to be detected.
[0202] As an example, the confidence level of the second detection result that is greater than or equal to a preset threshold, and the location information corresponding to the confidence level that is greater than or equal to the preset threshold, can be determined as the final detection result.
[0203] It is understandable that the above application scenarios can achieve target detection of target objects in target videos.
[0204] In some of the application scenarios described above, step three can be executed in the following manner (including sub-steps one through three):
[0205] Sub-step one: Determine the target size set of the first image set allocated to the preset number of sub-processes.
[0206] Here, the dimensions in the target size set can be the dimensions of the images in the first image set allocated to each subprocess. The dimensions of the images allocated to each subprocess can be determined after or before the images in the first image set allocated to each subprocess.
[0207] Sub-step two: For each of the preset number of sub-processes, based on the size of the image allocated to that sub-process from the target size set, determine the size of the image in the second image set to be allocated to that sub-process.
[0208] As an example, for each of the preset number of subprocesses, if the size of the image in the first image set that the subprocess recently processed is large, a smaller image can be allocated to the subprocess from the second image set.
[0209] Sub-step three: According to the size of the images in the determined second image set, allocate images from the second image set to the subprocesses of the preset number of subprocesses.
[0210] It is understood that, in the above case, the size of the image in the second image set to be allocated to the subprocess can be determined based on the size of the image allocated to the subprocess from the target size set. This makes the total number of images processed by each subprocess more similar, reducing the waiting time of the main process.
[0211] In some of the examples above, sub-step two can be performed as follows:
[0212] First, the dimensions in the target size set are sorted according to a second preset order to obtain a first size sequence.
[0213] The second preset order can be the same as or different from the first preset order. For example, the second preset order can be ascending or descending.
[0214] The first size sequence can be a size sequence obtained by sorting the sizes in the target size set according to a second preset order.
[0215] Next, the subprocess sequence corresponding to the first size sequence is determined. Specifically, there is a one-to-one correspondence between the first size in the first size sequence and the subprocesses in the subprocess sequence. The correspondence between the first size and the subprocess is as follows: an image of the first size has been assigned to the subprocess corresponding to that first size.
[0216] Then, the inverted sequence of the first size sequence is determined as the second size sequence.
[0217] Specifically, if the first size sequence is arranged in ascending order, then the second size sequence can be arranged in descending order. If the first size sequence is arranged in descending order, then the second size sequence can be arranged in ascending order.
[0218] Subsequently, for each subprocess in the subprocess sequence, the second size corresponding to that subprocess in the second size sequence is determined as the size of the image in the second image set to be allocated to that subprocess.
[0219] The correspondence between child processes and the second size is as follows: the index of the second size corresponding to the child process in the second size sequence is the same as the index of the child process in the child process sequence. An index represents the position of an element in a sequence. For example, the index of the child process in the child process sequence represents the position of the child process in the child process sequence. Similarly, the index of the second size in the second size sequence represents the position of the second size in the second size sequence.
[0220] It is understandable that in the example above, if the image size in the first image set most recently processed by the child process is large, a smaller image size can be allocated to the child process from the second image set, thereby further reducing the waiting time of the main process.
[0221] As an example, please refer to Figure 3 , Figure 3 To execute the target using a main process and child processes Figure 2 A flowchart illustrating the target detection method.
[0222] exist Figure 3In the process, after reading the image (i.e., the first image to be detected mentioned above), the main process can determine the eight pyramid proportions of the image (corresponding to the scale information set mentioned above). Message queue 1 can include the eight pyramid proportions. The order of the pyramid proportions in message queue 1 can be randomly determined, or arranged in ascending or descending order of size. Then, the main process delivers each pyramid proportion in message queue 1 to a subprocess. After each subprocess obtains the pyramid proportion, it can resize the image according to that pyramid proportion to obtain a pyramid image (corresponding to the first image set mentioned above). Then, the pyramid image is input into the P-net model to obtain the first detection result. Then, NMS (Non-Maximum Suppression) is performed on the first detection result. Specifically, the maxima in the first detection result can be stored in an npy file, and the nonmaxima in the first detection result can be discarded. The first detection result stored in the npy file is used as message queue 2. Next, the child process delivers message queue 2 to the main process, allowing the main process to read the first detection result from the pny file. The first detection result from the pny file is then sequentially input into the R-net (also known as R-Net) model and the O-net (also known as O-Net) model, thus determining the output of the O-net model as the final detection result.
[0223] The functions and structures of the P-net, R-net, and O-net models mentioned above will be described later and will not be repeated here.
[0224] It should be noted that, in addition to the contents described above, this embodiment may also include... Figure 1 The technical features described in the corresponding embodiments, thereby achieving Figure 1 For details on the technical effectiveness of the target detection method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0225] The target detection method provided in this disclosure improves the efficiency of target detection by creating sub-processes with a cardinality equal to that of the scale information set, so that each sub-process performs target detection on a single image in the first image set.
[0226] Figure 4This is a flowchart illustrating another target detection method provided in this disclosure. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0227] Specifically, such as Figure 4 As shown, the method specifically includes:
[0228] 301. Obtain the first image to be detected.
[0229] In this embodiment, step 301 and Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.
[0230] 302. Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected.
[0231] In this embodiment, step 302 and Figure 1 Step 102 in the corresponding embodiment is basically the same, and will not be repeated here.
[0232] 303. Create a subprocess so that the subprocess uses the P-Net (Proposal Network) model to determine the first detection result of the target object in the first image to be detected based on the images in the first image set.
[0233] In this embodiment, the basic structure of the P-Net model is a fully convolutional network. For the first image set constructed in the previous step, preliminary feature extraction and bounding box labeling are performed using an FCN (Fully Convolutional Network), followed by bounding-box regression to adjust the windows, and NMS (Non-Maximum Suppression) to filter most of the windows.
[0234] The P-Net model is a region proposal network for target regions (e.g., face regions). After the feature input results are processed through three convolutional layers, a target classifier (e.g., face classifier) is used to determine whether the region is a target object (e.g., face object). At the same time, bounding box regression and a keypoint (e.g., facial keypoint) locator are used to make preliminary proposals for target regions. This part will eventually output many regions that may contain target objects, and these regions are input into the R-Net model described in step 304 for further processing.
[0235] The purpose of using the P-Net model is to quickly generate target candidate windows (e.g., face candidate boxes) using a relatively shallow and simple CNN (Convolutional Neural Network). The process is as follows: Figure 7A As shown.
[0236] 304. Using the R-Net (Refine Network) model, based on the first detection result, determine the third detection result of the target object in the first image to be detected.
[0237] In this embodiment, the R-Net model is basically constructed as a convolutional neural network, adding a fully connected layer compared to the first layer of the P-Net model. Therefore, the selection of input data is more stringent. After the image (i.e., the first image to be detected mentioned above) passes through the P-Net model, many prediction windows (i.e., the first detection result mentioned above) are left. Here, all prediction windows are fed into the R-Net model. This network filters out a large number of candidate boxes with poor performance (corresponding to the position information of the target object in the first image to be detected mentioned above). Finally, the selected candidate boxes are further optimized by Bounding-BoxRegression and NMS to obtain the third detection result.
[0238] Because the output of the P-Net model is only a potentially plausible target region, this network refines the input selection, discarding most of the erroneous inputs. It then uses bounding box regression and keypoint (e.g., facial keypoint) locators again to perform bounding box regression and keypoint localization for the target region. Finally, it outputs a more plausible target region for use by the O-Net model in step 305. Compared to the P-Net model's use of 1x1x32 features from fully convolutional outputs, the R-Net model uses a 128-bit fully connected layer after the last convolutional layer, preserving more image features and achieving better accuracy than the P-Net model.
[0239] The goal of the R-Net model is to use a more complex network structure than the P-Net model to further select and adjust the potential target region windows generated by the P-Net model, thereby achieving high-precision filtering and target region optimization. Its process is as follows: Figure 7B As shown.
[0240] 305. Using the O-Net (Output Network) model, based on the third detection result, determine the final detection result of the target object in the first image to be detected.
[0241] In this embodiment, the basic structure of the O-Net model is a relatively complex convolutional neural network, which has one more convolutional layer than the R-Net model. The difference between the O-Net model and the R-Net model is that this layer structure identifies the target object region through more supervision, and it regresses feature points (such as facial feature points) and uses NMS to further optimize the prediction results. In the scenario of face detection, the O-Net model can finally output five facial feature points.
[0242] The O-Net model is a more complex convolutional network with more input features. At the end of its structure is a larger 256-segment fully connected layer, preserving even more image features. It then performs object discrimination, bounding box regression, and object feature localization, ultimately outputting the coordinates of the top-left and bottom-right corners of the target region along with five feature points. O-Net boasts more input features and a more complex network structure, resulting in better performance. The output of this layer serves as the final network model output, yielding the final detection result.
[0243] In some optional implementations of this embodiment, the P-Net model, the R-Net model, and the O-Net model are models optimized by TensorRT.
[0244] TensorRT optimization processes may include at least one of the following: identical model structure, fusion of model parameters, reduction of model parameter accuracy, and reuse of the same memory address.
[0245] It is understandable that among the above optional implementation methods, TensorRT is used to optimize the P-Net, R-Net and O-Net models, thereby further improving the efficiency of object detection.
[0246] As an example, please refer to Figure 5 , Figure 5 To execute the target using a main process and child processes Figure 4 A flowchart illustrating the target detection method.
[0247] exist Figure 5 In the process, after reading the image (i.e., the first image to be detected mentioned above), the main process can determine the eight pyramid proportions of the image (corresponding to the scale information set mentioned above). Message queue 1 can include the eight pyramid proportions. The order of the pyramid proportions in message queue 1 can be randomly determined, or arranged in ascending or descending order of size. Then, the main process submits each pyramid proportion in message queue 1 to a child process. After each child process obtains the pyramid proportion, it can resize the image according to that pyramid proportion to obtain a pyramid image (corresponding to the first image set mentioned above). Then, the pyramid image is input into the P-net model to obtain the first detection result. Then, NMS (Non-Maximum Suppression) is performed on the first detection result. Specifically, the maxima in the first detection result can be stored in an npy file, and the nonmaxima in the first detection result can be discarded. The first detection result stored in the npy file is used as message queue 2. Then, the child process submits message queue 2 to the main process. The main process checks if message queue 2 is full, i.e., whether all created child processes have returned the first detection result. If message queue 2 is full, the main process checks if message queue 1 is empty. If message queue 1 is empty (meaning the main process has delivered all pyramid proportions in message queue 1 to the child processes), the main process retrieves the next image (corresponding to the second image to be detected mentioned above). If message queue 2 is not full, the main process accepts the delivery from message queue 2 to read the first detection result from the pny file. The first detection result from the pny file is then sequentially input into the R-net (i.e., R-Net) model and the O-net (i.e., O-Net) model, thus determining the output of the O-net model as the final detection result.
[0248] Please continue to refer to this. Figure 6 , Figure 6 In response to Figure 4 A diagram illustrating an application scenario. In Figure 6 In the target video, there are images 0-6. The R-net model or O-net model is currently processing image 2, the P-Net model is currently processing image 3, and message queue 1 stores the 8 pyramid proportions of image 4.
[0249] The following describes the embodiments of this disclosure by way of example. However, it should be noted that the embodiments of this disclosure may have the features described below, but the following description does not constitute a limitation on the scope of protection of the embodiments of this disclosure.
[0250] Please refer to Figures 8A-8C , Figures 8A-8C This is a schematic flowchart of another target detection method provided in an embodiment of the present disclosure.
[0251] Here, eight pyramid scales can be generated (corresponding to the scale information set mentioned above). When the image (corresponding to the first-generation detection image and the second image to be detected) is input, the image size can be adjusted. Specifically, for target videos with a shortest side less than 500 pixels, the adjustment is performed by multiplying the side length by 0.3; for those greater than 500 pixels, the adjustment is performed by multiplying the side length by 0.15. After adjustment, when constructing the image pyramid, further size adjustments can be performed from the smaller image size, improving algorithm performance.
[0252] In addition, the results of video frame extraction (corresponding to the first-generation detection image and the second image to be detected mentioned above) are stored in memory, and then the target image (corresponding to the first-generation detection image and the second image to be detected mentioned above) is read directly from memory, which can reduce the time required to read the image.
[0253] For P-Net, R-Net, and O-Net models, optimization using TRT (such as TensorRT and TF-TRT) can further improve their performance.
[0254] Since the image pyramid constructs images with 8 different sizes, and the P-Net model's image processing time increases with the image size, when the previous image in a single process is the largest size, the next image will be scheduled to be the smallest size image, and so on. This way, the processing time can be averaged across the processes, avoiding the performance loss caused by continuously processing the largest size image in a single process.
[0255] In addition, because NMS processing is performed, not all the first detection results corresponding to the 8 scale images will be input into the R-net model.
[0256] exist Figures 8A-8C middle, Figure 8A The number 801 in the middle and Figure 8B The label 801 indicates the same connection point. Figure 8A Number 802 and Figure 8C The label 802 indicates the same connection point. Figure 8A Number 803 and Figure 8C The label 803 indicates the same connection point. Figure 8B The number 804 in the middle and Figure 8B The label 804 indicates the same connection point. For example, because... Figure 8A The number 801 in the middle and Figure 8BThe label 801 indicates the same connection point, so the main process can post message queue 1 to the subroutine.
[0257] The parameters of message queue 1 include: image matrix, pyramid ratio (used for image resizing), ratio marker (1-X), and image sequence number marker 0 or 1. For example, the Xth image is marked as the remainder of X divided by 2. X is used to identify the image and the storage path of the npy file. The parameters of message queue 2 include the ratio index number.
[0258] Specifically, face detection will be used as an example for illustration.
[0259] First, the images can be transformed at different scales to construct an image pyramid (e.g., 8 images) to accommodate the detection of faces of different sizes.
[0260] like Figures 8A-8CAs shown, after obtaining the dimensions of the video (corresponding to the target video mentioned above), the main process determines whether the shortest side of the video is greater than 500 pixels. If it is, the video size is scaled down to 15%; if it is less than or equal to 500 pixels, the video size is scaled down to 30% to pre-adjust the video size. Then, the main program uses the GPU (Graphics Processing Unit) to perform interval frame extraction on the pre-adjusted video to obtain a single image (corresponding to the first image to be detected mentioned above). After reading the image (i.e., the first image to be detected mentioned above), the main process can determine the eight pyramid proportions of the image (corresponding to the scale information set mentioned above). Then, the main process sends each pyramid proportion from message queue 1 to a subprocess. After each subprocess obtains the pyramid proportion, it can resize the image according to that proportion to obtain a pyramid image (corresponding to the first image set mentioned above). Then, the pyramid image is input into the P-Net model to obtain the first detection result. Then, NMS (Non-Maximum Suppression) is applied to the first detection result. Specifically, the maximum values in the first detection result can be stored in an npy file, and the non-maximum values in the first detection result can be discarded. The first detection result stored in the npy file is used as message queue 2. Then, the child process delivers message queue 2 to the main process. The main process checks whether message queue 2 has been delivered. If it has, it accepts the delivery of message queue 2; if it has not, it checks whether message queue 1 is empty. If message queue 1 is empty (that is, the main process delivers each pyramid ratio in message queue 1 to the child process), the main process obtains the next image (corresponding to the second image to be detected mentioned above). If message queue 1 is not empty, the main process checks whether message queue 2 has been delivered. The first detection results corresponding to the two images are maintained through list 0 and list 1 respectively. When list 0 or list 1 is full, the full list is cleared, and the first detection result in the pny file is read. The first detection result in the pny file is then sequentially input into the R-net (also known as R-Net) model and the O-net (also known as O-Net) model, so that the output result of the O-net model is determined as the final detection result.
[0261] The purpose of using NMS is to remove duplicate bounding boxes in the prediction results. Specifically, when an image is input, it is continuously downsampled using pyramids, and faces are detected on the sampled image each time. During detection, a P-Net network is usually used to slide across the image with a stride of 2. Because the stride is too small, a face may be bounded multiple times. Suppose we initially bound five faces on an image with confidence scores of 0.98, 0.83, 0.75, 0.81, and 0.67. The first three confidence scores indicate one face, and the last two indicate another face.
[0262] Here, the five bounding boxes are sorted according to their confidence scores. The box with the highest confidence score (0.98) is then compared with the remaining boxes using Intersection over Union (IoU). Boxes with IoU less than a threshold (set to 0.3 in the code) are retained. This leaves boxes 0.81 and 0.67. The process is repeated, taking the box with the highest confidence score (0.81) and comparing it with the remaining boxes using IoU, retaining boxes with IoU less than the threshold. Finally, only the two face bounding boxes, 0.98 and 0.81, remain.
[0263] The aforementioned IOU is used to calculate the degree of overlap between two images. In the P-net and R-net models, when calculating IOU, boxes with IOU values greater than a threshold are considered duplicates and discarded, retaining boxes with lower IOU values. However, if, after processing by the P-net and R-net models, a situation like the image below (where large boxes enclose smaller boxes) occurs in the final O-net, then even boxes with lower IOU values will be retained, which is undesirable. Therefore, we use a different approach to calculate IOU in the O-net network. Figure 8D The second method of IOU is used to improve the false detection rate.
[0264] thus, Figures 8A-8C The method shown improves object detection efficiency by combining multi-processing, pipelined tasks, image preprocessing, adaptive process scheduling, TRT optimization, and memory image reading. Furthermore, in some cases, images can be read directly from video memory to reduce the copying time from video memory to RAM and back to video memory, thereby further improving object detection efficiency.
[0265] It should be noted that, in addition to the contents described above, this embodiment may also include the technical features described in the above embodiments, thereby achieving the technical effects of the target detection method described above. For details, please refer to the relevant descriptions above. For the sake of brevity, these will not be elaborated here.
[0266] The target detection method provided in this disclosure improves the accuracy of target detection by employing P-Net, R-Net, and O-Ne models.
[0267] Figure 9 This is a schematic diagram of the structure of a target detection device provided in an embodiment of this disclosure. Specifically, it includes:
[0268] Acquisition unit 901 is used to acquire the first image to be detected;
[0269] The first transformation unit 902 is used to perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected.
[0270] A creation unit 903 is used to create a subprocess, so that the subprocess determines a first detection result of the target object in the first image to be detected based on the images in the first image set;
[0271] The first determining unit 904 is used to determine the final detection result of the target object in the first image to be detected based on the first detection result.
[0272] In one possible implementation, the multi-scale transformation of the first image to be detected includes:
[0273] Obtain a set of scale information with a preset base number;
[0274] Based on the scale information in the scale information set, scale transformation is performed on the first image to be detected; and
[0275] The creation of the subprocess includes:
[0276] Create the preset number of subprocesses;
[0277] The images in the first image set are allocated to the subprocesses of the preset number of subprocesses.
[0278] In one possible implementation, acquiring the first image to be detected includes:
[0279] Acquire the target video;
[0280] Determine the size information of the target video;
[0281] If the size represented by the size information is greater than or equal to a preset threshold, the size of the target video is reduced by a first percentage to obtain a size-adjusted video;
[0282] If the size represented by the size information is less than the preset threshold, the size of the target video is reduced by a second percentage to obtain a size-adjusted video, wherein the second percentage is greater than the first percentage;
[0283] The image is extracted from the resized video to obtain the first image to be detected.
[0284] In one possible implementation, the scale information set is determined in the following manner:
[0285] Determine the size information of the first image to be detected;
[0286] If the size represented by the size information is greater than or equal to a preset threshold, the scale information set is determined to be a first preset scale information set;
[0287] If the size represented by the size information is less than the preset threshold, the scale information set is determined to be a second preset scale information set;
[0288] Wherein, the second target scale in the second scale sequence is greater than the first target scale in the first scale sequence; the second scale sequence is a sequence obtained by sorting the scales in the second preset scale information set according to the first preset order; the first scale sequence is a sequence obtained by sorting the scales in the first preset scale information set according to the first preset order; the index of the second target scale in the second scale sequence is the same as the index of the first target scale in the first scale sequence.
[0289] In one possible implementation, the first image to be detected is extracted from the target video; and
[0290] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0291] Determine whether the preset number of first detection results have been obtained;
[0292] If the preset number of first detection results are obtained, the final detection result of the target object in the first image to be detected is determined based on the preset number of first detection results.
[0293] In one possible implementation, after determining whether the preset number of first detection results have been obtained, the device further includes:
[0294] An extraction unit (not shown in the figure) is used to extract a second image to be detected from the target video;
[0295] The second transformation unit (not shown in the figure) is used to perform scale transformation on the second image to be detected based on the scale information in the scale information set to obtain the second image set.
[0296] An allocation unit (not shown in the figure) is used to allocate images from the second image set to the subprocesses among the preset number of subprocesses, so that the subprocesses determine the second detection result of the target object in the second image to be detected based on the images in the second image set;
[0297] The second determining unit (not shown in the figure) is used to determine the final detection result of the target object in the second image to be detected based on the preset number of second detection results.
[0298] In one possible implementation, allocating images from the second image set to the subprocesses among the preset number of subprocesses includes:
[0299] Determine the target size set for the first image set allocated to the preset number of subprocesses;
[0300] For each of the preset number of subprocesses, based on the size of the image allocated to that subprocess from the target size set, the size of the image in the second image set to be allocated to that subprocess is determined;
[0301] Based on the dimensions of the images in the determined second image set, allocate images from the second image set to the subprocesses of the predetermined number of subprocesses.
[0302] In one possible implementation, determining the size of the image in the second image set allocated to each of the preset number of subprocesses, based on the size of the image allocated to that subprocess determined from the target size set, includes:
[0303] The dimensions in the target size set are sorted according to a second preset order to obtain a first size sequence;
[0304] Determine the subprocess sequence corresponding to the first size sequence, wherein the first size in the first size sequence corresponds one-to-one with the subprocess in the subprocess sequence, and the image of the first size has been assigned to the corresponding subprocess;
[0305] The inverted sequence of the first size sequence is determined as the second size sequence;
[0306] For each subprocess in the subprocess sequence, the second size corresponding to that subprocess in the second size sequence is determined as the size of the image in the second image set to be allocated to that subprocess; wherein the index of the second size corresponding to that subprocess in the second size sequence is the same as the index of that subprocess in the subprocess sequence.
[0307] In one possible implementation, acquiring the first image to be detected includes one of the following:
[0308] Retrieve the first image to be detected from memory;
[0309] Obtain the first image to be detected from the video memory.
[0310] In one possible implementation, the subprocess uses the P-Net model to determine a first detection result of the target object in the first image to be detected based on images in the first image set; and
[0311] The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes:
[0312] Using the R-Net model, based on the first detection result, a third detection result of the target object in the first image to be detected is determined;
[0313] Using the O-Net model, based on the third detection result, the final detection result of the target object in the first image to be detected is determined.
[0314] In one possible implementation, the P-Net model, the R-Net model, and the O-Net model are models optimized by TensorRT.
[0315] The target detection device provided in this embodiment can be as follows: Figure 9 The target detection device shown can execute all the steps of the target detection methods described above, thereby achieving the technical effects of the target detection methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.
[0316] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Figure 10 The illustrated electronic device 500 includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the electronic device 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to implement communication between these components. In addition to a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 10 The general designated all buses as Bus System 505.
[0317] The user interface 503 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).
[0318] It is understood that the memory 502 in this embodiment of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0319] In some implementations, memory 502 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 5021 and application program 5022.
[0320] The operating system 5021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 5022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 5022.
[0321] In this embodiment, by calling the program or instructions stored in memory 502, specifically the program or instructions stored in application program 5022, processor 501 executes the method steps provided in each method embodiment, including, for example:
[0322] Obtain the first image to be detected;
[0323] Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected;
[0324] A subprocess is created to determine a first detection result of a target object in the first image to be detected based on images in the first image set.
[0325] Based on the first detection result, the final detection result of the target object in the first image to be detected is determined.
[0326] The methods disclosed in the above embodiments of this disclosure can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in processor 501. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 502. Processor 501 reads the information in memory 502 and, in conjunction with its hardware, completes the steps of the above method.
[0327] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described above, or combinations thereof.
[0328] For software implementation, the techniques described herein can be implemented by units that perform the functions described above. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or external to the processor.
[0329] The electronic device provided in this embodiment may be as follows: Figure 10 The electronic device shown can perform all the steps of the target detection methods described above, thereby achieving the technical effects of the target detection methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.
[0330] This disclosure also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; the memory may also include combinations of the above types of memory.
[0331] When one or more programs in the storage medium can be executed by one or more processors to implement the target detection method described above that is executed on the electronic device side.
[0332] The processor described above is used to execute a target detection program stored in memory to implement the following steps of a target detection method executed on the electronic device side:
[0333] Obtain the first image to be detected;
[0334] Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected;
[0335] A subprocess is created to determine a first detection result of a target object in the first image to be detected based on images in the first image set.
[0336] Based on the first detection result, the final detection result of the target object in the first image to be detected is determined.
[0337] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0338] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0339] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this disclosure. It should be understood that the above descriptions are merely specific embodiments of this disclosure and are not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure. Furthermore, for the foregoing embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
Claims
1. A target detection method, characterized in that, The method includes: Acquire a first image to be detected, which is extracted from the target video; Perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected; A subprocess is created to determine a first detection result of a target object in the first image to be detected based on images in the first image set. Based on the first detection result, the final detection result of the target object in the first image to be detected is determined; The subprocess uses the P-Net model to determine the first detection result of the target object in the first image to be detected based on the images in the first image set; and The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes: using the R-Net model to determine the third detection result of the target object in the first image to be detected based on the first detection result; and using the O-Net model to determine the final detection result of the target object in the first image to be detected based on the third detection result. Extract the second image to be detected from the target video; Based on the scale information in the scale information set, the second image to be detected is scaled and transformed to obtain the second image set. The images in the second image set are allocated to a preset number of subprocesses, so that the subprocesses determine the second detection result of the target object in the second image to be detected based on the images in the second image set; Based on the preset number of second detection results, the final detection result of the target object in the second image to be detected is determined; The step of allocating images from the second image set to subprocesses in a preset number of subprocesses includes: determining a target size set for the first image set to be allocated to the preset number of subprocesses; for each subprocess in the preset number of subprocesses, determining the size of the images in the second image set to be allocated to that subprocess based on the size of the images allocated to that subprocess determined from the target size set; and allocating images from the second image set to the subprocesses in the preset number of subprocesses according to the determined size of the images in the second image set.
2. The method according to claim 1, characterized in that, The multi-scale transformation of the first image to be detected includes: Obtain a set of scale information with a preset base number; Based on the scale information in the scale information set, scale transformation is performed on the first image to be detected; and The creation of the subprocess includes: Create the preset number of subprocesses; The images in the first image set are allocated to the subprocesses of the preset number of subprocesses.
3. The method according to claim 2, characterized in that, The acquisition of the first image to be detected includes: Acquire the target video; Determine the size information of the target video; If the size represented by the size information is greater than or equal to a preset threshold, the size of the target video is reduced by a first percentage to obtain a size-adjusted video; If the size represented by the size information is less than the preset threshold, the size of the target video is reduced by a second percentage to obtain a size-adjusted video, wherein the second percentage is greater than the first percentage; The image is extracted from the resized video to obtain the first image to be detected.
4. The method according to claim 2, characterized in that, The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes: Determine whether the preset number of first detection results have been obtained; If the preset number of first detection results are obtained, the final detection result of the target object in the first image to be detected is determined based on the preset number of first detection results.
5. The method according to claim 1, characterized in that, For each of the preset number of subprocesses, determining the size of the image in the second image set to be allocated to that subprocess based on the size of the image allocated to that subprocess from the target size set includes: The dimensions in the target size set are sorted according to a second preset order to obtain a first size sequence; Determine the subprocess sequence corresponding to the first size sequence, wherein the first size in the first size sequence corresponds one-to-one with the subprocess in the subprocess sequence, and the image of the first size has been assigned to the corresponding subprocess; The inverted sequence of the first size sequence is determined as the second size sequence; For each subprocess in the subprocess sequence, the second size corresponding to that subprocess in the second size sequence is determined as the size of the image in the second image set to be allocated to that subprocess; wherein the index of the second size corresponding to that subprocess in the second size sequence is the same as the index of that subprocess in the subprocess sequence.
6. The method according to any one of claims 1-5, characterized in that, The acquisition of the first image to be detected includes one of the following: Retrieve the first image to be detected from memory; Obtain the first image to be detected from the video memory.
7. The method according to claim 1, characterized in that, The P-Net model, the R-Net model, and the O-Net model are models optimized by TensorRT.
8. A target detection device, characterized in that, The device includes: The acquisition unit is used to acquire a first image to be detected, which is extracted from the target video; The first transformation unit is used to perform multi-scale transformation on the first image to be detected to obtain a first image set of the first image to be detected. A creation unit is used to create a subprocess, so that the subprocess determines a first detection result of a target object in the first image to be detected based on the images in the first image set; The first determining unit is configured to determine the final detection result of the target object in the first image to be detected based on the first detection result; The subprocess uses the P-Net model to determine the first detection result of the target object in the first image to be detected based on the images in the first image set; and The step of determining the final detection result of the target object in the first image to be detected based on the first detection result includes: using the R-Net model to determine the third detection result of the target object in the first image to be detected based on the first detection result; and using the O-Net model to determine the final detection result of the target object in the first image to be detected based on the third detection result. An extraction unit is used to extract a second image to be detected from the target video; The second transformation unit is used to perform scale transformation on the second image to be detected based on the scale information in the scale information set, so as to obtain the second image set. An allocation unit is used to allocate images from the second image set to a preset number of subprocesses, so that the subprocesses can determine a second detection result of the target object in the second image to be detected based on the images in the second image set; The second determining unit is used to determine the final detection result of the target object in the second image to be detected based on the preset number of second detection results; The step of allocating images from the second image set to subprocesses in a preset number of subprocesses includes: determining a target size set for the first image set to be allocated to the preset number of subprocesses; for each subprocess in the preset number of subprocesses, determining the size of the images in the second image set to be allocated to that subprocess based on the size of the images allocated to that subprocess determined from the target size set; and allocating images from the second image set to the subprocesses in the preset number of subprocesses according to the determined size of the images in the second image set.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.
Citation Information
Patent Citations
Method and apparatus for smart television's video memory image recognition
CN106470358A
Image detection method and device, electronic equipment, and computer readable medium
CN108520229A
Face occlusion detection method and device, computer device and storage medium
CN110826519A
Image detection method, device and equipment and storage medium
CN111382638A
Target detection method and device, electronic equipment and storage medium
CN112507983A