Detection Method for Forward View Images of Intelligent Driving
By performing block detection and fusion processing on ultra-high-definition camera images, the accuracy and efficiency problems when detecting small targets in the prior art are solved, and more efficient detection results are achieved.
Patent Information
- Application Number
- CN202211004434.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-08-22
AI Technical Summary
In the prior art, images collected by ultra-high-definition cameras cannot take into account both detection accuracy and detection efficiency when detecting small objects, and conventional methods lead to increased calculation times or loss of information.
An intelligent driving forward-view image detection method is adopted. By acquiring multiple image blocks of different sizes, inputting the convolutional neural network model for detection, and combining the fusion strategy to improve detection accuracy and reduce the number of detections.
It improves the detection ability and accuracy of small distant objects in the image, while reducing the number of detections and improving the detection efficiency.
Smart Images

Figure CN115331200B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly to a method for detecting a forward view image of intelligent driving. Background Art
[0002] The development of autonomous driving technology is in full swing, and advanced driver assistance systems (ADAS) have gradually become standard equipment for automobiles. People's requirements for the forward view perception technology are also getting higher and higher, hoping that the forward view camera can "see" farther and more clearly. Therefore, ultra-high definition cameras are gradually applied to the vehicle perception system. After the high-definition images collected by the ultra-high definition camera are detected, features such as people and objects in the images can be recognized, providing information about obstacles in front of the vehicle for autonomous driving and improving the safety of autonomous driving.
[0003] Currently, convolutional neural networks are widely used in the detection of high-definition images. Compared with traditional vision algorithms, convolutional neural networks require higher computing power of the chip. Generally speaking, the larger the network model of the convolutional neural network and the larger the resolution of the input picture, the better the detection result for identifying small targets. However, the computing power of in-vehicle chips in the current market is generally not high. The conventional approach is to design a lightweight network model, and then fix the size of the picture input into the network model at about 600,000 pixels. There are mainly two methods: cutting the picture and not cutting the picture.
[0004] (1) Not cutting the picture: Suppose an ultra-high definition camera captures a picture of 8 million (3840×2160) pixels. Then, the 8 million pixel picture needs to be first reduced to about 600,000 pixels and then sent into the convolutional neural network model for calculation. This results in the inability to effectively utilize the local information of small targets in the ultra-high definition picture, thus affecting the detection effect of small targets.
[0005] (2) Cutting the picture: First, the original picture is successively cut into multiple small pictures of the same size, and the size of the small pictures is the same as the input size of the network model. Then, each small picture is sent into the convolutional neural network model for detection, and the ultra-high definition original picture is reduced to the input size of the network model and sent into the network for detection to obtain the final detection result. Taking the above 3840×2160 picture as an example, the original picture needs to be cut into 16 small pictures of about 600,000 pixels, and finally 16 + 1 = 17 detections are required. This picture cutting detection method, although the detection ability of distant small targets is improved, the number of calculations is significantly increased. Even if the size of the network model is reduced, the total detection time will be much higher than that of the non-picture-cutting detection method, and the detection efficiency is low. Summary of the Invention
[0006] The technical problem to be solved by the present invention is: to solve the technical problem that the existing detection methods for ultra-high-definition pictures cannot balance the detection accuracy and detection efficiency. The present invention provides a detection method for the front-view image of intelligent driving, which can not only improve the detection accuracy of small targets, but also reduce the number of detections, thereby improving the detection efficiency.
[0007] The technical solution adopted by the present invention to solve its technical problems is: a detection method for the front-view image of intelligent driving, including the following steps: S1. Obtain the front-view image collected by the front-view camera; S2. Obtain multiple first pictures Xi and multiple second pictures Yj from the front-view image, and the size of the second picture Yj is larger than that of the first picture Xi; S2. Process the second picture Yj to obtain a third picture Yj'; S4. Input the first picture Xi and the third picture Yj' into a convolutional neural network model to output multiple first detection results; S5. Map the first detection results to the front-view image to obtain second detection results; S6. Perform a fusion process on the multiple second detection results to obtain the final detection result.
[0008] Further, the first detection result is an object detection result or a semantic segmentation detection result.
[0009] Further, the size of the first picture Xi is w×h, and the size of the second picture Yj is 2w×2h.
[0010] Further, there is an overlap of 0.2w between adjacent first pictures Xi, an overlap of 0.2h between the first picture Xi and the second picture Yj, and an overlap of 0.6w between adjacent second pictures Yj.
[0011] Further, reduce the size of the front-view image to w×h and input it into the convolutional neural network model to output the first detection result; the first detection result is an object detection result.
[0012] Further, when the first detection result is an object detection result, the fusion process in step S6 specifically includes:
[0013] Traverse all the second detection results to determine whether there is an edge rectangle;
[0014] If not, perform non-maximum suppression processing on the second detection results to obtain the final detection result;
[0015] If so, calculate the iog value between the edge rectangle and other rectangles; if the iog value is greater than the threshold roi_threshold, determine that the edge rectangle is a redundant box and delete the redundant box; until all redundant boxes are deleted, perform non-maximum suppression processing on the second detection results to obtain the final detection result.
[0016] Further, when the first detection result is a semantic segmentation detection result, the fusion process in step S6 specifically includes:
[0017] When the second detection results of two adjacent first pictures Xi are fused, two class probability values will be obtained for the overlapping area, and the maximum value of the two class probability values is taken as the final class probability value of the overlapping area;
[0018] When the second detection result of the first picture Xi and the second detection result of the second picture Yj are fused, more than two class probability values will be obtained for the overlapping area, and the maximum value of the class probabilities of the first picture Xi is taken as the final class probability value of the overlapping area.
[0019] Further, the first picture Xi is the upper area of the front view image; the second picture Yj is the lower area of the front view image.
[0020] Further, the size of the third picture Yj' is w×h.
[0021] Further, assuming the number of the first pictures Xi is n and the number of the second pictures Yj is m, the number of detections of the detection method is n + m + 1 or n + m.
[0022] The beneficial effect of the present invention is that the detection method of the intelligent driving front view image of the present invention can not only improve the detection ability and detection accuracy of small targets in the distance in the image, but also reduce the number of detections and improve the detection efficiency by improving the cutting method and combining the fusion strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present invention will be further described below with reference to the drawings and embodiments.
[0024] Figure 1 is a flowchart of the detection method of the intelligent driving front view image of the present invention.
[0025] Figure 2 is a schematic diagram of cutting the front view image of the present invention.
[0026] Figure 3 is a schematic diagram of non-overlapping cutting in the prior art.
[0027] Figure 4 is a schematic diagram of the detection process of the object detection of the present invention.
[0028] Figure 5 is a schematic diagram of the coordinate transformation of the present invention.
[0029] Figure 6 is a flowchart of the fusion of the object detection results of the present invention.
[0030] Figure 7 It is a schematic diagram of the redundant frame of the present invention.
[0031] Figure 8 It is a schematic diagram of the detection process of semantic segmentation detection of the present invention. Detailed implementation manners
[0032] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0033] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation to the present invention. In addition, the features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.
[0034] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "mounted", "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0035] As Figure 1 shown, the detection method of the forward view image of intelligent driving of the present invention includes the following steps:
[0036] S1. Obtain the forward view image collected by the forward view camera.
[0037] S2. Obtain multiple first pictures Xi and multiple second pictures Yj from the forward view image, and the size of the second picture Yj is larger than that of the first picture Xi.
[0038] S2. Process the second picture Yj to obtain a third picture Yj'.
[0039] S4. Input the first image Xi and the third image Yj' into the convolutional neural network model to output multiple first detection results.
[0040] S5. Map the first detection results to the front view image to obtain the second detection results.
[0041] S6. Perform a fusion process on multiple second detection results to obtain the final detection results.
[0042] It should be noted that the front view image is a super high-definition image, and the size of the front view image is, for example, 34000×28000. Multiple first images Xi and multiple second images Yj can be segmented from the front view image. In other words, the first image Xi and the second image Yj are local images of the front view image. Among them, the size of the second image Yj is larger than that of the first image Xi. The first image Xi is the upper region of the front view image, and the second image Yj is the lower region of the front view image. Because in the front view perception scenario, small targets in the distance are located in the upper half of the image, and large targets in the vicinity are located in the lower half of the image. Since the pixel information of small targets in the distance is less than that of large targets in the vicinity, if the images of small targets in the distance are further reduced, it will lead to further loss of pixel information; while for the images of large targets in the vicinity, even if the length and width are reduced by half respectively, the original features and categories of the targets can still be judged. Therefore, when cutting the images, the upper region of the front view image is cut with the first image Xi (i.e., the small image), and the lower region of the front view image X is cut with the second image Yj (i.e., the medium image). For example, the size of the first image Xi is w×h, and the size of the second image Yj is 2w×2h, that is, the size of the second image Yj is four times that of the first image Xi. Before detection, the size of the second image Yj also needs to be reduced to the same as that of the first image Xi, that is, the size of the third image Yj' is also w×h. Then the first image Xi and the third image Yj are respectively input into the convolutional neural network model for detection. In this way, the detection ability of small targets in the distance can be improved while reducing the detection operation amount and improving the detection efficiency. It should be noted that the convolutional neural network used in the present invention has been pre-trained.
[0043] Specifically, there is an overlap of 0.2w between two adjacent first images Xi, an overlap of 0.2h between the first image Xi and the second image Yj, and an overlap of 0.6w between two adjacent second images Yj. In other words, there is a 20% overlapping area between small images, a 20% overlapping area between small images and medium images, and a 30% overlapping area between medium images. For example, Figure 2For example, the front view image is divided into four small images and two medium-sized images, denoted as X1, X2, X3, X4, Y1, and Y2 respectively. There is a 20% area overlap between small image X1 and small image X2, a 20% area overlap between small image X2 and small image X3, a 20% area overlap between small image X3 and small image X4, and small images X1, X2, X3, and X4 are all located in the upper half of the front view image. There is a 30% area overlap between medium-sized image Y1 and medium-sized image Y2, and the upper part of medium-sized image Y1 overlaps with the lower parts of small images X1, X2, and X3, and the upper part of medium-sized image Y2 overlaps with the lower parts of small images X2, X3, and X4. When cutting the images in the present invention, the overlapping areas are set to improve the detection rate and detection accuracy of the targets at the edges of the images. For example, please refer to Figure 3 , assuming that the target is exactly located at the tangent line. At this time, half of the pixel information of the target is in small image X1 and the other half is in small image X2. If the overlapping area is not set, then the convolutional neural network model will detect this target in both small images X1 and X2, but it is impossible to determine whether these two detection results are for the same target or two targets, resulting in the final detection result not conforming to the actual situation.
[0044] It should be noted that the first detection result of the present invention is a target detection result or a semantic segmentation detection result. In other words, the present invention can perform two types of detections on the front view image, namely, target detection and semantic segmentation detection. Target detection means bounding the target with a rectangular box in the image, and semantic segmentation detection means assigning the same color to the pixel points belonging to the same category in the image. The convolutional neural network model adopted by the present invention can simultaneously have the functions of target detection and semantic segmentation detection. The difference between target detection and semantic segmentation detection is that for the target detection task, the front view image also needs to be input into the convolutional neural network model to output the first detection result. The specific processes of the two detections are introduced separately below.
[0045] (1) Target detection
[0046] As Figure 4As shown, in the object detection task, first, the acquired front view image is segmented into n first pictures Xi and m second pictures Yj, where i = 1, 2,..., n; j = 1, 2,..., m. The sizes of the m second pictures Yj are reduced to be the same as those of the first pictures Xi, obtaining m third pictures Yj'. The size of the front view image (i.e., the original image) is reduced to be the same as that of the first pictures Xi, obtaining a reduced image. The n first pictures Xi, the m second pictures Yj, and a reduced image of the front view image are respectively input into the convolutional neural network model, obtaining n + m + 1 object detection results. The object detection results of the n first pictures Xi and the m second pictures Yj are subjected to coordinate transformation and mapped onto the front view image, obtaining n + m second detection results. Finally, the n + m second detection results (i.e., the object detection results) and the first detection result of the front view image are subjected to fusion processing to obtain the final object detection result. In the object detection task, the reduced image of the front view image is also input into the convolutional neural network and finally undergoes fusion processing together, in order to enhance the global prediction ability of the convolutional neural network model and at the same time avoid the situation where large objects are not accurately detected.
[0047] Please refer to Figure 5 , the object detection results (i.e., the first detection results) of the n first pictures Xi and the m second pictures Yj are subjected to coordinate transformation and mapped onto the front view image. The coordinate transformation specifically includes: taking the upper left corner of the picture (including the first picture Xi and the second picture Yj) as the origin, with the horizontal right direction as the positive x-axis direction and the vertical downward direction as the positive y-axis direction to establish a two-dimensional coordinate system, so that the coordinates (x, y) of each pixel point in the picture can be obtained. The first detection result is a rectangular box, which is also composed of multiple pixel points. Thus, the coordinates of all pixel points of the rectangular box can be obtained. Mapping the rectangular boxes of the first pictures Xi and the second pictures Yj onto the front view image is a mapping from point coordinates to point coordinates. Each first picture Xi and each second picture Yj can establish its own two-dimensional coordinate system. Since the positions of each picture in the front view image are different, the mapping relationships of the point coordinates are also different.
[0048] For example, the n first pictures Xi are X1, X2, X3,..., Xn from left to right in sequence. Let the coordinates of a certain pixel point U in the first picture Xi be (x ai , y ai ). The mapping formula (1) for mapping the U point onto the front view image is:
[0049] x' ai = 0.8w×(i - 1) + x ai
[0050] y' ai = y ai
[0051] Among them, a represents the number of pixel points of the small image, i represents the serial number of the small image, and (x' ai , y' ai ) represents the position of the U point in the front view image after mapping. Through this mapping formula (1), the rectangular frames of all small images can be mapped to the original image.
[0052] For example, m second pictures Yj are Y1, ..., Ym from left to right in sequence. Let the coordinates of a certain pixel point R in the second picture Yj be (x bj , y bj ). The mapping formula (2) for mapping the R point to the front view image is:
[0053] x' bj = 1.4w×(j - 1) + x bj
[0054] y' bj = 0.8h×(j - 1) + y bj
[0055] Among them, b represents the number of pixel points of the middle image, j represents the serial number of the middle image, and (x' bj , y' bj ) represents the position of the R point in the front view image after mapping. Through this mapping formula (2), the rectangular frames of all middle images can be mapped to the original image.
[0056] According to the mapping formulas (1) and (2), it can be known that the mapping relationships between the small image, the middle image and the original image are different. Due to the particularity of the cut images, the positions of the small image and the middle image in the original image are different. In order to accurately reflect the rectangular frames in the original image, it is necessary to perform coordinate mapping on the small image and the middle image respectively.
[0057] Please refer to Figure 6 . In the object detection task, the specific process of fusing the second detection results (i.e., the mapped rectangular frames) of the small image and the middle image with the first detection result of the original image includes: traversing all the second detection results to determine whether there are edge rectangular frames. If not, perform non-maximum suppression processing on the second detection results to obtain the final detection result. If so, calculate the iog value between the edge rectangular frame and other rectangular frames; if the iog value is greater than the threshold roi_threshold, it is considered that the edge rectangular frame is a redundant frame, and delete this redundant frame; until all the redundant frames are deleted, perform non-maximum suppression processing on the second detection results to obtain the final detection result.
[0058] Since there are overlapping regions when cutting the images in the present invention, when detecting the small images and medium images, it is equivalent to detecting the overlapping regions twice. Moreover, when cutting the images, the cutting line may pass through the target, resulting in only a part of the target being detected in some images, thus leading to inaccurate recognition of the target. Therefore, when performing the fusion process, it is necessary to screen the second detection results to exclude redundant bounding boxes. For example Figure 7 As shown Figure 1 the small Figure 2 and the small Figure 1 have 20% overlap. The small Figure 2 can only detect the left half (ABEF) of the target, while the small Figure 1 detects the whole target (ACDF). At this time, the detection result of the small
[0059] for the target (ABEF) is redundant. When performing the fusion process, first traverse all the second detection results to determine whether there are edge bounding boxes. An edge bounding box refers to a bounding box located near the cutting line of the image. For example, the distance d between the second detection result and the cutting line of the image where it is located can be calculated, and a distance judgment threshold side_threshold is set. If the distance d is less than the distance judgment threshold side_threshold (the distance judgment threshold is 5, unit: pixel for example), then mark this second detection result as an edge bounding box; if the distance d is greater than or equal to the distance judgment threshold side_threshold, then mark this second detection result as a normal bounding box. After identifying the edge bounding box, it is necessary to determine whether this edge bounding box is a redundant box. Calculate the iog value between this edge bounding box and other bounding boxes. The calculation formula of the iog value is iog(G,P)=(G∩P) / G, where G represents the area of the edge bounding box, P represents the area of other bounding boxes, and G∩P represents the intersection of the two areas. If the iog value is greater than the threshold roi_threshold (the threshold roi_threshold is 0.8 for example), then it is considered that this edge bounding box is a redundant box. Taking the figure as an example, G = the area of ABEF, P = the area of ACDF, and G∩P is equal to G, then the iog value is equal to 1 (greater than 0.8). Therefore, this edge bounding box is a redundant box and needs to be deleted. Until all the redundant boxes in all the second detection results are deleted, perform non-maximum suppression (NMS) processing on the remaining second detection results to obtain the final detection result. Since there may be overlapping situations in the remaining second detection results (that is, there are multiple bounding boxes at the same target position), after non-maximum suppression processing, only the bounding box with the highest confidence is retained among the overlapping bounding boxes, and other non-overlapping bounding boxes are directly retained to obtain the final detection result, which can further improve the accuracy of target detection.
[0060] (2) Semantic segmentation detection
[0061] As Figure 8 shown, in the semantic segmentation detection task, first, the acquired front view image is segmented into n first pictures Xi and m second pictures Yj, where i = 1, 2,..., n; j = 1, 2,..., m. The sizes of the m second pictures Yj are reduced to be the same as those of the first pictures Xi, obtaining m third pictures Yj'. The n first pictures Xi and the m second pictures Yj are respectively input into the convolutional neural network model to obtain n + m semantic segmentation detection results. Then, the n + m second detection results (i.e., the semantic segmentation detection results) are fused to obtain the final object detection result.
[0062] In the semantic segmentation detection task, the obtained detection results show the features of the same category in the same color. The semantic segmentation detection results of the n first pictures Xi and the m second pictures Yj are subjected to coordinate transformation and mapped to the front view image, that is, coordinate transformation is performed on the pixel points. The specific coordinate mapping formula can be processed using mapping formulas (1) and (2), which will not be elaborated here.
[0063] Since semantic segmentation detection is a judgment of feature categories, the fusion process specifically includes two cases: (1) When fusing the second detection results of two adjacent first pictures Xi, two category probability values will be obtained in the overlapping area, and the maximum value of the two category probability values is taken as the final category probability value of the overlapping area. (2) When fusing the second detection result of the first picture Xi and the second detection result of the second picture Yj, more than two category probability values will be obtained in the overlapping area, and the maximum value of the category probability of the first picture Xi is taken as the final category probability value of the overlapping area. In other words, since there are overlapping areas between the pictures when cutting the pictures, during the fusion process, it is necessary to judge and select the categories detected in the overlapping areas. For example, two category probabilities are detected in the overlapping area between small pictures. For example, the small picture X1 detects "person", and the small picture X2 detects "car". Then, when performing the fusion process, select the one with the larger probability value between "person" and "car" as the final category of this feature. For example, in the overlapping area between a small picture and a medium picture, since this overlapping area is the overlap of multiple small pictures and the medium picture, more than two (including two) category probability values can be obtained during detection. Because the small pictures input into the convolutional neural network model are not scaled, while the medium pictures are scaled down, during the fusion process, among the multiple category probability values in the same overlapping area, select the maximum value of the category probability value of the small picture as the final category of this feature, which can improve the detection accuracy. For example, the small picture X1 detects "person", the small picture X2 detects "car", the small picture X3 detects "roadblock", and the medium picture Y1 detects "car". Assuming that the probability value obtained from the small picture X1 is the largest, then the final detected category is determined to be "person".
[0064] For example, the size of the front view image is 3840×2160 (consistent with the background art). Using this method, four small images and two medium-sized images can be cut out. The number of pixels in the small images is about 600,000 (meeting the calculation requirements of the convolutional neural network model). When performing the object detection task, a total of 4 + 2 + 1 detections are carried out. When performing the semantic segmentation task, a total of 4 + 2 detections are carried out. The number of detections is much less than that of the prior art. Using the detection method of the present invention can not only improve the detection rate and detection accuracy of small targets in the distance in the image, but also reduce the number of detections, shorten the detection time, and improve the detection effect.
[0065] Enlightened by the above ideal embodiments according to the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A detection method for the front view image of intelligent driving, characterized in that, It includes the following steps: S1. Obtain the front view image collected by the front view camera; S2. Obtain multiple first pictures Xi and multiple second pictures Yj from the front view image, where the size of the second picture Yj is larger than that of the first picture Xi; there is an overlap of 0.2w between two adjacent first pictures Xi, an overlap of 0.2h between the first picture Xi and the second picture Yj, and an overlap of 0.6w between two adjacent second pictures Yj; w represents the width and h represents the height; S3. Process the second picture Yj to obtain a third picture Yj'; S4. Input the first picture Xi and the third picture Yj' into the convolutional neural network model to output multiple first detection results; S5. Map the first detection results to the front view image to obtain second detection results; S6. Perform a fusion process on multiple second detection results to obtain a final detection result; When the first detection result is an object detection result, the fusion process in step S6 specifically includes: Traverse all the second detection results to determine whether there is an edge rectangle; If not, perform non-maximum suppression processing on the second detection results to obtain the final detection result; If so, calculate the iog value between the edge rectangle and other rectangles; if the iog value is greater than the threshold roi_threshold, determine that the edge rectangle is a redundant rectangle and delete the redundant rectangle; until all redundant rectangles are deleted, perform non-maximum suppression processing on the second detection results to obtain the final detection result; When the first detection result is a semantic segmentation detection result, the fusion process in step S6 specifically includes: When fusing the second detection results of two adjacent first pictures Xi, two class probability values will be obtained in the overlapping area, and the maximum value of the two class probability values is taken as the final class probability value of the overlapping area; When fusing the second detection result of the first picture Xi and the second detection result of the second picture Yj, more than two class probability values will be obtained in the overlapping area, and the maximum value of the class probabilities of the first picture Xi is taken as the final class probability value of the overlapping area.
2. The detection method of the intelligent driving front view image according to claim 1, characterized in that, The size of the first picture Xi is w×h, and the size of the second picture Yj is 2w×2h.
3. The detection method of the forward view image for intelligent driving according to claim 1, characterized in that, Reduce the size of the front view image to w×h and input it into the convolutional neural network model to output the first detection result; the first detection result is an object detection result.
4. The detection method of the intelligent driving front view image according to claim 1, characterized in that, The first picture Xi is the upper area of the front view image; the second picture Yj is the lower area of the front view image.
5. The detection method of the intelligent driving front view image according to claim 2, wherein, The size of the third picture Yj' is w×h.
6. The detection method of the forward view image for intelligent driving according to claim 3, wherein Assume the number of the first pictures Xi is n and the number of the second pictures Yj is m, then the number of detections of the detection method is n + m + 1 or n + m.
Citation Information
Patent Citations
Population statistics system based on video surveillance image processing
CN109376637A
Improved target detection method, system and device
CN109741333A