Data processing method, device and storage medium

By identifying and mapping road areas into target images of a set size in intelligent transportation scenarios, the problem of low effective information density in video surveillance systems is solved, improving the accuracy of object detection and recognition results.

CN115049990BActive Publication Date: 2026-03-20ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-16
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In intelligent transportation scenarios, the effective information density of road images captured by video surveillance systems is low, which makes it difficult for computers to effectively identify objects in distant areas and affects the coverage of intelligent traffic perception systems.

Method used

By identifying road areas and mapping them to target images of a set size, a mapping algorithm is used to appropriately enlarge smaller objects in the distant area and shrink larger objects in the near area, thereby increasing the effective information density of the image. Deep learning algorithms are then used for object detection.

Benefits of technology

It improves the accuracy of object detection and recognition results, and enhances the coverage of the intelligent traffic perception system for distant areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049990B_ABST
    Figure CN115049990B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, device and storage medium. In the data processing method, after a road region in a to-be-processed image is recognized, the road region is mapped to a target image of a set size, so that redundant information in the to-be-processed image can be removed, and the effective information density of the target image obtained by mapping is improved. Object detection is performed based on the target image with higher effective information density, which is beneficial to improving the accuracy of the recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a data processing method, device and storage medium. Background Technology

[0002] In intelligent transportation scenarios, video surveillance systems can be deployed on roads to capture images, and specific algorithms can be used to identify objects within these images. Based on the identified objects, object tracking or anomaly detection can be performed.

[0003] Typically, road images captured by video surveillance systems have low effective information density, making it difficult to accurately identify target objects within them. Therefore, a new solution is needed. Summary of the Invention

[0004] This application provides a data processing method, apparatus, and storage medium to improve the effective information density of an image.

[0005] This application provides a data processing method, including: acquiring an image to be processed; identifying a road region in the image to be processed; mapping the road region to a target image of a set size according to a set mapping algorithm; and performing object detection on the target image to identify objects contained in the road region.

[0006] This application embodiment also provides a data processing method, including: acquiring a road region in a road image, wherein objects in the road region are labeled with a first bounding box; performing mapping processing on the road region and the first bounding box according to a set mapping algorithm to obtain a training sample of a set size and a second bounding box on the training sample; inputting the training sample into a neural network model, and using the second bounding box as a supervision signal to train the object detection capability of the neural network model.

[0007] This application also provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the data processing method provided in this application.

[0008] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the data processing method provided in this application.

[0009] In the data processing method provided by the embodiment of the present application, after the road region in the image to be processed is recognized, the road region is mapped to a target image with a set size, so that the redundant information in the image to be processed can be removed, and the effective information density of the target image obtained by mapping is improved. Object detection is performed based on the target image with higher effective information density, which is beneficial to improving the accuracy of the recognition result. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the principles of the present application, and do not limit the present application in any manner. In the drawings:

[0011] Figure 1 A flowchart of a data processing method provided by an exemplary embodiment of the present application is shown in FIG. 2;

[0012] Figure 2a A flowchart of a data processing method provided by another exemplary embodiment of the present application is shown in FIG. 3;

[0013] Figure 2b A schematic diagram of the geometric features of a road region provided by an exemplary embodiment of the present application is shown in FIG. 4;

[0014] Figure 2c A schematic diagram of the geometric features of a road region provided by another exemplary embodiment of the present application is shown in FIG. 5;

[0015] Figure 2d A schematic diagram of the image mapping effect provided by an exemplary embodiment of the present application is shown in FIG. 6;

[0016] Figure 2e A flowchart of a data processing method provided by yet another exemplary embodiment of the present application is shown in FIG. 7;

[0017] Figure 3 A flowchart of a data processing method provided by yet another exemplary embodiment of the present application is shown in FIG. 8;

[0018] Figure 4 A schematic diagram of an application scenario instance provided by an exemplary embodiment of the present application is shown in FIG. 9;

[0019] Figure 5 A structural schematic diagram of an electronic device provided by an exemplary embodiment of the present application is shown in FIG. 10. DETAILED DESCRIPTION

[0020] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in connection with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0021] In the field of intelligent traffic perception, in order to realize intelligent monitoring, a public traffic video monitoring system can be deployed on the road to shoot the road. Based on the shot road image, the computer can replace the manual operation to identify, track and find the abnormal event of the target object. Generally, when the computer identifies the target object, it can analyze the video image collected by the monitoring device based on the method of deep learning. In the field of traffic perception, the computer can analyze the objects on the road based on the road video shot by the monitoring system, such as motor vehicles, non-motor vehicles, pedestrians, etc.

[0022] However, due to the scene live, the pixel ratio of the object on the long-range road in the road image shot by the monitoring system is low, which makes the effective information density of the road image low. In this case, the computer cannot effectively analyze the object on the long-range road, which makes the intelligent traffic perception system unable to effectively cover the long-range area. In view of the above technical problems, the present application provides a data processing method, which will be described below in connection with the drawings.

[0023] Figure 1 The flowchart of the data processing method provided by an exemplary embodiment of the present application is shown in FIG. 1, which comprises the following steps. Figure 1

[0024] Step 101, acquiring a to-be-processed image.

[0025] Step 102, identifying the road area in the to-be-processed image.

[0026] Step 103, mapping the road area to a target image with a set size according to a set mapping algorithm.

[0027] Step 104, detecting the object of the target image to identify the object contained in the road area.

[0028] ​The to-be-processed image contains a road region obtained by imaging a road, and the road region contains objects such as vehicles and pedestrians. The to-be-processed image can be obtained by a public transportation video monitoring system arranged on the road, or can be obtained by a vehicle-mounted event data recorder of a vehicle driving on the road, or can be obtained by a pedestrian on the road using an image collection device such as a mobile phone or a camera, and the present embodiment is not limited in this regard.

[0029] Object detection refers to perceiving and analyzing objects in a picture or a video stream by a computer and marking the objects. The object detection on the to-be-processed image can identify objects such as vehicles and pedestrians on the road and analyze the flow characteristics of the vehicles and pedestrians. When detecting objects in the to-be-processed image, the effective information required for detection is concentrated in the road region.

[0030] The road is in an extended state, and for an image collection device, the road has a far view and a near view. In the to-be-processed image obtained by the image collection device, the road region usually has a state that the near view region accounts for a large proportion and the far view region accounts for a small proportion. When the proportion of the far view region is small, the density of effective information is low, which is not conducive to object detection.

[0031] In the present embodiment, the road region can be identified from the to-be-processed image, and the road region is mapped into an image of a set size. For the convenience of description and distinction, the image obtained by mapping the to-be-processed image is described as a target image. The image information contained in the target image is composed of the image information contained in the road region. In the mapping process, the smaller objects in the far view region can be appropriately enlarged, and the larger objects in the near view region can be appropriately reduced, so that the objects in the target image obtained by mapping are approximately in the same order of magnitude of pixel size, which greatly improves the effective information density.

[0032] The mapping algorithm can be determined according to the size ratio of the road region, or can be determined according to the geometric characteristics of the road region, or can be determined based on the energy distribution characteristics of the road region, or can be determined based on the texture characteristics of the road region, and the present embodiment is not limited in this regard.

[0033] After the target image is obtained by mapping, object detection can be performed based on the target image. The target image has a high effective information density, which is conducive to improving the accuracy of object detection. It should be understood that the target image is mapped from the road region, and therefore, the objects on the target image and the objects contained in the road region have a corresponding relationship. After the objects on the target image are detected, the objects contained in the road region can be analyzed.

[0034] In the embodiment, after the road region in the image to be processed is recognized, the road region is mapped to a target image with a set size, so that the redundant information in the image to be processed can be removed, and the effective information density of the target image obtained by mapping is improved. Based on the target image with higher effective information density, object detection is performed, which is beneficial to improving the accuracy of the recognition result.

[0035] In the above and the following embodiments of the present application, when the road region is mapped to a target image with a set size, various mapping algorithms can be used, such as the mapping algorithm based on the size ratio of the road region, the mapping algorithm based on the geometric features of the road region, the mapping algorithm based on the energy distribution features of the road region, the mapping algorithm based on the texture features of the road region, and the like. The present embodiment is not limited.

[0036] Optionally, in the algorithm based on the size ratio of the road region, the size ratio of the road region can be calculated, and different local positions in the road region can be enlarged or reduced according to the size ratio, so as to obtain the target image.

[0037] Optionally, in the algorithm based on the geometric features of the road region, the road region can be mapped to a target image with the same size as the image to be processed according to the geometric features of the road region and the geometric features of the image to be processed.

[0038] Optionally, in the algorithm based on the energy distribution features of the road region, the energy distribution of the road region can be detected, the local region with higher energy distribution can be enlarged, and the local region with lower energy distribution can be reduced, so as to obtain the target image.

[0039] Optionally, in the algorithm based on the texture features of the road region, the texture distribution features in the road region can be recognized, the local region with concentrated texture distribution can be enlarged, and the local region with dispersed texture distribution can be reduced, so as to obtain the target image.

[0040] In different embodiments of the present application, the mapping processing of the road region can be implemented based on any one of the above-mentioned embodiments or a combination of the above-mentioned embodiments. The present embodiment is not limited. In the following embodiments, the optional implementation of the mapping processing based on the geometric features of the road region will be exemplarily described. Figure 2a 、 Figure 2b 、 Figure 2c 、 Figure 2d and Figure 2e .

[0041] Figure 2a The flowchart of the data processing method provided by another exemplary embodiment of the present application is shown in FIG. 6. Figure 2aAs shown, the method comprises:

[0042] Step 201, obtaining a to-be-processed image.

[0043] Step 202, identifying a road region in the to-be-processed image.

[0044] Step 203, determining a mapping relationship according to the to-be-processed image and geometric features of the road region.

[0045] Step 204, performing coordinate mapping on pixels in the road region according to the mapping relationship, to obtain a target image; the target image is adapted in size to the to-be-processed image.

[0046] Step 205, identifying an object in the target image and calculating a first bounding box of the object.

[0047] Step 206, performing reverse mapping processing on the first bounding box according to a reverse mapping relationship corresponding to the mapping relationship, to map the first bounding box to a second bounding box on the to-be-processed image.

[0048] Step 207, displaying the second bounding box on the to-be-processed image.

[0049] In step 201, a to-be-processed image is obtained. Optionally, the to-be-processed image can be obtained by sampling a road monitoring video.

[0050] In step 202, when identifying the road region in the to-be-processed image, the to-be-processed image can be subjected to edge detection to obtain edge information contained in the to-be-processed image; optionally, when the to-be-processed image is subjected to edge detection, a search-based edge detection method or a zero-crossing-based edge detection method can be used, and the present embodiment does not limit the same.

[0051] After obtaining the edge information contained in the to-be-processed image, a road contour can be determined from the to-be-processed image according to the edge information. In some cases, the road contour is a trapezoid or a polygon. To facilitate subsequent processing, optionally, an inscribed trapezoid of the road contour can be calculated to obtain an image region with a relatively regular shape and convenient for calculation, and the image region corresponding to the inscribed trapezoid is taken as the road region.

[0052] Figure 2b and Figure 2c Two optional embodiments of calculating the inscribed rectangle are illustrated. As shown in Figure 2b The road contour in the to-be-processed image is relatively regular, and the shape of the inscribed trapezoid is close to the road contour. As shown in Figure 2cAs shown, the road profile in the image to be processed is an irregular polygon, and the inscribed trapezoid is slightly smaller than the road profile. In practice, to obtain a more regular road profile, the road image can be avoided to be taken at the road corner, so as to improve the object detection effect.

[0053] In step 203, the mapping relationship refers to the mapping relationship between the coordinates of the pixel on the image to be processed and the coordinates of the pixel on the target image. Optionally, the mapping coefficient corresponding to the coordinates of the pixel can be calculated first, and then the mapping relationship is determined according to the mapping coefficient and the coordinates of the pixel. Optionally, the mapping coefficient includes a horizontal coordinate mapping coefficient, a vertical coordinate mapping coefficient, and a nonlinear mapping coefficient.

[0054] The nonlinear mapping coefficient is used to realize nonlinear mapping processing. In some scenarios, when the road in the captured road image extends along the vertical direction, the nonlinear mapping coefficient can be realized as a nonlinear mapping coefficient of the vertical coordinate. When the road extends along the vertical direction, the objects on the long-range road are small, and the objects on the short-range road are large. Based on the nonlinear mapping coefficient of the vertical coordinate, the proportion of the short-range objects in the road region can be reasonably compressed, and the proportion of the long-range objects in the road region can be expanded, so as to improve the effective information density of the target image.

[0055] The following will take any pixel in the road region as an example to illustrate the optional implementation of calculating the mapping coefficient and the mapping relationship. The geometric features of the image to be processed can include the length and height of the image to be processed. The geometric features of the road region can include at least one of the horizontal coordinate range of the row where each pixel in the road region is located in the road region and the vertical coordinate range of the road region. Generally, the vertical coordinate range of the road region is the range of the vertical coordinates spanned by the upper base and the lower base of the inscribed trapezoid corresponding to the road region.

[0056] Optionally, when calculating the horizontal coordinate mapping coefficient of the pixel, the horizontal coordinate range of the row where the pixel is located in the road region can be obtained, and then the horizontal coordinate mapping coefficient of the pixel is calculated according to the horizontal coordinate range and the length of the image to be processed. Optionally, the ratio of the length of the horizontal coordinate range to the length of the image to be processed can be calculated as the horizontal coordinate mapping coefficient. For details, reference can be made to the following formula:

[0057]

[0058] wherein S1 represents the horizontal coordinate mapping coefficient of the pixel, x max represents the maximum horizontal coordinate of the row where the pixel is located in the road region, x min represents the minimum horizontal coordinate of the row where the pixel is located in the road region, and w represents the length of the image to be processed, as Figure 2b and Figure 2c shown.

[0059] Accordingly, after the abscissa mapping coefficient is obtained, the abscissa mapping relationship of the pixel can be calculated according to the abscissa mapping coefficient and the abscissa of the pixel. Alternatively, the product of the abscissa of the pixel and the abscissa mapping coefficient can be calculated, and the relationship between the minimum abscissa of the row in which the pixel is located in the road region and the product is summed up as the abscissa mapping relationship of the pixel. Assuming that the coordinates of the pixel in the image to be processed are (x, y), the abscissa mapping relationship of the pixel can be recorded according to the following formula:

[0060]

[0061] Wherein, X represents the mapped abscissa.

[0062] Alternatively, when calculating the ordinate mapping coefficient of the pixel, the ordinate range of the road region can be obtained; and the ordinate mapping coefficient of the pixel can be calculated according to the ordinate range and the height of the image to be processed. Alternatively, the ratio of the height of the ordinate range to the height of the image to be processed can be calculated as the ordinate mapping coefficient. The specific calculation can be recorded according to the following formula:

[0063]

[0064] Wherein, S2 represents the ordinate mapping coefficient of the pixel, y max represents the maximum ordinate of the road region, y min represents the minimum ordinate of the road region, and h represents the height of the image to be processed, as shown in the following formula: Figure 2b and Figure 2c

[0065] Alternatively, when calculating the non-linear mapping coefficient of the pixel, in order to ensure the rationality and high availability of the mapping result, the road region can be further divided according to the extension trend of the road, and different non-linear mapping coefficients can be calculated for the pixels in different parts obtained by the division. The advantage of the division is that the road region is divided into a near view region and a far view region, based on the non-linear mapping coefficient of the near view region, the pixels in the near view region can be compressed and processed by non-linear mapping; based on the non-linear mapping coefficient of the far view region, the pixels in the far view region can be expanded and processed by non-linear mapping.

[0066] Alternatively, the road region can be divided into two, three or four parts according to the center line of the image to be processed, which is not limited in the embodiment. Alternatively, when divided into two parts, one part includes pixels with ordinate less than h / 2, and the other part includes pixels with ordinate greater than or equal to h / 2.

[0067] ​Next, the non-linear mapping coefficient of each pixel can be calculated based on the range of its ordinate and the mapping coefficient of its abscissa. One possible method for calculating the non-linear mapping coefficient is described in the following formula:

[0068]

[0069] Accordingly, when calculating the ordinate mapping relationship of a pixel, the product of the pixel's ordinate, the ordinate mapping coefficient, and the non-linear mapping coefficient can be calculated. Then, the summation relationship between the minimum ordinate value of the column containing the pixel and this product is determined, and this summation is used as the ordinate mapping relationship for that pixel. See the following formula for details:

[0070]

[0071] Where Y represents the mapped ordinate.

[0072] After obtaining the horizontal and vertical coordinate mapping relationships, step 204 can be executed to perform coordinate mapping on the pixels in the road area according to these mapping relationships, thereby obtaining the target image. The size of the target image is adapted to the image to be processed.

[0073] When mapping pixels in a road area, the mapped x-coordinate value of each pixel can be calculated based on the x-coordinate mapping relationship of each pixel and the x-coordinate value of each pixel in the image to be processed; similarly, the mapped y-coordinate value of each pixel can be calculated based on the y-coordinate mapping relationship of each pixel and the y-coordinate value of each pixel in the image to be processed. As shown in the following formula:

[0074] I`(x,y)=I[X(x,y),Y(x,y)] Formula 6

[0075] Where I`(x,y) represents the target image obtained by mapping, I represents the image to be processed, X represents the horizontal coordinate mapping value, Y represents the horizontal coordinate mapping value, and (x,y) represents the coordinate value of the pixel in the image to be processed.

[0076] A typical mapping result can be as follows Figure 2d As shown, in Figure 2d In this process, the trapezoidal road region in the image is mapped to a rectangular image, and the rectangular image has the same size as the original image, which greatly improves the effective information density in the mapped image.

[0077] Optionally, after obtaining the target image, step 205 can be performed to detect objects in the target image in order to identify objects in the target image.

[0078] Optionally, in the present embodiment, the object detection can be implemented based on a deep learning algorithm. Deep learning is a branch of machine learning, and is an algorithm for learning the representation of data based on an artificial neural network. Deep learning can use multiple processing layers containing complex structures or composed of multiple nonlinear transformations to abstract the data at a high level, and perform object recognition based on the features obtained at the high level.

[0079] Optionally, the algorithm model for object detection can be trained in advance based on a training sample using a deep learning algorithm, i.e., a deep learning detector in the algorithm model. Figure 2e Optionally, the algorithm model can be a neural network model (NN). As shown in Figure 2e In the training stage, the road region can be identified from an existing road image and mapped according to the method described in the foregoing embodiments, and the training sample can be obtained. The training sample is an image with a high density of information. The mapping algorithm used for mapping the road region in the road image is the same as the mapping algorithm used for mapping the road region in the image to be processed, so as to ensure that the algorithm model can learn the features well, and thus the description is omitted.

[0080] After the road region is mapped, the road region can be mapped into an image with the same size and shape as the original road image. In this process, the small objects in the distance in the road image can be enlarged, and the relatively large objects in the near distance can be appropriately reduced, i.e., the objects in the mapped training sample are almost in the same order of pixel size. In this case, when the algorithm model performs object detection learning, more effective information can be extracted, and better learning performance can be achieved.

[0081] Optionally, the present embodiment uses a supervised learning model training method. In the supervised learning process, the bounding box of the object is labeled on the training sample (the bounding box is the detection box of the object, and is used to identify the position of the object), and the labeled bounding box can be used as the bounding box ground truth to participate in the supervised process of deep learning.

[0082] Optionally, as shown in Figure 2eAs shown, the bounding box ground truth in the training sample is obtained by mapping the bounding box ground truth labeled in the road image. The mapping algorithm used for mapping the bounding box ground truth in the road image is the same as the mapping algorithm used for mapping the road region in the image to be processed, and will not be described again. For example, in the process of preparing the training sample, the road region can be extracted from a road image, and the bounding box of the object is labeled on the road region according to the position of the object on the road region. Then, according to the set mapping algorithm, the road region and the bounding box therein are mapped to obtain the training sample and the bounding box ground truth therein.

[0083] Based on the above, after the road region in the image to be processed is mapped into the target image by the mapping algorithm, the target image can be input into the algorithm model, the object in the target image is recognized by the algorithm model, and the bounding box of the object is calculated, as shown. Figure 2e For ease of description and differentiation, the bounding box of the object recognized from the target image is described as a first bounding box.

[0084] In step 206, after the first bounding box is obtained, the first bounding box can be reversely mapped to map the first bounding box on the target image to a bounding box on the image to be processed. For ease of description and differentiation, the bounding box on the image to be processed is described as a second bounding box.

[0085] The mapping algorithm used for reversely mapping the first bounding box is the reverse mapping algorithm of the mapping algorithm used for mapping the road region in the image to be processed, which will not be described again. Based on the reverse mapping, the object detection result on the target image can be corresponded to the image to be processed. Next, step 207 can be performed to display the second bounding box on the image to be processed, and the second bounding box is the detection box of the object on the image to be processed.

[0086] In this embodiment, after the road region in the image to be processed is recognized, the road region is mapped into a target image with a set size, so that the redundant information in the image to be processed can be removed, and the effective information density of the target image obtained by mapping is improved. Based on the target image with higher effective information density, the object detection is beneficial to improving the accuracy of the recognition result.

[0087] Figure 3 The flowchart of the data processing method provided by another exemplary embodiment of the present application is shown in the figure, and the method comprises:

[0088] In step 301, a road region in a road image is obtained, and an object in the road region is labeled with a first bounding box.

[0089] Step 302, according to the set mapping algorithm, the road area and the first bounding box are mapped to obtain a training sample of a set size and a second bounding box on the training sample.

[0090] Step 303, input the training sample into the neural network model, and take the second bounding box as a supervision signal to train the object detection capability of the neural network model.

[0091] The mapping algorithm used for mapping the road area and the first bounding box can refer to the description of the foregoing embodiments, and will not be described here. The training sample obtained by mapping is an image with a high density information proportion, which is beneficial to better feature learning of the neural network.

[0092] Optionally, the neural network model (Neural Networks, NN) can be implemented as one or more of a convolutional neural network (Convolutional Neural Networks, CNN), a deep neural network (Deep Neural Network, DNN), a graph convolutional neural network (Graph Convolutional Networks, GCN), a recurrent neural network (Recurrent Neural Network, RNN), and a long short-term memory neural network (Long Short-Term Memory, LSTM), or can be obtained by deforming one or more of the above neural networks, and the present embodiment is not limited.

[0093] In the present embodiment, when training the deep learning detection model, the image with high information density is used as the training sample, which greatly improves the learning efficiency of the neural network model in the offline training stage and the detection accuracy in the online inference stage.

[0094] Figure 4 A typical application scenario of the present application is illustrated, which is a typical application scenario of the present application, and the data processing method provided by the present application can be applied to the intelligent traffic monitoring system. Figure 4 In the schematic diagram, the data processing method provided by the present application can be deployed in an intelligent traffic monitoring system. The intelligent traffic monitoring system can include a monitoring device 41 installed above the road, a server 42, and a management terminal 43. The monitoring device 41 can be implemented as a high-speed camera; the server 42 can be implemented as a cloud server, a data center, etc.; and the management terminal 43 can be implemented as a user terminal of a traffic management unit, such as a computer, a smart phone, a smart display, etc. The present embodiment includes but is not limited to the above.

[0095] After the monitoring device 41 captures the road monitoring video, the road monitoring video can be sent to the server 42. The server 42 performs sampling processing on the road monitoring video to obtain multiple frames of road images. Then, the server 42 can identify the contour of the road according to the distribution rule and characteristics of the road in the road images, and extract the road region from the road images according to the contour of the road. The contour of the road is usually trapezoidal or polygonal. Then, the server 42 can use the mapping method described in the foregoing embodiments to map the road region into a target image with the same size as the original road image.

[0096] Then, the server 42 can input the target image into the deep learning detector and obtain the detection result output by the deep learning detector. Based on the detection result, the server 42 can use the inverse mapping algorithm to map the detection result into the original road image to determine the target object in the road image. Next, the server 42 can distribute the detection result of the target object in the road image to the management terminal 43 to display to the management personnel. Alternatively, the server 42 can also calculate the moving track, motion speed, etc. of the target object according to the target objects detected in multiple frames of continuous road images, and can distribute the calculation result to the management terminal 43, which will not be described herein again.

[0097] The foregoing embodiments describe the application of the data processing method provided by the present application in the field of intelligent transportation. It should be understood that, in addition to the field of intelligent transportation, the data processing method can also be applied to image processing processes in other fields.

[0098] For example, in some scenarios, the data processing method can be applied to the field of face recognition. In the field of face recognition, due to the shooting angle and facial features, different local regions on the face image have different proportions. For example, the nose and forehead parts have a large proportion in the image, and the chin region has a small proportion in the image, which is not conducive to face recognition. In order to improve the accuracy of face recognition, the face region can be identified from the image, and the face region can be mapped into an image with a set size according to the mapping algorithm. The mapping operation can remove information irrelevant to face recognition, expand the face region with a small proportion, and compress the face region with a large proportion. Face recognition based on the mapped image effectively reduces the difficulty of recognizing small targets.

[0099] For another example, in some scenarios, the data processing method can be applied to the field of unmanned aerial vehicle target detection. In the field of unmanned aerial vehicle target detection, the images captured by the unmanned aerial vehicle usually contain far-range objects and near-range objects. The far-range objects take up a small proportion in the image, which is not conducive to target detection. Therefore, the data processing method provided in the embodiments of the present application can be used to identify a to-be-detected region in which a target object may exist from the images captured by the unmanned aerial vehicle, and perform mapping processing on the to-be-detected region according to a mapping algorithm. The mapping processing operation can expand the local region that takes up a small proportion in the to-be-detected region, and compress the local region that takes up a large proportion, so as to balance the information amount of the far-range and the near-range, and remove irrelevant information. Based on the images obtained through mapping, the far-range and the near-range targets can be more accurately detected, which is conducive to providing better obstacle avoidance basis and target tracking basis for the unmanned aerial vehicle.

[0100] Of course, in addition to the above-mentioned fields, it can also be extended to other fields that need to detect small targets, which will not be described here.

[0101] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 201 to 203 can be device A; for another example, the execution subject of steps 201 and 202 can be device A, and the execution subject of step 203 can be device B; and so on.

[0102] In addition, in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appearing in a specific order are included, but it should be clear that these operations can be executed in the order appearing in this document or in parallel, and the serial numbers of the operations, such as 201, 202, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second", etc. described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order, nor do "first" and "second" represent different types.

[0103] Figure 5 FIG. 1 is a structural schematic diagram of an electronic device provided by an example embodiment of the present application, as shown in the figure, the electronic device includes a memory 501 and a processor 502. Figure 5

[0104] The memory 501 is used to store computer programs and can be configured to store other various data to support operations on the electronic device. Examples of these data include instructions of any application program or method for operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.​

[0105] The memory 501 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0106] The processor 502 is coupled to the memory 501 and is configured to execute computer programs in the memory 501 to acquire a to-be-processed image, identify a road region in the to-be-processed image, map the road region into a target image of a set size according to a set mapping algorithm, and perform object detection on the target image to identify objects contained in the road region.

[0107] Further, when mapping the road region into the target image of the set size according to the set mapping algorithm, the processor 502 is specifically configured to determine a mapping relationship according to geometric features of the to-be-processed image and the road region, perform coordinate mapping on pixels in the road region according to the mapping relationship to obtain the target image, and adapt the size of the target image to the to-be-processed image.

[0108] Further, when determining the mapping relationship according to the geometric features of the to-be-processed image and the road region, the processor 502 is specifically configured to, for any pixel in the road region, calculate a mapping coefficient of the pixel according to a coordinate of the pixel, a coordinate range of the road region and / or a size of the to-be-processed image, and calculate a coordinate mapping relationship of the pixel according to the mapping coefficient of the pixel and the coordinate of the pixel; wherein the mapping coefficient includes at least one of a horizontal coordinate mapping coefficient, a vertical coordinate mapping coefficient and a non-linear mapping coefficient.

[0109] Further, when calculating the mapping coefficient of the pixel, the processor 502 is specifically configured to calculate a horizontal coordinate mapping coefficient of the pixel according to a horizontal coordinate range of a row where the pixel is located in the road region and a length of the to-be-processed image.

[0110] Further, when calculating the mapping coefficient of the pixel according to the horizontal coordinate range of the row where the pixel is located in the road region and the length of the to-be-processed image, the processor 502 is specifically configured to take a ratio of a length of the horizontal coordinate range to the length of the to-be-processed image as the horizontal coordinate mapping coefficient.

[0111] Further, the processor 502 is further configured to calculate the coordinate mapping relationship of the pixel according to the mapping coefficient of the pixel and the coordinate of the pixel, and specifically configured to: calculate a product of the horizontal coordinate of the pixel and the horizontal coordinate mapping coefficient; and determine a summation relationship between a horizontal minimum value of the row in which the pixel is located in the road region and the product as the horizontal coordinate mapping relationship of the pixel.

[0112] Further, the processor 502 is further configured to calculate the mapping coefficient of the pixel, and specifically configured to: calculate the vertical coordinate mapping coefficient of the pixel according to the vertical coordinate range of the road region and the height of the to-be-processed image.

[0113] Further, the processor 502 is further configured to calculate the mapping coefficient of the pixel, and specifically configured to: calculate the horizontal coordinate mapping coefficient of the pixel according to the horizontal coordinate range of the row in which the pixel is located in the road region and the length of the to-be-processed image; and calculate the nonlinear mapping coefficient of the pixel according to the vertical coordinate of the pixel, the range to which the vertical coordinate of the pixel belongs, and the horizontal coordinate mapping coefficient of the pixel.

[0114] Further, the processor 502 is further configured to calculate the coordinate mapping relationship of the pixel according to the mapping coefficient of the pixel and the coordinate of the pixel, and specifically configured to: calculate a product of the vertical coordinate of the pixel, the vertical coordinate mapping coefficient and the nonlinear mapping coefficient; and determine a summation relationship between the vertical minimum value of the column in which the pixel is located and the product as the vertical coordinate mapping relationship of the pixel.

[0115] Further, the processor 502 is further configured to identify the road region in the to-be-processed image, and specifically configured to: perform edge detection on the to-be-processed image to obtain edge information contained in the to-be-processed image; determine a road contour according to the edge information; calculate an inscribed trapezoid of the road contour, and take an image region corresponding to the inscribed trapezoid as the road region.

[0116] Further, the processor 502 is further configured to perform object detection on the target image to identify an object contained in the road region, and specifically configured to: identify the object in the target image and calculate a first bounding box of the object; perform reverse mapping processing on the first bounding box according to a reverse mapping algorithm corresponding to the mapping algorithm, to map the first bounding box to a second bounding box on the to-be-processed image; and display the second bounding box on the to-be-processed image.

[0117] Further, the processor 502 is configured to identify the object in the target image and calculate the first bounding box of the object by inputting the target image into an algorithm model, and identifying the object in the target image and calculating the first bounding box of the object by the algorithm model. The training sample used for training the algorithm model is obtained by mapping the road region in the road image according to the mapping algorithm, and the ground truth of the bounding box labeled in the training sample is obtained by mapping the ground truth of the bounding box labeled in the road image according to the mapping algorithm.

[0118] Further, as shown in Figure 5 the electronic device further includes a communication component 503, a display component 504, a power supply component 505, an audio component 506, and other components. Figure 5 Some components are only schematically shown in the electronic device, and it does not mean that the electronic device only includes Figure 5 the components shown.

[0119] The communication component 503 is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G or 5G, or a combination thereof. In an example embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component can be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0120] The display component 504 includes a screen, which can include a liquid crystal display component (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect a duration and a pressure related to the touch or slide action.

[0121] The power supply component 505 provides power to various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.

[0122] In this embodiment, after identifying the road region in the image to be processed, the road region is mapped to a target image of a set size. This removes redundant information from the image to be processed and improves the effective information density of the mapped target image. Object detection based on a target image with higher effective information density is beneficial to improving the accuracy of the recognition results.

[0123] In addition to the execution logic described in the foregoing embodiments, Figure 5 The illustrated electronic device can also perform the following data processing logic: acquire a road region in a road image through processor 502, wherein objects in the road region are marked with a first bounding box; perform mapping processing on the road region and the first bounding box according to a set mapping algorithm to obtain a training sample of a set size and a second bounding box on the training sample; input the training sample into a neural network model, and use the second bounding box as a supervision signal to train the object detection capability of the neural network model.

[0124] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by an electronic device in the above method embodiments.

[0125] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0127] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.

[0129] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0130] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) and / or cache memory, non-volatile memory, such as read-only memory (ROM), EPROM, and / or flash memory, etc. The memory is an example of computer readable media.

[0131] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.

[0132] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0133] The above description is merely illustrative of the application, and not restrictive. Various modifications and changes can become apparent to those skilled in the art. Incorporating any modification, equivalent substitution, improvement, etc. within the spirit and principle of the application, shall be included in the scope of the claims of the application.

Claims

1. A data processing method, characterized in that, include: Obtain the image to be processed; Identify road regions in the image to be processed; According to the set mapping algorithm, the road area is mapped to a target image of a set size; A neural network model is used to perform object detection on the target image to identify objects contained in the road area; The mapping algorithm is determined based on the geometric features and size ratio of the road area; the geometric features are used to determine the mapping relationship including nonlinear mapping coefficients, and the nonlinear mapping coefficients of the near and far areas in the road area are different; the size ratio is used to enlarge or reduce different local locations in the road area.

2. The method according to claim 1, characterized in that, According to the set mapping algorithm, the road area is mapped to a target image of a set size, including: The mapping relationship is determined based on the geometric features of the image to be processed and the road area; Based on the mapping relationship, the pixels in the road area are mapped to obtain the target image; the size of the target image is adapted to the image to be processed.

3. The method according to claim 2, characterized in that, Determining the mapping relationship based on the geometric features of the image to be processed and the road region includes: For any pixel in the road region, a mapping coefficient for the pixel is calculated based on the pixel's coordinates, the coordinate range of the road region, and / or the size of the image to be processed. Calculate the coordinate mapping relationship of the pixel based on the mapping coefficient of the pixel and the coordinates of the pixel; The mapping coefficients include at least one of the following: horizontal coordinate mapping coefficient, vertical coordinate mapping coefficient, and nonlinear mapping coefficient.

4. The method according to claim 3, characterized in that, Calculating the mapping coefficients of the pixels includes: The horizontal coordinate mapping coefficient of the pixel is calculated based on the range of the horizontal coordinates of the row where the pixel is located in the road area and the length of the image to be processed.

5. The method according to claim 4, characterized in that, The mapping coefficient of the pixel is calculated based on the range of the horizontal coordinates of the row containing the pixel in the road region and the length of the image to be processed, including: The ratio of the length of the horizontal coordinate range to the length of the image to be processed is used as the horizontal coordinate mapping coefficient.

6. The method according to claim 3, characterized in that, Based on the mapping coefficients of the pixel and the coordinates of the pixel, the coordinate mapping relationship of the pixel is calculated, including: Calculate the product of the x-coordinate of the pixel and the x-coordinate mapping coefficient; The summation of the minimum horizontal coordinate of the row containing the pixel in the road region and the product of these values ​​is determined as the horizontal coordinate mapping relationship of the pixel.

7. The method according to claim 3, characterized in that, Calculating the mapping coefficients of the pixels includes: The ordinate mapping coefficient of the pixel is calculated based on the ordinate range of the road area and the height of the image to be processed.

8. The method according to claim 3, characterized in that, Calculating the mapping coefficients of the pixels includes: The horizontal coordinate mapping coefficient of the pixel is calculated based on the range of the horizontal coordinates of the row where the pixel is located in the road area and the length of the image to be processed. The nonlinear mapping coefficient of the pixel is calculated based on the pixel's ordinate, the range to which the pixel's ordinate belongs, and the pixel's abscissa mapping coefficient.

9. The method according to claim 3, characterized in that, Based on the mapping coefficients of the pixel and the coordinates of the pixel, the coordinate mapping relationship of the pixel is calculated, including: Calculate the product of the pixel's ordinate, the ordinate mapping coefficient, and the nonlinear mapping coefficient; The minimum value of the vertical coordinate of the column containing the pixel and the sum of their products are determined as the vertical coordinate mapping relationship of the pixel.

10. The method according to any one of claims 1-9, characterized in that, Identifying road regions in the image to be processed includes: Edge detection is performed on the image to be processed to obtain the edge information contained in the image to be processed; Based on the edge information, the road outline is determined; Calculate the inscribed trapezoid of the road profile, and use the image region corresponding to the inscribed trapezoid as the road region.

11. The method according to any one of claims 1-9, characterized in that, Performing object detection on the target image to identify objects contained in the road area includes: Identify objects in the target image and calculate the first bounding box of the objects; According to the reverse mapping algorithm corresponding to the mapping algorithm, the first bounding box is reverse mapped to the second bounding box on the image to be processed. The second bounding box is displayed on the image to be processed.

12. The method according to claim 11, characterized in that, Identifying objects in the target image and calculating the first bounding box of the objects includes: The target image is input into the algorithm model to identify objects in the target image and calculate the first bounding box of the objects. The training samples used to train the algorithm model are obtained by mapping road regions in road images according to the mapping algorithm; the ground truth values ​​of the bounding boxes labeled in the training samples are obtained by mapping the ground truth values ​​of the bounding boxes labeled in the road images according to the mapping algorithm.

13. A data processing method, characterized in that, include: Obtain the road region in the road image, wherein objects in the road region are marked with a first bounding box; According to the set mapping algorithm, the road area and the first bounding box are mapped to obtain a training sample of a set size and a second bounding box on the training sample. The training samples are input into the neural network model, and the second bounding box is used as a supervision signal to train the object detection capability of the neural network model. The mapping algorithm is determined based on the geometric features and size ratio of the road area; the geometric features are used to determine the mapping relationship including nonlinear mapping coefficients, and the nonlinear mapping coefficients of the near and far areas in the road area are different; the size ratio is used to enlarge or reduce different local locations in the road area.

14. An electronic device, characterized in that, include: Memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute one or more computer instructions for: performing the data processing method according to any one of claims 1-13.

15. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it is able to implement the data processing method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Image processing method, device and equipment and computer readable storage medium

    CN110458164A