Image processing method and device
By acquiring image data in scenarios such as autonomous driving, extracting and enhancing features in bird's eye view space, and fusing features for target prediction, the problem of poor target prediction results caused by insufficient clear and comprehensive images in the prior art is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202210118870.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-02-08
AI Technical Summary
When the prior art uses image data to predict targets in scenarios such as autonomous driving, the target prediction effect is limited by the problem of insufficient clear and comprehensive images.
By acquiring image data, extracting image perspective features, and performing feature enhancement processing in the bird's-eye view space, integrating bird's-eye view features for target prediction.
The goal prediction effect from a bird's eye perspective is improved, the accuracy of goal prediction is improved, and problems caused by the enhancement of a single feature are avoided by fusion of different features.
Smart Images

Figure CN114612353B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image technology, and in particular to an image processing method and device. Background Art
[0002] At present, the realization of more and more functions requires the use of image data for target prediction. For example, in the process of autonomous driving, image data needs to be collected to predict environmental conditions such as vehicles and pedestrians.
[0003] In the prior art, the acquired image data is usually directly used for target prediction. However, due to the problems that the acquired images are not clear and comprehensive enough, the effect of target prediction using this method is not good. Summary of the invention
[0004] In view of the above problems, a method and apparatus for image processing is proposed to overcome the above problems or at least partially solve the above problems, including:
[0005] A method for image processing, the method comprising:
[0006] Acquire first image data, and obtain first image viewing angle features according to the first image data;
[0007] A first bird's-eye view feature is obtained according to the first image view feature, and a feature enhancement process is performed on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature;
[0008] Determining a third bird's-eye view perspective feature for the first image perspective feature;
[0009] The second bird's-eye view feature and the third bird's-eye view feature are fused to obtain a fourth bird's-eye view feature, and target prediction is performed based on the fourth bird's-eye view feature.
[0010] Optionally, determining a third bird's-eye view perspective feature for the first image perspective feature includes:
[0011] The first bird's-eye view perspective feature is determined as the third bird's-eye view perspective feature.
[0012] Optionally, determining a third bird's-eye view perspective feature for the first image perspective feature includes:
[0013] Performing feature enhancement processing on the first image perspective feature in the image perspective space to obtain a second image perspective feature;
[0014] A third bird's-eye view feature is obtained according to the second image view feature.
[0015] Optionally, obtaining a third bird's-eye view feature according to the second image view feature includes:
[0016] Determining depth information corresponding to the viewing angle feature of the second image;
[0017] According to the depth information corresponding to the second image perspective feature, the second image perspective feature is transformed to obtain a third bird's-eye view perspective feature.
[0018] Optionally, obtaining a first bird's-eye view feature according to the first image view feature includes:
[0019] Determining depth information corresponding to the viewing angle feature of the first image;
[0020] According to the depth information corresponding to the first image perspective feature, the first image perspective feature is transformed to obtain a first bird's-eye view perspective feature.
[0021] Optionally, target prediction is performed based on the fourth bird's-eye view feature, including:
[0022] Based on the fourth bird's-eye view feature, attribute information of the corresponding target in the bird's-eye view space is predicted, where the attribute information includes any one or more of the following:
[0023] Category information, size information, orientation information, speed information.
[0024] Optionally, the first image data is image data collected during the autonomous driving process.
[0025] Optionally, the third bird's-eye view feature is a bird's-eye view feature that has not been subjected to feature enhancement processing in the bird's-eye view space.
[0026] An image processing device, comprising:
[0027] A first image viewing angle feature obtaining module, used to obtain first image data and obtain first image viewing angle features according to the first image data;
[0028] A second bird's-eye view feature obtaining module is used to obtain a first bird's-eye view feature according to the first image view feature, and perform feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature;
[0029] A third bird's-eye view feature determination module, used to determine a third bird's-eye view feature for the first image view feature;
[0030] The target prediction module is used to perform feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and perform target prediction based on the fourth bird's-eye view feature.
[0031] An electronic device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the above-mentioned image processing method when executed by the processor.
[0032] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method described above is implemented.
[0033] The embodiments of the present invention have the following advantages:
[0034] In an embodiment of the present invention, by acquiring first image data, and obtaining a first image perspective feature based on the first image data, obtaining a first bird's-eye view feature based on the first image perspective feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature, and then determining a third bird's-eye view feature for the first image perspective feature, and then performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature, it is achieved that feature enhancement technology is used to improve the target prediction effect under the bird's-eye view, and the accuracy of target prediction is improved, and by fusing features that have been feature enhanced and features that have not been feature enhanced, problems caused by only using features that have been feature enhanced are avoided. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0036] Figure 1a is a flowchart of a method for image processing provided by one embodiment of the present invention;
[0037] Figure 1b is a schematic diagram of an image processing example provided by an embodiment of the present invention;
[0038] Figure 2a is a flowchart of another image processing method provided by an embodiment of the present invention;
[0039] Figure 2b is a schematic diagram of another image processing example provided by an embodiment of the present invention;
[0040] Figure 3a is a flowchart of another image processing method provided by an embodiment of the present invention;
[0041] Figure 3b is a schematic diagram of another image processing example provided by an embodiment of the present invention;
[0042] Figure 4 It is a structural block diagram of an image processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] Reference Figure 1a , shows a flowchart of a method for image processing provided by an embodiment of the present invention, which may specifically include the following steps:
[0045] Step 101: Acquire first image data, and obtain first image viewing angle features based on the first image data.
[0046] Among them, the first image data may be image data collected during the process of autonomous driving, which may be collected by a camera deployed in a smart cockpit (such as a vehicle).
[0047] In some scenarios, there is a need to predict targets in the environment based on the collected image data. For example, during autonomous driving, it is necessary to predict the category, size, direction, speed and other attributes of targets outside the vehicle. In this case, the first image data collected for the environment can be obtained.
[0048] After obtaining the first image data, it can be encoded using a preset image space encoder. The image space encoder can be an encoder based on a deep neural network, such as a resnet deep neural network. The first image data is input into the image space encoder to obtain the first image viewing angle feature.
[0049] Step 102: obtaining a first bird's-eye view feature according to the first image view feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature.
[0050] Since the image perspective feature is a feature in the image perspective space (2D), the features in the image perspective space are expressed in units of pixels in the image. This method will cause the same target to have different positions, sizes and other information depending on the distance at which it is photographed in the image. The bird's-eye view features in the bird's-eye view space (3D) are expressed in units in a three-dimensional coordinate system, such as meters. The position, size and other information of the same target are fixed, which is more suitable for target prediction.
[0051] Based on this, the first image perspective feature in the image perspective space may be transformed to obtain the first bird's-eye view perspective feature in the bird's-eye view perspective space.
[0052] After obtaining the first bird's-eye view feature, a preset bird's-eye view space encoder can be used to perform feature enhancement processing on it in the bird's-eye view space. The bird's-eye view space encoder can be an encoder based on a deep neural network. The first bird's-eye view feature is input into an image space encoder to obtain a second bird's-eye view feature.
[0053] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the value of the corresponding first bird's-eye view feature is 0.5. After being processed by the bird's-eye view spatial encoder, its value can be enhanced to 0.8, which makes it clear that its type is a vehicle.
[0054] In an embodiment of the present invention, obtaining the first bird's-eye view feature according to the first image view feature may include:
[0055] Sub-step 11: determining depth information corresponding to the viewing angle feature of the first image.
[0056] In a specific implementation, a depth prediction module can be set in advance, and the depth prediction module can be used to perform dense depth estimation to obtain depth information corresponding to the perspective feature of the first image. The dense depth estimation is to estimate the distance from the position of the feature corresponding to each pixel of the image in three-dimensional space to the optical center of the camera.
[0057] Sub-step 12: performing a perspective transformation on the first image perspective feature according to the depth information corresponding to the first image perspective feature to obtain a first bird's-eye view perspective feature.
[0058] Since there is a mapping relationship between the coordinates in the image perspective space and the coordinates in the three-dimensional space, after obtaining the depth information, the preset perspective conversion module can be used to transform the first image perspective feature in combination with the depth information to obtain the first bird's-eye view perspective feature.
[0059] Step 103, determining a third bird's-eye view feature for the first image view feature;
[0060] The third bird's-eye view feature may be a bird's-eye view feature that has not been subjected to feature enhancement processing in the bird's-eye view space.
[0061] In some scenarios, such as Figure 1b After the first image data is processed by the image space encoder, a first image perspective feature is obtained. The first image perspective feature is passed through a depth prediction module to obtain dense depth information. Combining the first image perspective feature and the depth information, a first bird's-eye view perspective feature can be obtained through a perspective conversion module. The first bird's-eye view perspective feature is then processed by a bird's-eye view perspective space encoder to obtain a feature-enhanced second bird's-eye view perspective feature, which can then be used for target prediction through a prediction module to directly obtain a prediction result.
[0062] However, through actual data analysis, it is found that the bird's-eye view spatial encoder will forget the image spatial features during the encoding process, resulting in poor performance in tasks such as target classification. The classification task refers to judging the category attributes of the target. For example, if there are pedestrians, vehicles and other targets in the scene, judging whether a target is a pedestrian or a vehicle is a classification task.
[0063] Based on this, on the basis of using the bird's-eye view spatial encoder for feature enhancement, the third bird's-eye view feature for the first image perspective feature can be combined. It can be a bird's-eye view feature that has not been feature enhanced in the bird's-eye view space by the bird's-eye view spatial encoder, so as to make up for the defects of the bird's-eye view spatial encoder in the encoding process and increase the accuracy of subsequent target prediction.
[0064] Step 104 , performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature.
[0065] After obtaining the second bird's-eye view feature and the third bird's-eye view feature after feature enhancement processing, the two can be fused using a preset feature fusion module. Specifically, the two features can be added or combined to obtain a fourth bird's-eye view feature.
[0066] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the corresponding value of the second bird's-eye view feature is 0.8, and the value of the third bird's-eye view feature is 0.6. Then, it can be determined that the value of the fourth bird's-eye view feature is 0.7.
[0067] After obtaining the fourth bird's-eye view feature, a preset prediction module can be used to perform target prediction on it. The prediction module can be a module based on a deep neural network. The fourth bird's-eye view feature prediction module is used to obtain a prediction result to better assist functions such as autonomous driving.
[0068] In one embodiment of the present invention, performing target prediction based on the fourth bird's-eye view feature may include:
[0069] Sub-step 21, based on the fourth bird's-eye view feature, predicting the attribute information of the corresponding target in the bird's-eye view space, the attribute information includes any one or more of the following:
[0070] Category information, size information, orientation information, speed information.
[0071] As an example, the target may include a movable target and a stationary target. For example, the movable target may be a vehicle, a pedestrian, etc., and the stationary target may be a signpost, a traffic obstacle, etc.
[0072] In a specific implementation, the attribute information of the target corresponding to the fourth bird's-eye view feature can be predicted, for example, determining the category of the target as a pedestrian, vehicle or other object, or determining the length, width, height and other dimensions of the target, or determining whether the target is facing the current vehicle or facing away from the current vehicle, or determining the speed value of the current vehicle.
[0073] In an embodiment of the present invention, by acquiring first image data, and obtaining a first image perspective feature based on the first image data, obtaining a first bird's-eye view feature based on the first image perspective feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature, and then determining a third bird's-eye view feature for the first image perspective feature, and then performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature, it is achieved that feature enhancement technology is used to improve the target prediction effect under the bird's-eye view, and the accuracy of target prediction is improved, and by fusing features that have been feature enhanced and features that have not been feature enhanced, problems caused by only using features that have been feature enhanced are avoided.
[0074] Reference Figure 2a , shows a flowchart of another image processing method provided by an embodiment of the present invention, which may specifically include the following steps:
[0075] Step 201: Acquire first image data, and obtain first image viewing angle features based on the first image data.
[0076] Among them, the first image data may be image data collected during the process of autonomous driving, which may be collected by a camera deployed in a smart cockpit (such as a vehicle).
[0077] In some scenarios, there is a need to predict targets in the environment based on the collected image data. For example, during autonomous driving, it is necessary to predict the category, size, direction, speed and other attributes of targets outside the vehicle. In this case, the first image data collected for the environment can be obtained.
[0078] After obtaining the first image data, it can be encoded using a preset image space encoder. The image space encoder can be an encoder based on a deep neural network, such as a resnet deep neural network. The first image data is input into the image space encoder to obtain the first image viewing angle feature.
[0079] Step 202: obtaining a first bird's-eye view feature according to the first image view feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature.
[0080] Since the image perspective feature is a feature in the image perspective space (2D), the features in the image perspective space are expressed in units of pixels in the image. This method will cause the same target to have different positions, sizes and other information depending on the distance at which it is photographed in the image. The bird's-eye view features in the bird's-eye view space (3D) are expressed in units in a three-dimensional coordinate system, such as meters. The position, size and other information of the same target are fixed, which is more suitable for target prediction.
[0081] Based on this, the first image perspective feature in the image perspective space may be transformed to obtain the first bird's-eye view perspective feature in the bird's-eye view perspective space.
[0082] After obtaining the first bird's-eye view feature, a preset bird's-eye view space encoder can be used to perform feature enhancement processing on it in the bird's-eye view space. The bird's-eye view space encoder can be an encoder based on a deep neural network. The first bird's-eye view feature is input into an image space encoder to obtain a second bird's-eye view feature.
[0083] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the value of the corresponding first bird's-eye view feature is 0.5. After being processed by the bird's-eye view spatial encoder, its value can be enhanced to 0.8, which makes it clear that its type is a vehicle.
[0084] Step 203: determine the first bird's-eye view feature as the third bird's-eye view feature.
[0085] In some scenarios, such as Figure 1b After the first image data is processed by the image space encoder, a first image perspective feature is obtained. The first image perspective feature is passed through a depth prediction module to obtain dense depth information. Combining the first image perspective feature and the depth information, a first bird's-eye view perspective feature can be obtained through a perspective conversion module. The first bird's-eye view perspective feature is then processed by a bird's-eye view perspective space encoder to obtain a feature-enhanced second bird's-eye view perspective feature, which can then be used for target prediction through a prediction module to directly obtain a prediction result.
[0086] However, through actual data analysis, it is found that the bird's-eye view spatial encoder will forget the image spatial features during the encoding process, resulting in poor performance in tasks such as target classification. The classification task refers to judging the category attributes of the target. For example, if there are pedestrians, vehicles and other targets in the scene, judging whether a target is a pedestrian or a vehicle is a classification task.
[0087] Based on this, on the basis of using the bird's-eye view spatial encoder for feature enhancement, the third bird's-eye view feature can be combined, which can be a bird's-eye view feature that has not been enhanced by the bird's-eye view spatial encoder in the bird's-eye view space, to make up for the defects of the bird's-eye view spatial encoder in the encoding process and increase the accuracy of subsequent target prediction.
[0088] Specifically, the first bird's-eye view feature obtained by directly converting the first image view feature may be used as the third bird's-eye view feature that has not been subjected to feature enhancement processing.
[0089] Step 204 , feature fusion is performed on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and target prediction is performed based on the fourth bird's-eye view feature.
[0090] After obtaining the second bird's-eye view feature and the third bird's-eye view feature after feature enhancement processing, the two can be fused using a preset feature fusion module. Specifically, the two features can be added or combined to obtain a fourth bird's-eye view feature.
[0091] like Figure 2b ,exist Figure 1b On the basis of the present invention, a feature fusion module is added to fuse the first bird's-eye view feature obtained by the perspective transformation module from the first image perspective feature with the feature-enhanced second bird's-eye view feature to obtain a fourth bird's-eye view feature.
[0092] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the corresponding value of the second bird's-eye view feature is 0.8, and the value of the third bird's-eye view feature is 0.6. Then, it can be determined that the value of the fourth bird's-eye view feature is 0.7.
[0093] After obtaining the fourth bird's-eye view feature, a preset prediction module can be used to perform target prediction on it. The prediction module can be a module based on a deep neural network. The fourth bird's-eye view feature prediction module is used to obtain a prediction result to better assist functions such as autonomous driving.
[0094] Reference Figure 3a , shows a flowchart of another image processing method provided by an embodiment of the present invention, which may specifically include the following steps:
[0095] Step 301: Acquire first image data, and obtain first image viewing angle features based on the first image data.
[0096] Among them, the first image data may be image data collected during the process of autonomous driving, which may be collected by a camera deployed in a smart cockpit (such as a vehicle).
[0097] In some scenarios, there is a need to predict targets in the environment based on the collected image data. For example, during autonomous driving, it is necessary to predict the category, size, direction, speed and other attributes of targets outside the vehicle. In this case, the first image data collected for the environment can be obtained.
[0098] After obtaining the first image data, it can be encoded using a preset image space encoder. The image space encoder can be an encoder based on a deep neural network, such as a resnet deep neural network. The first image data is input into the image space encoder to obtain the first image viewing angle feature.
[0099] Step 302: Obtain a first bird's-eye view feature according to the first image view feature, and perform feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature.
[0100] Since the image perspective feature is a feature in the image perspective space (2D), the features in the image perspective space are expressed in units of pixels in the image. This method will cause the same target to have different positions, sizes and other information depending on the distance at which it is photographed in the image. The bird's-eye view features in the bird's-eye view space (3D) are expressed in units in a three-dimensional coordinate system, such as meters. The position, size and other information of the same target are fixed, which is more suitable for target prediction.
[0101] Based on this, the first image perspective feature in the image perspective space may be transformed to obtain the first bird's-eye view perspective feature in the bird's-eye view perspective space.
[0102] After obtaining the first bird's-eye view feature, a preset bird's-eye view space encoder can be used to perform feature enhancement processing on it in the bird's-eye view space. The bird's-eye view space encoder can be an encoder based on a deep neural network. The first bird's-eye view feature is input into an image space encoder to obtain a second bird's-eye view feature.
[0103] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the value of the corresponding first bird's-eye view feature is 0.5. After being processed by the bird's-eye view spatial encoder, its value can be enhanced to 0.8, which makes it clear that its type is a vehicle.
[0104] Step 303: Perform feature enhancement processing on the first image viewing angle feature in the image viewing angle space to obtain a second image viewing angle feature.
[0105] In some scenarios, such as Figure 1b After the first image data is processed by the image space encoder, a first image perspective feature is obtained. The first image perspective feature is passed through a depth prediction module to obtain dense depth information. Combining the first image perspective feature and the depth information, a first bird's-eye view perspective feature can be obtained through a perspective conversion module. The first bird's-eye view perspective feature is then processed by a bird's-eye view perspective space encoder to obtain a feature-enhanced second bird's-eye view perspective feature, which can then be used for target prediction through a prediction module to directly obtain a prediction result.
[0106] However, through actual data analysis, it is found that the bird's-eye view spatial encoder will forget the image spatial features during the encoding process, resulting in poor performance in tasks such as target classification. The classification task refers to judging the category attributes of the target. For example, if there are pedestrians, vehicles and other targets in the scene, judging whether a target is a pedestrian or a vehicle is a classification task.
[0107] Based on this, on the basis of using the bird's-eye view spatial encoder for feature enhancement, the third bird's-eye view feature can be combined, which can be the bird's-eye view feature that has not been enhanced by the bird's-eye view spatial encoder to make up for the defects of the bird's-eye view spatial encoder in the encoding process and increase the accuracy of subsequent target prediction.
[0108] Specifically, a preset image space encoder may be used to perform feature enhancement processing on the first image perspective feature in the image perspective space to obtain the second image perspective feature. The feature enhancement processing process may refer to the feature enhancement processing process on the bird's-eye view feature described above.
[0109] Step 304: Obtain a third bird's-eye view feature according to the second image view feature.
[0110] After the second image viewing angle feature is obtained, the second image viewing angle feature in the image viewing angle space may be transformed to obtain a third bird's-eye view viewing angle feature in the bird's-eye view viewing angle space.
[0111] In one embodiment of the present invention, obtaining the third bird's-eye view feature according to the second image view feature may include:
[0112] Sub-step 31, determining depth information corresponding to the viewing angle feature of the second image.
[0113] In a specific implementation, a depth prediction module can be set in advance, and the depth prediction module can be used to perform dense depth estimation to obtain depth information corresponding to the perspective feature of the second image. The dense depth estimation is to estimate the distance from the position of the feature corresponding to each pixel of the image in three-dimensional space to the optical center of the camera.
[0114] It should be noted that the second image viewing angle feature is obtained by performing feature enhancement processing on the first image viewing angle feature, but the depth information of the two is the same, so the depth information of the first image viewing angle feature can be directly used as the depth information of the second image viewing angle feature.
[0115] Sub-step 32: performing a perspective transformation on the second image perspective feature according to the depth information corresponding to the second image perspective feature to obtain a third bird's-eye view perspective feature.
[0116] Since there is a mapping relationship between the coordinates in the image perspective space and the coordinates in the three-dimensional space, after obtaining the depth information, the preset perspective conversion module can be used to transform the perspective of the second image perspective feature in combination with the depth information to obtain the second bird's-eye view perspective feature.
[0117] Step 305 , performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature.
[0118] After obtaining the second bird's-eye view feature and the third bird's-eye view feature after feature enhancement processing, the two can be fused using a preset feature fusion module. Specifically, the two features can be added or combined to obtain a fourth bird's-eye view feature.
[0119] like Figure 3b ,exist Figure 1b On the basis of, on the one hand, after the first image data is processed by the image space encoder to obtain the first image perspective feature, the image space encoder can be used to process the first image perspective feature to obtain the feature-enhanced second image perspective feature. On the other hand, the depth information obtained from the first image perspective feature can be used to transform the feature-enhanced second image perspective feature through the perspective transformation module to obtain the third bird's-eye view feature, and then the feature fusion module can be used to fuse the third bird's-eye view feature with the feature-enhanced second bird's-eye view feature to obtain the fourth bird's-eye view feature.
[0120] For example, features can be expressed using a multidimensional array. Suppose the target category represented by the value 0 is a pedestrian, and the target category represented by the value 1 is a vehicle. For a certain vehicle target, the corresponding value of the second bird's-eye view feature is 0.8, and the value of the third bird's-eye view feature is 0.6. Then, it can be determined that the value of the fourth bird's-eye view feature is 0.7.
[0121] After obtaining the fourth bird's-eye view feature, a preset prediction module can be used to perform target prediction on it. The prediction module can be a module based on a deep neural network. The fourth bird's-eye view feature prediction module is used to obtain a prediction result to better assist functions such as autonomous driving.
[0122] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0123] Reference Figure 4 , shows a schematic diagram of the structure of an image processing device provided by an embodiment of the present invention, which may specifically include the following modules:
[0124] The first image viewing angle feature obtaining module 401 is used to obtain first image data and obtain the first image viewing angle feature according to the first image data.
[0125] The second bird's-eye view feature obtaining module 402 is used to obtain a first bird's-eye view feature according to the first image view feature, and perform feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature.
[0126] The third bird's-eye view feature determination module 403 is used to determine a third bird's-eye view feature for the first image view feature.
[0127] The target prediction module 404 is used to perform feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and perform target prediction based on the fourth bird's-eye view feature.
[0128] In one embodiment of the present invention, the third bird's-eye view feature determination module 403 may include:
[0129] The submodule for determining the first feature as the third feature is used to determine the first bird's-eye view feature as the third bird's-eye view feature.
[0130] In one embodiment of the present invention, the third bird's-eye view feature determination module 403 may include:
[0131] The second image viewing angle feature obtaining submodule is used to perform feature enhancement processing on the first image viewing angle feature in the image viewing angle space to obtain the second image viewing angle feature.
[0132] The third feature submodule is obtained according to the second feature, and is used to obtain a third bird's-eye view feature according to the second image view feature.
[0133] In one embodiment of the present invention, obtaining the third feature submodule according to the second feature may include:
[0134] The second feature depth information determining unit is used to determine the depth information corresponding to the second image viewing angle feature.
[0135] The second feature perspective transformation unit is used to transform the perspective of the second image perspective feature according to the depth information corresponding to the second image perspective feature to obtain a third bird's-eye view perspective feature.
[0136] In one embodiment of the present invention, the second bird's-eye view feature obtaining module 402 may include:
[0137] The first feature depth information determination submodule is used to determine the depth information corresponding to the first image viewing angle feature.
[0138] The first feature perspective transformation submodule is used to transform the perspective of the first image perspective feature according to the depth information corresponding to the first image perspective feature to obtain a first bird's-eye view perspective feature.
[0139] In one embodiment of the present invention, the target prediction module 404 may include:
[0140] The attribute information prediction submodule is used to predict the attribute information of the corresponding target in the bird's-eye view space based on the fourth bird's-eye view feature, and the attribute information includes any one or more of the following:
[0141] Category information, size information, orientation information, speed information.
[0142] In one embodiment of the present invention, the first image data is image data collected during the process of autonomous driving.
[0143] In one embodiment of the present invention, the third bird's-eye view feature is a bird's-eye view feature that has not been subjected to feature enhancement processing in the bird's-eye view space.
[0144] In an embodiment of the present invention, by acquiring first image data, and obtaining a first image perspective feature based on the first image data, obtaining a first bird's-eye view feature based on the first image perspective feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature, and then determining a third bird's-eye view feature for the first image perspective feature, and then performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature, it is achieved that feature enhancement technology is used to improve the target prediction effect under the bird's-eye view, and the accuracy of target prediction is improved, and by fusing features that have been feature enhanced and features that have not been feature enhanced, problems caused by only using features that have been feature enhanced are avoided.
[0145] An embodiment of the present invention further provides an electronic device, which may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and the computer program implements the above image processing method when executed by the processor.
[0146] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above image processing method is implemented.
[0147] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0148] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0149] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0150] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0151] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0152] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0153] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0154] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0155] The above is a detailed introduction to the provided image processing method and device. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for image processing, It is characterized in that The method comprises: Acquire first image data, and obtain first image viewing angle features according to the first image data; Obtaining a first bird's-eye view feature according to the first image view feature, and performing feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature; Determining a third bird's-eye view perspective feature for the first image perspective feature; Performing feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and performing target prediction based on the fourth bird's-eye view feature; The determining of a third bird's-eye view feature for the first image view feature includes: Performing feature enhancement processing on the first image perspective feature in the image perspective space to obtain a second image perspective feature; Obtaining a third bird's-eye view feature according to the second image view feature; The step of obtaining a third bird's-eye view feature according to the second image view feature includes: Determining depth information corresponding to the second image viewing angle feature; According to the depth information corresponding to the second image perspective feature, the second image perspective feature is transformed to obtain a third bird's-eye view perspective feature.
2. The method according to claim 1, It is characterized in that The determining of a third bird's-eye view feature for the first image view feature comprises: The first bird's-eye view perspective feature is determined as a third bird's-eye view perspective feature.
3. The method according to claim 1 or 2, It is characterized in that The obtaining a first bird's-eye view feature according to the first image view feature comprises: Determining depth information corresponding to the first image viewing angle feature; According to the depth information corresponding to the first image perspective feature, the first image perspective feature is transformed to obtain a first bird's-eye view perspective feature.
4. The method according to claim 1, It is characterized in that The performing target prediction based on the fourth bird's-eye view feature includes: Based on the fourth bird's-eye view feature, attribute information of the corresponding target in the bird's-eye view space is predicted, where the attribute information includes any one or more of the following: Category information, size information, orientation information, speed information.
5. The method according to claim 1, It is characterized in that The first image data is image data collected during the autonomous driving process.
6. The method according to claim 1 or 2, It is characterized in that The third bird's-eye view feature is a bird's-eye view feature that has not been subjected to feature enhancement processing in the bird's-eye view space.
7. An image processing device, It is characterized in that The device comprises: A first image viewing angle feature obtaining module, used to obtain first image data, and obtain a first image viewing angle feature according to the first image data; a second bird's-eye view feature obtaining module, configured to obtain a first bird's-eye view feature according to the first image view feature, and perform feature enhancement processing on the first bird's-eye view feature in a bird's-eye view space to obtain a second bird's-eye view feature; A third bird's-eye view feature determination module, used to determine a third bird's-eye view feature for the first image view feature; a target prediction module, configured to perform feature fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a fourth bird's-eye view feature, and perform target prediction based on the fourth bird's-eye view feature; Wherein, the third bird's-eye view feature determination module may include: A second image viewing angle feature obtaining submodule is used to perform feature enhancement processing on the first image viewing angle feature in the image viewing angle space to obtain a second image viewing angle feature; A submodule for obtaining a third feature according to the second feature, used for obtaining a third bird's-eye view feature according to the second image view feature; Wherein, obtaining the third characteristic submodule according to the second characteristic may include: A second feature depth information determining unit, used to determine the depth information corresponding to the second image viewing angle feature; The second feature perspective transformation unit is used to transform the perspective of the second image perspective feature according to the depth information corresponding to the second image perspective feature to obtain a third bird's-eye view perspective feature.
8. An electronic device, It is characterized in that The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the image processing method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 6 is implemented.