A processing method and device for image feature encoding based on camera internal and external parameters

By combining the camera's internal and external parameters for deep feature encoding and cascade processing, the problem of cross-camera visual model adaptation is solved, and the camera calibration prior is eliminated, which improves the model's adaptability and multiplexed rate.

CN114937092BActive Publication Date: 2025-05-16SUZHOU QINGZHOU ZHIHANG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210536997.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-05-16
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to adapt across camera vision models, mainly because the visual model relies on the priori of camera calibration parameters, which makes it impossible to effectively handle the viewing angle differences of different cameras.

Method used

By combining the camera's internal and external parameters for depth feature encoding, an image combining depth features is generated, and the depth features are added to the image through cascade, thereby getting rid of the dependence on the camera calibration prior and achieving cross-camera adaptation.

Benefits of technology

This method can reduce the difficulty of model maintenance, improve model reuse rate, reduce the working complexity of the environment perception module, improve work efficiency, and improve model adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937092B_ABST
    Figure CN114937092B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention relates to a processing method and device for encoding image features based on camera internal and external parameters, the method comprising: obtaining a captured image of a first camera as a first original image, and obtaining the position coordinates of the first camera as the first coordinates; initializing a feature map with the same height and width as the first original image but with a channel number of 1 as the first feature map; taking the position point corresponding to the first coordinate as the center point of the bottom edge on the map, selecting a rectangular area along the radial direction of the first camera as the first area map; performing feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; cascading the second feature map with the first original image to obtain a corresponding third feature map. Through the present invention, an image combining the internal and external parameter features of the camera can be obtained, and the image output by the present invention can be sent to the visual model to get rid of the dependence on the camera calibration prior and help the model complete cross-camera adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a processing method and device for image feature encoding based on camera internal and external parameters. Background Art

[0002] In the field of autonomous driving, different cameras have different viewing angles, and visual models (especially 3D-related visual models) rely heavily on the learned priors of camera calibration parameters. Therefore, a visual model often cannot be well adapted across cameras. If the input of the model can be increased by deep features calculated from camera internal and external parameters, it can help the visual model get rid of the camera calibration priors and complete cross-camera adaptation. Here, we propose a method to encode deep features in combination with camera internal and external parameters, and use the deep features together with the image as the input of the visual model. Summary of the invention

[0003] The purpose of the present invention is to address the defects of the prior art and provide a processing method, device, electronic device and computer-readable storage medium for encoding image features based on camera internal and external parameters, encode the depth features of ground points in the image based on the true coordinates of the map according to the camera internal and external parameters, and add the depth features to the image in a cascade manner. Through the present invention, an image combined with depth features can be obtained, and the image output by the present invention can be sent to the visual model to get rid of the dependence on the camera calibration prior and help the model complete cross-camera adaptation.

[0004] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present invention provides a processing method for encoding image features based on camera internal and external parameters, the method comprising:

[0005] An image captured by a first camera is obtained as a first original image, and the position coordinates of the first camera are obtained as first coordinates; the tensor shape of the first original image is H*W*C, where H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is a map coordinate system;

[0006] Initialize a feature map with the same height and width as the first original map but with 1 channel number, record it as the first feature map; the shape of the first feature map is H*W*1;

[0007] On the map, taking the position point corresponding to the first coordinate as the center point of the bottom edge, a rectangular area is selected along the radial direction of the first camera as a first area map; the coordinate system of the first area map is the vehicle coordinate system;

[0008] Perform feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; the shape of the second feature map is H*W*1;

[0009] The second feature map is cascaded with the first original map to obtain a corresponding third feature map; the shape of the third feature map is H*W*(C+1).

[0010] Preferably, the initializing a feature map having the same height and width as the first original map but with 1 channel number as the first feature map specifically includes:

[0011] Constructing a feature map with a shape of H*W*1 as the first feature map; the first feature map includes H*W first pixel points; the first pixel point corresponds to a first pixel coordinate (u1, v1) and a first channel data;

[0012] The first channel data of all the first pixels of the first feature map are uniformly initialized to a preset default channel value.

[0013] Furthermore, the default channel value is 255 or a negative number.

[0014] Preferably, the length and width of the first area map are set according to preset long side parameters and wide side parameters;

[0015] The first area map includes a plurality of first position points; the first position points correspond to a set of first vehicle coordinates (x w ,y w , z w ).

[0016] Preferably, performing feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain the corresponding second feature map specifically includes:

[0017] The first vehicle coordinate z on the first area map w The first position point whose value is 0 is recorded as the corresponding first ground point P w ;

[0018] For the first ground point P on the first area map w Perform dense sampling to obtain a corresponding first sampling ground point set;

[0019] According to the external parameter matrices R and T of the first camera, each of the first ground points P in the first sampling ground point set w The first ego-vehicle coordinate (x w ,y w , z w ) to convert the camera coordinates to obtain the corresponding first camera coordinates (x c ,y c , z c ),

[0020] According to the intrinsic parameter matrix K of the first camera, each of the first ground points P in the first sampling ground point set w The first camera coordinates (x c ,y c , z c ) to obtain the corresponding second pixel coordinates (u2, v2), z c [u2,v2,1] T =K[x c ,y c ,z c ,1] T ;

[0021] Each first pixel point of the first feature map is traversed; during the traversal, the first pixel point currently traversed is recorded as the current pixel point, and the first pixel coordinate (u1, v1) of the current pixel point is recorded as the current pixel coordinate, and the first ground point P whose second pixel coordinate (u2, v2) in the first sampling ground point set matches the current pixel coordinate w as the current matching ground point and use the first camera coordinate z of the current matching ground point c The first channel data of the current pixel is set; wherein the first feature map includes H*W first pixel points; the first pixel point corresponds to a first pixel coordinate (u1, v1) and a first channel data;

[0022] Normalization and uniform quantization are performed on all the first channel data that are not preset default channel values ​​in the first feature map to obtain the corresponding second feature map.

[0023] Furthermore, the normalizing and uniformly quantizing all the first channel data that are not preset default channel values ​​in the first feature map to obtain the corresponding second feature map specifically includes:

[0024] According to the preset first parameter Z max Normalize each of the first channel data that is not a preset default channel value in the first feature map, specifically: if the first channel data is less than or equal to the first parameter Z max , then divide the first channel data by the first parameter Z max as the normalized first channel data; if the first channel data is greater than the first parameter Z max , then the first channel data is set to 1; the value range of the normalized first channel data is [0, 1];

[0025] According to the preset value range [0, 255], uniform quantization processing is performed on each normalized first channel data, specifically: the product of multiplying the first channel data by 255 is rounded, and the rounded result is used as the first channel data after uniform quantization processing; the value range of the first channel data after uniform quantization processing is [0, 255].

[0026] A second aspect of an embodiment of the present invention provides a device for implementing the processing method for image feature encoding based on camera internal and external parameters as described in the first aspect, the device comprising: an acquisition module, a first processing module, a second processing module, a third processing module and a fourth processing module;

[0027] The acquisition module is used to acquire an image captured by a first camera as a first original image, and acquire the position coordinates of the first camera as first coordinates; the tensor shape of the first original image is H*W*C, where H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is a map coordinate system;

[0028] The first processing module is used to initialize a feature map with the same height and width as the first original map but with 1 channel number, which is recorded as the first feature map; the shape of the first feature map is H*W*1;

[0029] The second processing module is used to select a rectangular area on the map along the radial direction of the first camera with the position point corresponding to the first coordinate as the base center point as the first area map; the coordinate system of the first area map is the vehicle coordinate system;

[0030] The third processing module is used to perform feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; the shape of the second feature map is H*W*1;

[0031] The fourth processing module is used to cascade the second feature map and the first original map to obtain a corresponding third feature map; the shape of the third feature map is H*W*(C+1).

[0032] A third aspect of an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0033] The processor is used to be coupled to the memory, read and execute instructions in the memory, so as to implement the method steps described in the first aspect above;

[0034] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

[0035] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions. When the computer instructions are executed by a computer, the computer executes the instructions of the method described in the first aspect above.

[0036] The embodiments of the present invention provide a processing method, device, electronic device and computer-readable storage medium for encoding image features based on camera internal and external parameters, and perform depth feature encoding on ground points in an image based on the true coordinates of a map according to camera internal and external parameters, and add the depth features to the image in a cascade manner. Through the present invention, an image combined with depth features can be obtained, and the image output by the present invention can be sent to a visual model to get rid of the dependence on camera calibration priors, help the model complete cross-camera adaptation, reduce the complexity of model processing, and improve the adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a processing method for image feature encoding based on camera internal and external parameters provided in Embodiment 1 of the present invention;

[0038] Figure 2 A module structure diagram of a processing device for encoding image features based on camera internal and external parameters provided in Embodiment 2 of the present invention;

[0039] Figure 3 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] Before explaining the method provided by the embodiment of the present invention, a brief introduction is first given to the use of the visual model by the environmental perception module in the unmanned or automatic driving system. One of the functions of the environmental perception module is to identify obstacle targets (cars, motorcycles, bicycles, buildings, people, animals, static obstacles, traffic signs, etc.) in the driving environment of the vehicle and perform multi-target tracking based on the recognition results to analyze the motion trajectory of each obstacle target. When performing target recognition, the environmental perception module will use some 2D or 3D visual models (such as YOLO, YOLOv3, SSD, Faster RCNN, deep CNN, etc.) to locate and classify the camera images. Specifically, the environmental perception module directly inputs the camera images into each visual model for calculation and the model outputs a classified semantic image with target recognition frame information. As explained above, this traditional approach has a flaw. The viewing angles of different cameras vary greatly, and visual models (especially 3D-related visual models) rely heavily on the learned camera calibration parameters. This means that even if the same visual model is used, the model must be trained based on different cameras and the corresponding model version must be configured for each camera. As a result, the environmental perception module also needs to call different versions of the model to process images from different cameras based on the corresponding relationship between the camera and the model version.

[0042] To solve this problem, a first embodiment of the present invention provides a processing method for encoding image features based on internal and external parameters of the camera; the environmental perception module uses this method to encode the depth features of the image based on the internal and external parameters of the camera corresponding to the image before sending the image taken by the camera into the visual model, and then sends the image with depth features into the visual model; correspondingly, the visual model can directly restore the depth data according to the corresponding encoding rules. Then, for the same visual model, there is no need to conduct multiple trainings based on different cameras to generate multiple versions; for practical applications, the environmental perception module does not need to frequently switch model versions during work. In other words, the method of the embodiment of the present invention can not only reduce the difficulty of model maintenance and improve the model reuse rate, but also reduce the work complexity of the environmental perception module and improve work efficiency.

[0043] Figure 1 A schematic diagram of a processing method for encoding image features based on camera internal and external parameters provided in the first embodiment of the present invention is shown in FIG. Figure 1 As shown, this method mainly includes the following steps:

[0044] Step 1, obtaining an image captured by a first camera as a first original image, and obtaining position coordinates of the first camera as first coordinates;

[0045] Among them, the tensor shape of the first original image is H*W*C, H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is the map coordinate system.

[0046] Here, the environment perception module obtains the image captured by the first camera as the first original image. If the first camera is an ordinary camera, the number of channels C of the first original image is usually 3, that is, the pixel values ​​of the three channels of red, green and blue (RGB) in the pixel. If the first camera has other functions or more pixel channels, the number of channels C of the first original image can also be other values, which are not listed here one by one; the first coordinate is the map coordinate of the vehicle, which the environment perception module can obtain from the positioning module or map module of the vehicle.

[0047] Step 2: Initialize a feature map with the same height and width as the first original map but with 1 channel number, which is recorded as the first feature map;

[0048] Here, the first feature map is actually the feature map of the camera's internal and external parameter encoding, and the current step is to initialize it;

[0049] Specifically comprising: step 21, constructing a feature map with a shape of H*W*1 as a first feature map;

[0050] The shape of the first feature map is H*W*1; the first feature map includes H*W first pixel points; the first pixel point corresponds to a first pixel coordinate (u1, v1) and a first channel data;

[0051] Here, because the camera internal and external parameter encoding feature map, that is, the first feature map, will be added to the first original map later, the height and width of the two must be consistent during initialization; the position coordinates of each pixel point in the first feature map are calibrated based on the pixel coordinate system, that is, the first pixel coordinates (u1, v1);

[0052] Step 22, uniformly initializing the first channel data of all first pixel points of the first feature map to a preset default channel value;

[0053] The default channel value is 255 or a negative number.

[0054] Step 3, taking the position point corresponding to the first coordinate on the map as the center point of the bottom edge, and selecting a rectangular area along the radial direction of the first camera as the first area map;

[0055] The coordinate system of the first area map is the vehicle coordinate system; the length and width of the first area map are set according to the preset long side parameter and wide side parameter; the first area map includes a plurality of first position points; the first position point corresponds to a set of first vehicle coordinates (x w ,y w, z w );

[0056] Specifically, the method includes: firstly, taking the position point corresponding to the first coordinate on the map as the center point of the bottom edge, selecting a rectangular area along the radial direction of the first camera as the map of the area to be extracted, and performing self-vehicle coordinate conversion on the map coordinates of each position point in the map of the area to be extracted, and finally taking the map of the area to be extracted after the coordinate conversion as the first area map.

[0057] Here, the environmental perception module actually constructs a map around the vehicle based on the vehicle positioning to obtain a first area map, and ensures that the first area map has a spatial intersection with the first original map; when constructing the first area map, it is specifically constructed based on the map output by the map module.

[0058] Step 4, performing feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map;

[0059] Among them, the shape of the second feature map is H*W*1;

[0060] Specifically, step 41 includes: w The first position point where the value is 0 is recorded as the corresponding first ground point P w ;

[0061] Here, because the position coordinates of the first area map follow the characteristics of the vehicle coordinate system, the first vehicle coordinate z in the first area map is w Points with a value of 0 are considered ground points;

[0062] Step 42: For the first ground point P on the first area map w Perform dense sampling to obtain a corresponding first sampling ground point set;

[0063] Here, the embodiment of the present invention provides a plurality of dense sampling modes for the first ground point P on the first area map. w Dense sampling is performed, one of the sampling modes is full sampling, that is, the first ground point P on the first area map w When performing dense sampling, all the first ground points P on the first area map are w All of them are included in the first sampling ground point set; another sampling mode is the equal interval sampling mode, that is, the first ground point P on the first area map w When dense sampling is performed, the first ground point P on each row or column on the first area map is sampled according to the preset row sampling interval or column sampling interval. w Sampling is performed to obtain the corresponding first sampling ground point set; another sampling mode is the center point dense sampling mode, that is, the first ground point P on the first area mapw When performing dense sampling, first scan the first ground point P of each row in a row-by-row manner. w Clustering is performed to obtain the corresponding first-row ground point sets, and then the first ground point P at the center of each first-row ground point set is clustered. w The first ground point P is defined as the center point and located at the left and right boundaries. w As the corresponding left and right boundary points, the row spacing from each center point to the corresponding left and right boundary points is equally divided to obtain multiple equally divided intervals, and then the corresponding sampling rate is set according to the distance from each equally divided interval to its corresponding center point (the sampling rate of each equally divided interval is inversely proportional to the distance from the equally divided interval to the corresponding center point), and then the first ground point P in the equally divided interval is sampled according to the sampling rate of each equally divided interval w Sampling is performed to obtain the corresponding first row of ground point sets;

[0064] Step 43: According to the external parameter matrices R and T of the first camera, each first ground point P in the first sampling ground point set is sampled. w The first self-vehicle coordinate (x w ,y w , z w ) to convert the camera coordinates to obtain the corresponding first camera coordinates (x c ,y c , z c );

[0065] in,

[0066] Here, each camera has a pair of external parameter matrices R and T. The definition of the external parameter matrix can be referred to in the public technical literature, and will not be repeated here. The vehicle coordinate system is sometimes also called the environment coordinate system in the field of autonomous driving technology. In the public technical literature, the conversion relationship between the environment coordinate system (the vehicle coordinate system) and the camera coordinate system is Therefore, the external parameter matrix R, T of the first camera and the first ground point P w The first self-vehicle coordinate (x w ,y w , z w ) is known, we can substitute it into the above conversion formula to get the first ground point P w The coordinates in the camera coordinate system are the first camera coordinates (x c ,y c , z c );

[0067] Step 44: According to the intrinsic parameter matrix K of the first camera, each first ground point P in the first sampling ground point set is sampled. w The first camera coordinate (x c ,y c , zc ) to obtain the corresponding second pixel coordinates (u2, v2);

[0068] Among them, z c [u2,v2,1] T =K[x c ,y c ,z c ,1] T ;

[0069] Here, each camera has an intrinsic parameter matrix K. The definition of the intrinsic parameter matrix can be found in the public technical literature, which will not be described in detail here. In the public technical literature, the conversion relationship between the camera coordinate system and the pixel coordinate system is z c [u2,v2,1] T =K[x c ,y c ,z c ,1] T , therefore, the intrinsic parameter matrix K of the first camera and the first ground point P w The first camera coordinate (x c ,y c , z c ) is known, we can substitute it into the above conversion formula to get the first ground point P w The coordinates in the pixel coordinate system are the second pixel coordinates (u2, v2);

[0070] Step 45, traverse each first pixel point of the first feature map; during the traversal, record the first pixel point currently traversed as the current pixel point, and record the first pixel coordinate (u1, v1) of the current pixel point as the current pixel coordinate, and record the first ground point P whose second pixel coordinate (u2, v2) in the first sampling ground point set matches the current pixel coordinate w As the current matching ground point, and use the first camera coordinate z of the current matching ground point c Set the first channel data of the current pixel;

[0071] Here, although the first area map must have a large intersection with the first original map, it cannot be guaranteed that all points on the first area map can be captured in the first original map. In other words, the first ground point P of the first area map w There must be a corresponding pixel point in the first original image, and the pixel coordinates (u, v) of the first original image and the first feature image are completely consistent, that is, the first ground point P of the first area map w There must be corresponding pixels in the first feature map; therefore, the first ground point P that matches each other based on the pixel coordinates wSet the first pixel point; when setting the channel data of the first pixel point, the first ground point P is actually used w The depth information is also the first camera coordinate z c The coordinate value of is used as the channel data of the first pixel;

[0072] Step 46, normalizing and uniformly quantizing all first channel data that are not preset default channel values ​​in the first feature map to obtain a corresponding second feature map;

[0073] Specifically, step 461, according to the preset first parameter Z max Normalizing each first channel data that is not a preset default channel value in the first feature map;

[0074] Specifically: if the first channel data is less than or equal to the first parameter Z max , then divide the first channel data by the first parameter Z max The quotient is taken as the normalized first channel data; if the first channel data is greater than the first parameter Z max , then set the first channel data to 1;

[0075] Among them, the value range of the normalized first channel data is [0, 1];

[0076] Here, the first parameter Z max is a preset large value; for example, the first channel data is d1, and the normalized first channel data is d2, then, when d1≤Z max When d2=d1 / Z is obtained by the maximum value normalization method max ; in d>Z max When , we get d2=1;

[0077] Step 462, performing uniform quantization processing on each normalized first channel data according to a preset value range [0, 255];

[0078] Specifically, the product of multiplying the first channel data by 255 is rounded, and the rounded result is used as the first channel data after uniform quantization processing;

[0079] The value range of the first channel data after uniform quantization processing is [0, 255].

[0080] Here, the first channel data after uniform quantization processing is actually a camera internal and external parameter feature encoding that is standardized between [0, 255]; for example, the first channel data after normalization is d2, and the first channel data after uniform quantization processing is d3, then d3=int(d2*255), int() is a rounding function.

[0081] It should be noted that the default channel value of the first channel data initialized in the first feature map mentioned in step 22 is 255 or a negative number; if 255 is pre-selected as the default channel value, then after the second feature map is obtained in step 4, go directly to step 5 for the next step of processing; if a negative number is pre-selected as the default channel value, then after the second feature map is obtained in step 4, all first channel data that are specifically negative numbers must be set to 255. This is mainly to unify data expression.

[0082] Step 5: cascade the second feature map and the first original map to obtain a corresponding third feature map; wherein the shape of the third feature map is H*W*(C+1).

[0083] Here, the embodiment of the present invention adds the channel data in the second feature map to the camera image, that is, the first original map, by means of tensor cascade to obtain the third feature map; therefore, the length and width of the third feature map are consistent with those of the first original map, but the dimension of the data channel is increased by 1. The third feature map here is actually a fusion feature map that adds depth features to the ground points of the first original map.

[0084] Subsequently, the environmental perception module sends the third feature map to the corresponding visual model for analysis. Because the third feature map incorporates the ground point depth features, the embodiment of the present invention also needs to modify the parsing process of the visual model so that it no longer obtains the ground point depth features by learning the camera's internal and external parameter matrices K, R, and T, but directly extracts the ground point depth features from the second feature map. The specific modification method is to provide a ground point depth feature decoding process for the visual model, that is, the method of the embodiment of the present invention also includes: decomposing the second feature map from the third feature map; and recording the pixel points in the second feature map whose channel data is not 255 as ground points; and converting the channel data of each ground point, the converted channel data = int(channel data before conversion * Z max / 255), int() is the rounding function.

[0085] In this way, the visual model based on the embodiment of the present invention can directly obtain the depth information of the ground points in the image from the input third feature map, and use the obtained ground point depth information as a reference to calculate the depth information of the remaining non-ground points; it is no longer necessary to use the internal and external parameter matrices K, R, and T of different cameras obtained through training and learning to perform depth estimation on each point in the image. Therefore, for the same visual model, it is no longer necessary to conduct multiple trainings based on different cameras to generate multiple versions; for practical applications, the environmental perception module does not need to frequently switch model versions during operation. In other words, the method of the embodiment of the present invention can reduce the difficulty of model maintenance and improve the model reuse rate, and can also reduce the working complexity of the environmental perception module and improve work efficiency.

[0086] Figure 2 The module structure diagram of a processing device for image feature encoding based on camera internal and external parameters provided in the second embodiment of the present invention, the device is a terminal device or server that implements the aforementioned method embodiment, and can also be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiment, for example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 2 As shown, the device includes: an acquisition module 201 , a first processing module 202 , a second processing module 203 , a third processing module 204 and a fourth processing module 205 .

[0087] The acquisition module 201 is used to acquire the image captured by the first camera as the first original image, and acquire the position coordinates of the first camera as the first coordinates; the tensor shape of the first original image is H*W*C, H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is the map coordinate system.

[0088] The first processing module 202 is used to initialize a feature map with the same height and width as the first original map but with 1 channel number, which is recorded as the first feature map; the shape of the first feature map is H*W*1.

[0089] The second processing module 203 is used to select a rectangular area along the radial direction of the first camera as a first area map with the position point corresponding to the first coordinate as the base center point on the map; the coordinate system of the first area map is the vehicle coordinate system.

[0090] The third processing module 204 is used to perform feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; the shape of the second feature map is H*W*1.

[0091] The fourth processing module 205 is used to cascade the second feature map and the first original map to obtain a corresponding third feature map; the shape of the third feature map is H*W*(C+1).

[0092] It should be noted that, corresponding to the implementation method of the ground point depth feature decoding process provided for the visual model in the first embodiment of the present invention, the second embodiment of the present invention also provides a ground point depth feature decoding module for the visual model. The ground point depth feature decoding module is used to decompose the second feature map from the third feature map; and record the pixel points whose channel data in the second feature map is not 255 as ground points; and convert the channel data of each ground point, and the converted channel data = int(channel data before conversion * Z max / 255), int() is the rounding function.

[0093] An embodiment of the present invention provides a processing device for encoding image features based on camera internal and external parameters, which can execute the method steps in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here.

[0094] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the acquisition module can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a processing element of the above device. The function of the above-mentioned module is determined. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.

[0095] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0096] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the above method embodiments are generated. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above-mentioned computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) methods. The above-mentioned computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The above-mentioned available medium can be a magnetic medium (such as a floppy disk, a hard disk, a tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0097] Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be the aforementioned terminal device or server, or may be a terminal device or server connected to the aforementioned terminal device or server to implement the method of the embodiment of the present invention. Figure 3 As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the aforementioned method embodiment. Preferably, the electronic device involved in the embodiment of the present invention also includes: a power supply 304, a system bus 305 and a communication port 306. The system bus 305 is used to realize the communication connection between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.

[0098] exist Figure 3The system bus 305 mentioned in the figure can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The term "communication interface" is represented by only one thick line, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM) and may also include non-volatile memory (Non-Volatile Memory), such as at least one disk storage.

[0099] The above-mentioned processor can be a general-purpose processor, including a central processing unit CPU, a network processor (Network Processor, NP), a graphics processing unit (Graphics Processing Unit, GPU), etc.; it can also be a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0100] It should be noted that an embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the method and processing process provided in the above embodiments.

[0101] An embodiment of the present invention further provides a chip for executing instructions, wherein the chip is used to execute the processing steps described in the aforementioned method embodiment.

[0102] The embodiments of the present invention provide a processing method, device, electronic device and computer-readable storage medium for encoding image features based on camera internal and external parameters, and perform depth feature encoding on ground points in an image based on the true coordinates of a map according to camera internal and external parameters, and add the depth features to the image in a cascade manner. Through the present invention, an image combined with depth features can be obtained, and the image output by the present invention can be sent to a visual model to get rid of the dependence on camera calibration priors, help the model complete cross-camera adaptation, reduce the complexity of model processing, and improve the adaptability of the model.

[0103] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0104] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0105] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A processing method for image feature encoding based on camera internal and external parameters, characterized in that: The method comprises: An image captured by a first camera is obtained as a first original image, and the position coordinates of the first camera are obtained as first coordinates; the tensor shape of the first original image is H*W*C, where H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is a map coordinate system; Initialize a feature map with the same height and width as the first original map but with 1 channel number, record it as the first feature map; the shape of the first feature map is H*W*1; On the map, taking the position point corresponding to the first coordinate as the center point of the bottom edge, a rectangular area is selected along the radial direction of the first camera as a first area map; the coordinate system of the first area map is the vehicle coordinate system; Perform feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; the shape of the second feature map is H*W*1; The second feature map is cascaded with the first original map to obtain a corresponding third feature map; the shape of the third feature map is H*W*(C+1); The length and width of the first area map are set according to preset long side parameters and wide side parameters; The first area map includes a plurality of first position points; the first position points correspond to a set of first vehicle coordinates (x w ,y w , z w ); The performing feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain the corresponding second feature map specifically includes: The first vehicle coordinate z on the first area map w The first position point whose value is 0 is recorded as the corresponding first ground point P w ; For the first ground point P on the first area map w Perform dense sampling to obtain a corresponding first sampling ground point set; According to the external parameter matrices R and T of the first camera, each of the first ground points P in the first sampling ground point set w The first ego-vehicle coordinate (x w ,y w , z w ) to convert the camera coordinates to obtain the corresponding first camera coordinates (x c ,y c , z c ), ; According to the intrinsic parameter matrix K of the first camera, each of the first ground points P in the first sampling ground point set w The first camera coordinates (x c ,y c , z c ) to obtain the corresponding second pixel coordinates (u2, v2), ; The first feature map is subjected to ground point channel data setting and all channel data are normalized and uniformly quantized to obtain the second feature map.

2. The method for encoding image features based on camera internal and external parameters according to claim 1, characterized in that: Initializing a feature map having the same height and width as the first original map but with 1 channel number as the first feature map specifically includes: Construct a feature map with a shape of H*W*1 as the first feature map; the first feature map includes H*W first pixel points; the first pixel point corresponds to a first pixel coordinate (u1, v1) and a first channel data; The first channel data of all the first pixels of the first feature map are uniformly initialized to a preset default channel value.

3. The method for encoding image features based on camera internal and external parameters according to claim 2, characterized in that: The default channel value is 255 or a negative number.

4. The method for encoding image features based on camera internal and external parameters according to claim 1, characterized in that: The step of setting ground point channel data for the first feature map and normalizing and uniformly quantizing all channel data to obtain the second feature map specifically includes: traversing each first pixel point of the first feature map; during the traversal, recording the first pixel point currently traversed as the current pixel point, and recording the first pixel coordinate (u1, v1) of the current pixel point as the current pixel coordinate, and recording the first ground point P whose second pixel coordinate (u2, v2) in the first sampling ground point set matches the current pixel coordinate w as the current matching ground point and use the first camera coordinate z of the current matching ground point c The first channel data of the current pixel is set; wherein the first feature map includes H*W first pixel points; the first pixel point corresponds to a first pixel coordinate (u1, v1) and a first channel data; Normalization and uniform quantization are performed on all the first channel data that are not preset default channel values ​​in the first feature map to obtain the corresponding second feature map.

5. The method for encoding image features based on camera internal and external parameters according to claim 4, characterized in that: The step of normalizing and uniformly quantizing all the first channel data that are not preset default channel values ​​in the first feature map to obtain the corresponding second feature map specifically includes: According to the preset first parameter Z max Normalize each of the first channel data that is not a preset default channel value in the first feature map, specifically: if the first channel data is less than or equal to the first parameter Z max , then divide the first channel data by the first parameter Z max as the normalized first channel data; if the first channel data is greater than the first parameter Z max , then the first channel data is set to 1; the value range of the normalized first channel data is [0, 1]; According to the preset value range [0, 255], uniform quantization processing is performed on each normalized first channel data, specifically: the product of multiplying the first channel data by 255 is rounded, and the rounded result is used as the first channel data after uniform quantization processing; the value range of the first channel data after uniform quantization processing is [0, 255].

6. A device for implementing the processing method for image feature encoding based on camera internal and external parameters as described in any one of claims 1 to 5, characterized in that: The device comprises: an acquisition module, a first processing module, a second processing module, a third processing module and a fourth processing module; The acquisition module is used to acquire an image captured by a first camera as a first original image, and acquire the position coordinates of the first camera as first coordinates; the tensor shape of the first original image is H*W*C, where H is the image height, W is the image width, and C is the number of channels; the coordinate system of the first coordinate is a map coordinate system; The first processing module is used to initialize a feature map with the same height and width as the first original map but with 1 channel number, which is recorded as the first feature map; the shape of the first feature map is H*W*1; The second processing module is used to select a rectangular area on the map along the radial direction of the first camera with the position point corresponding to the first coordinate as the base center point as the first area map; the coordinate system of the first area map is the vehicle coordinate system; The third processing module is used to perform feature encoding processing on the first feature map according to the internal and external parameters of the first camera and the first area map to obtain a corresponding second feature map; the shape of the second feature map is H*W*1; The fourth processing module is used to cascade the second feature map and the first original map to obtain a corresponding third feature map; the shape of the third feature map is H*W*(C+1).

7. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 5; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which, when executed by a computer, enable the computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • A monocular SLAM initialization method and system based on a wheel type encoder

    CN109671120A

  • Camera parameter acquisition method and device, equipment and storage medium

    CN111583345A