Image processing method, readable storage medium, program product and vehicle-mounted device
By storing the obstacle model in the vehicle-mounted device and optimizing its feature point representation, the problem of excessive computing power requirements on the vehicle-side platform is solved, and efficient and accurate environmental image generation is achieved.
Patent Information
- Application Number
- CN202510621997.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-05-14
AI Technical Summary
As the vehicle's demand for surrounding environment detection range increases, the size of the BEV space grows quadratically, resulting in excessively high computing power requirements for the vehicle-side platform and difficulty in effectively deploying image processing technology.
The on-board device stores the first spatial data of N first obstacle models. By obtaining features from multiple images, it selects N1 obstacle models that are highly similar to the actual obstacle features, and optimizes and generates an environmental image with the historical obstacle models to reduce computing power requirements.
By optimizing the feature point representation of the obstacle model, the computing power requirements of the on-board equipment are reduced, and the accuracy and speed of generating environmental images are improved.
Smart Images

Figure CN120148009B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, a readable storage medium, a program product, and a vehicle-mounted device. Background Art
[0002] Multi-view image three-dimensional object detection is currently widely used in many fields and application scenarios, such as autonomous driving scenarios, intelligent transportation scenarios, industrial automation, as well as virtual reality and augmented reality scenarios.
[0003] For example, in assisted driving technology, bird's eye view (BEV) technology can provide multi-perspective image features of the vehicle's surrounding environment projected into the BEV space, and then estimate the global perspective of the position, size and attributes of dynamic obstacles, thereby facilitating path planning and decision-making for users.
[0004] However, as vehicles' demands for surrounding environment detection range grow, the size of the BEV space grows quadratically, and the amount of data that needs to be processed to generate the BEV space increases, thereby increasing the computing power requirements of the vehicle-side platform, making it difficult to deploy the corresponding technology on the vehicle-side platform. Summary of the Invention
[0005] Embodiments of the present application provide an image processing method, a readable storage medium, a program product, and an in-vehicle device.
[0006] In a first aspect, an embodiment of the present application provides an image processing method for an on-board device on a vehicle. The on-board device stores first spatial data of N first obstacle models, where the first spatial data of each first obstacle model includes first position data of M feature points of the first obstacle model. The method includes: the on-board device, in response to a request to generate a first environmental image of the vehicle, obtains multiple images, where the multiple images are images from different perspectives captured by multiple sensors of the vehicle, and the multiple images include images of actual obstacles around the vehicle. The on-board device samples features of the multiple images based on the first position data of each feature point in the first spatial data to obtain first features of the N first obstacle models. The on-board device selects N1 first obstacle models from the N first obstacle models based on the multiple images, where the similarity between the first features of the N1 first obstacle models and the features of the actual obstacles is greater than the similarity between the first features of the remaining first obstacle models and the features of the actual obstacles. The on-board device optimizes the first spatial data of N1 first obstacle models and the second spatial data of N2 historical obstacle models to obtain third spatial data of N1 first obstacle models and N2 historical obstacle models, where N = N1 + N2, and each piece of third spatial data includes second position data of M feature points of the corresponding obstacle model. Furthermore, the similarity between second features obtained by sampling features of multiple images based on the second position data of each feature point in the third spatial data and features of actual obstacles is greater than a set threshold. Thus, the on-board device can generate a first environment image based on the second features.
[0007] In some embodiments of the present application, the vehicle-mounted device stores first spatial data for N first obstacle models, including position data for M feature points of the first obstacle model. During the process of generating the first environment image, the vehicle-mounted device can optimize the spatial data of the first obstacle model using spatial data of actual obstacles from multiple captured images, enabling the first obstacle model to replace the actual obstacles. The vehicle-mounted device then generates the first environment image based on the M feature points of the first obstacle model, thereby reducing the computing power required by the vehicle-mounted device.
[0008] In some embodiments of the present application, since the vehicle needs to continuously generate an environmental image while driving, the on-board device can also use N2 historically optimized historical obstacle models to participate in the calculation process of generating the current first environmental image. It is understood that the on-board device can estimate the position of the historical obstacle model at the current moment based on the historical obstacle model relative to the vehicle's speed, thereby participating in the generation of the first environmental image at the current moment. Since the size of the historical obstacle model is already optimized, incorporating the historical obstacle model into the process of generating the first environmental image can improve the accuracy and speed of generating the first environmental image.
[0009] In a possible implementation of the first aspect, the M feature points include M points on three coordinate axes that are orthogonal to each other in the coordinate space of the first obstacle model.
[0010] The first position data of the M feature points are used to represent the initial spatial size, position, and rotation angle of the first obstacle model.
[0011] In some embodiments of the present application, the positions of the M feature points can represent information such as the spatial size, position, and rotation angle of the first obstacle model. Optimizing the first spatial data of the first obstacle model actually involves adjusting the positions of the M feature points of the first obstacle model.
[0012] For example, in some embodiments of the present application, the value of M may be 13. For example, the coordinate space of the first obstacle model includes five feature points on each of the three coordinate axes, three of which pass through the coordinate origin, which may be the geometric center of the first obstacle. In other embodiments, M may also take other values. It will be appreciated that a smaller value of M reduces the computing power required to generate the first environment image based on the first obstacle model.
[0013] In a possible implementation of the first aspect, selecting N1 first obstacle models from N first obstacle models based on the multiple images includes:
[0014] The vehicle-mounted device selects, according to the preset arrangement positions of the N first obstacle models, a first obstacle model having a high similarity between the corresponding first feature and the feature of the actual obstacle from between every two adjacent first obstacle models as the obstacle model among the N1 first obstacle models.
[0015] It can be understood that by selecting N1 first obstacle models with higher similarity, the on-board device does not need to sort the N first obstacle models according to similarity, thereby further reducing the computing power requirement of the on-board device.
[0016] In one possible implementation of the first aspect, the on-board device includes third features of N third obstacle models corresponding to a historically generated second environment image, and fourth spatial data of the N third obstacle models, wherein the similarity between the third features and image features of actual obstacles corresponding to the second environment image is greater than a set threshold, and the third features are obtained by sampling the fourth spatial data of the N third obstacle models from image features of actual obstacles corresponding to the second environment image. The second spatial data of the N2 historical obstacle models is determined by selecting, from the third features of the N third obstacle models, the N2 third features having the highest similarity to the features of the image of the actual obstacle corresponding to the second environment image; and using the fourth spatial data of the third obstacle models corresponding to the N2 third features as the second spatial data of the N2 historical obstacle models.
[0017] In some embodiments of the present application, when the on-board device acquires N2 historical obstacle models, it can select, from the N third obstacle models corresponding to the environment image generated in the previous frame, the N2 third obstacle models corresponding to the N2 third features with the highest similarity to the real obstacle in the previous frame as the N2 historical obstacle models. Similarly, the fourth spatial data corresponding to the N2 historical obstacle models can be used as the second spatial data of the N2 historical obstacle models. The N2 third features with the highest similarity to the real obstacle can be selected, for example, by sorting the N third features by their similarity to the real obstacle and selecting the top N2 third features with the highest similarity.
[0018] In one possible implementation of the first aspect, optimizing the first spatial data of N1 first obstacle models and the second spatial data of N2 historical obstacle models to obtain third spatial data of the N1 first obstacle models and the N2 historical obstacle models includes adjusting positions of M feature points of the N1 first obstacle models and the N2 historical obstacle models at least once based on spatial data of actual obstacles in multiple images, thereby obtaining the third spatial data.
[0019] In some embodiments of the present application, the vehicle-mounted device may perform multiple optimizations on the first spatial data of the determined N1 first obstacle models and the N2 spatial data of the N2 historical obstacle models based on the spatial data of the multiple images, thereby obtaining third spatial data for the N1 first obstacle models and the N2 historical obstacle models. It is understood that the optimization process is, for example, a process of reducing errors in size, position, and rotation angle between the N1 first obstacle models and the N2 historical obstacle models and the actual obstacles in the multiple images.
[0020] In a possible implementation of the first aspect, the N2 historical obstacle models further include a corresponding third feature; and generating the first environment image based on the second feature includes: performing at least one self-attention adjustment on the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models to obtain a fourth feature; and using a fifth feature formed by fusing the fourth feature with the second feature as a feature for generating the first environment image; wherein a degree of similarity between the fifth feature and the features of the multiple images is higher than a degree of similarity between the second feature and the features of the multiple images.
[0021] In some embodiments of the present application, the vehicle-mounted device can also optimize the feature data corresponding to the N1 first obstacle models and the N2 historical obstacle models through a self-attention mechanism, thereby improving the accuracy of generating the first environment image based on the N1 first obstacle models and the N2 historical obstacle models.
[0022] In one possible implementation of the first aspect, the multiple images include distorted images; the onboard device stores distortion errors resulting from mapping the spatial size, position, and rotation angle of the first obstacle model from the undistorted image to the distorted image; and the third spatial data includes distortion-error-adjusted first spatial data of the first obstacle model projected into the coordinate space of the multiple images.
[0023] In some embodiments of the present application, the images captured by the cameras on the vehicle may be distorted images. Therefore, in some embodiments, the on-board device is required to perform distortion processing on the first obstacle model or the historical obstacle model mapped to the coordinate space of multiple images. However, the process of calculating the distortion offset through the distortion formula will increase the computing power required by the on-board device. Therefore, by storing the distortion errors of each camera in the on-board device, the on-board device only needs to query to obtain the distortion error, so as to adjust the positions of the N1 first obstacle models and the N2 historical obstacle models mapped to the coordinate space of multiple images, thereby further reducing the computing power required by the on-board device.
[0024] In a second aspect, the present application provides an in-vehicle device comprising: a memory for storing instructions; and at least one processor for executing the instructions to cause the device to implement the image processing method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the second aspect can be referenced to the beneficial effects of the image processing method provided in any embodiment of the first aspect and will not be further elaborated here.
[0025] In a third aspect, the present application provides a computer-readable storage medium storing instructions that, when executed by a device, cause a computer to implement the image processing method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achieved in the third aspect can be referenced to the beneficial effects of the image processing method provided in any embodiment of the first aspect and are not further elaborated here.
[0026] In a fourth aspect, the present application provides a computer program product that, when executed on a device, causes the device to implement the image processing method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achieved in the fourth aspect can be referenced to the beneficial effects of the image processing method provided in any embodiment of the first aspect and are not further elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1A A flow chart for generating a BEV space is shown;
[0028] Figure 1B A schematic diagram of a vehicle during driving is shown;
[0029] Figure 2 According to some embodiments of the present application, a schematic diagram of a feature point is shown;
[0030] Figure 3 According to an embodiment of the present application, a flowchart of an implementation of generating a vehicle environment image is shown;
[0031] Figure 4A According to some embodiments of the present application, a schematic diagram of obtaining distortion error is shown;
[0032] Figure 4B According to some embodiments of the present application, a schematic diagram of obtaining first features of N first obstacle models is shown;
[0033] Figure 5 A schematic diagram of a process for processing distortion errors is shown according to some embodiments of the present application;
[0034] Figure 6 According to some embodiments of the present application, a process of selecting N1 first obstacle models is shown;
[0035] Figure 7 According to some embodiments of the present application, a schematic diagram of a vehicle-mounted device generating a first environment image is shown;
[0036] Figure 8A A schematic diagram of inter-frame interaction is shown according to some embodiments of the present application;
[0037] Figure 8BA schematic diagram of inter-frame interaction is shown according to some embodiments of the present application;
[0038] Figure 8C According to some embodiments of the present application, a schematic diagram of an optimization process of a self-attention module is shown;
[0039] Figure 9 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device 100 is shown. DETAILED DESCRIPTION
[0040] The illustrative embodiments of the present application include, but are not limited to, an image processing method, a readable storage medium, a program product, and an in-vehicle device.
[0041] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0042] To facilitate understanding, some of the terms and related technologies involved in this application are explained below.
[0043] BEV Technology:
[0044] BEV technology uses a neural network to convert image information from image space to BEV space. BEV technology can simplify complex three-dimensional environments into two-dimensional images. For example, in the field of autonomous driving, BEV technology can generate a panoramic view from above the vehicle based on the spatial information around the vehicle, thus displaying the vehicle's surroundings in all directions, including the front, rear, left, right, and top. This enables the autonomous driving system to better understand the surrounding environment and improve the accuracy of perception and decision-making.
[0045] Next, we will introduce the process of generating BEV space using the vehicle's onboard equipment.
[0046] For example, Figure 1A A flow chart for generating a BEV space is shown.
[0047] It is understood that the following processes can be executed by an onboard device. The onboard device in the embodiments of this application can also be referred to as an onboard terminal. The onboard terminal can be a mobile phone, a vehicle computer, a terminal in a self-driving car, a wireless terminal in transportation safety, a terminal in a smart city, etc. The following description will take the vehicle computer as an example. However, it is understood that the technical solution described in this application is applicable to the various onboard electronic devices for three-dimensional target detection mentioned above, and is not limited to vehicle computers.
[0048] S101: Acquire sensor data collected by multiple sensors.
[0049] In some embodiments of the present application, a vehicle computer deployed on the vehicle 10 can obtain sensor data collected by multiple sensors.
[0050] For example, Figure 1B A schematic diagram of a vehicle during driving is shown.
[0051] It is understood that the vehicle 10 is generally equipped with multiple sensors, such as cameras (including front-view, side-view, and rear-view cameras, etc.), radars, laser radars, etc. Figure 1B As shown, while the vehicle 10 is driving, multiple sensor data can be collected at the same time through a timestamp synchronization mechanism or a hardware synchronization mechanism. The sensor data may be, for example, image data from different perspectives of the vehicle 10 collected by a camera and / or point cloud data of various obstacles collected by a radar.
[0052] For example, the vehicle 10 can collect environmental data near the vehicle 10 through sensors, referring to Figure 1B The environmental data collected by the sensors may include, for example, sensor data of the first vehicle 01, pedestrian 02, first building 03, second vehicle 04, and second building 05. The vehicle computer on the vehicle 10 may obtain environmental data from various perspectives collected by the sensors on the vehicle 10 at the same time.
[0053] S102: Preprocess the sensor data.
[0054] For example, after acquiring sensor data, the onboard computer on vehicle 10 can pre-process the sensor data. For example, during image processing, the images captured by the camera can be subjected to denoising and distortion correction. During point cloud processing, the point cloud data collected by the lidar and / or radar can be filtered, denoised, and segmented. The data from different sensors can then be aligned to the same point in time to ensure data consistency.
[0055] S103: Perform coordinate conversion on the sensor data.
[0056] For example, the vehicle computer can also perform coordinate conversion on sensor data collected by the sensor, thereby converting the data collected by the sensor from the sensor coordinate system to the vehicle coordinate system. For example, the data of the camera, radar, or lidar can be converted to the vehicle coordinate system centered on the vehicle 10.
[0057] Then, the sensor data converted to the vehicle coordinate system is transferred from the vehicle coordinate system to the BEV coordinate system. For example, the sensor data in the vehicle coordinate system is converted into a bird's-eye view coordinate system (top view).
[0058] Data collected by cameras can be transformed from a perspective to a bird's-eye view through inverse perspective mapping (IPM). Radar and lidar data can be projected to map 3D point cloud data to a 2D bird's-eye view.
[0059] It can be understood that when processing sensor data, the vehicle computer on the vehicle 10 can only process sensor data within a preset range of the vehicle 10. The preset range can be, for example, a circular range with a radius R centered on the vehicle 10.
[0060] It can be understood that in other embodiments, the range of the vehicle 10 sensor collecting data can also be other shapes, and the embodiments of the present application do not limit the shape of the range of the vehicle 10 sensor collecting data.
[0061] S104: Fusing sensor data.
[0062] For example, during the sensor data fusion process, the vehicle computer can perform feature extraction on the sensor data. For example, it can extract features such as lane lines and obstacles from camera images, or extract information such as object position and speed from radar and lidar point clouds. The vehicle computer then performs data fusion to integrate the features of different sensors.
[0063] S105: Generate a BEV map based on the fused sensor data.
[0064] For example, in some embodiments, the vehicle computer can map the fused data into a two-dimensional grid map, where each grid represents a certain range of space (e.g., 0.1m×0.1m). In some embodiments, the vehicle computer can also add semantic information (e.g., marking lane lines, obstacles, pedestrians, vehicles, and other semantic information in the BEV map) and dynamically update (e.g., dynamically adjust the BEV map based on real-time updates of sensor data). For example, the BEV map generated by the vehicle computer can also refer to Figure 1B view.
[0065] It will be appreciated that during the BEV map generation process, only sensor data within a radius R of vehicle 10 is processed. Therefore, if the detection range of vehicle 10 needs to be increased (R becomes larger), more sensor data within that space needs to be processed, thereby increasing the computing power required by the vehicle 10's computer. For example, if the detection range of vehicle 10 increases, the amount of data processed during the above steps S102 to S105 will increase, thereby increasing the computing power required by the vehicle 10's computer.
[0066] For example, in step S104, to increase the detection range of vehicle 10, the vehicle computer needs to sample sensor data more densely to obtain more accurate sensor data features. Therefore, the vehicle computer requires higher sensor computing power. In step S105, because the vehicle computer extracts more features from the sensor data, the vehicle computer also requires higher computing power to project the features of the fused sensor data into the BEV space.
[0067] As mentioned earlier, as vehicles' demands for detecting the surrounding environment grow, they need to process more data during target detection in three-dimensional space, resulting in higher computing power requirements for vehicle computers, making it difficult to deploy corresponding technologies on the vehicle-side platform.
[0068] To address the high computing power required by vehicle computers to generate three-dimensional spatial data, this application proposes an image processing method. The vehicle-mounted device stores first spatial data of N first obstacle models, each of which includes first position data of M feature points of the first obstacle model. The method includes:
[0069] In response to a request to generate a first environment image of the vehicle, the vehicle-mounted device obtains multiple images of different perspectives collected by multiple sensors of the vehicle, where the multiple images include images of actual obstacles around the vehicle.
[0070] The on-board device projects the first spatial data of the N first obstacle models into the coordinate space of the multiple images and samples features of the multiple images based on the first spatial data to obtain first features of the N first obstacle models. Based on the multiple images, the on-board device selects N1 first obstacle models from the N first obstacle models, where the similarity between the first features of the N1 first obstacle models and features of the actual obstacle is greater than the similarity between the first features of the remaining first obstacle models and features of the actual obstacle.
[0071] The on-board device optimizes the first spatial data of N1 first obstacle models and the second spatial data of N2 historical obstacle models to obtain third spatial data of N1 first obstacle models and N2 historical obstacle models, where N = N1 + N2, and each piece of third spatial data includes second position data of M feature points of the corresponding obstacle model, and a similarity between a second feature obtained by sampling features of multiple images based on the second position data of each feature point in the third spatial data and a feature of the actual obstacle is greater than a set threshold. The on-board device generates a first environment image based on the second feature.
[0072] With this solution, the onboard device doesn't need to generate the first environmental image based on features corresponding to a large amount of sensor data. Instead, it generates the first environmental image based on M optimized feature points from N1 preset first obstacle models and M optimized feature points from N2 historical obstacle models, collected from multiple images using second features. Because each obstacle model (including the first obstacle model and the historical obstacle model) can represent the complete obstacle model's spatial information (such as the obstacle model's spatial size, position, and rotation angle) with a relatively small number of feature points, the onboard device doesn't require significant computing power to acquire the second features, thereby reducing the computing power required.
[0073] In some embodiments of the present application, the multiple first obstacle models stored in the onboard device may include common objects encountered during traffic, such as buildings, people, animals, traffic lights, vehicles (including cars, motorcycles, bicycles, etc.), trees, rivers, etc. The first spatial data of the first obstacle may be the first positions of M feature points of the first obstacle model. The M feature points of the first obstacle model may be, for example, M points on three mutually orthogonal coordinate axes in the coordinate space of the first obstacle model. The first position data of the M feature points represents the initial spatial dimensions (e.g., length, width, and height), position, and rotation angle of the first obstacle model. The spatial dimensions of the first obstacle model may correspond to the average dimensions of common objects encountered during traffic. For example, the first obstacle model may be a human. The average height of an adult male is 1.75 meters, and the average height of an adult female is 1.62 meters. The first obstacle model may be a small car. The length of a small car is approximately 4.0 to 4.5 meters, the width is approximately 1.7 to 1.8 meters, and the height is approximately 1.4 to 1.5 meters.
[0074] For example, Figure 2 According to some embodiments of the present application, a schematic diagram of a feature point is shown.
[0075] like Figure 2As shown, in the coordinate space of the first obstacle model, the first obstacle model can be considered as a rectangular block, and the geometric center of the rectangular block can be the coordinate center of the first obstacle model. In some embodiments of the present application, taking M=13 as an example, 13 feature points can be established based on the three-dimensional dimensions of the first obstacle model. These 13 feature points can include five feature points equally spaced along the length, width, and height axes of the first obstacle model. Of the five feature points along each coordinate axis, the distance between the two feature points at the ends represents the dimension of the first obstacle model in that coordinate direction. It can be understood that because the feature point at the geometric center of the first obstacle passes through three coordinate axes, the feature point at the geometric center of the first obstacle is calculated twice, resulting in a total of 13 feature points for the first obstacle model.
[0076] It is understood that in other embodiments, the feature points of the first obstacle model may be of another number, for example, M may be 7, 19, etc. Alternatively, the feature points of the first obstacle model may be distributed in other ways, as long as spatial data such as the size and rotation angle of the first obstacle model can be represented. The embodiments of the present application do not limit the number and location of the feature points of the first obstacle model.
[0077] The following describes a process in which the vehicle-mounted terminal generates a first environment image based on data collected by sensors on the vehicle in an embodiment of the present application.
[0078] Figure 3 According to an embodiment of the present application, a flowchart for generating a vehicle environment image is shown.
[0079] For example, in some embodiments of the present application, the in-vehicle device stores first spatial data for N first obstacles. The first spatial data for each first obstacle model includes first position data for M feature points of the first obstacle model. For example, in embodiments of the present application, N is greater than 384. The first position data is used to identify the initial spatial dimensions (e.g., length, width, and height), position, and rotation angle of the first obstacle model. The spatial dimensions of the first obstacle model may correspond to the average dimensions of common objects encountered during traffic. In the initial state, the positions of the N first obstacles may be, for example, evenly distributed, and the rotation angles of the N first obstacles may be preset values, such as 0°.
[0080] It is understood that if N is large, the vehicle-mounted device will take longer to generate the vehicle's environmental data, but the generated vehicle environmental data will be more accurate. Therefore, when setting the specific value of N, the value of N can be determined based on the computing power and accuracy requirements of the vehicle-mounted device. The embodiments of this application do not limit the value of N. For example, N can also be 400, 500, or 600.
[0081] It can be understood that the execution entities of the following processes can all be on-board equipment arranged on the vehicle. In the process of introducing the following processes, there is no limitation on the execution entities of each process.
[0082] like Figure 3 As shown, the process includes:
[0083] S301 , in response to a request to generate a first environment image of a vehicle, obtain a plurality of images.
[0084] For example, in some embodiments of the present application, the multiple images are images from different perspectives captured by multiple sensors of the vehicle, and the multiple images include images of actual obstacles around the vehicle.
[0085] For example, refer to Figure 1B During driving, vehicle 10 can simultaneously capture multiple images using multiple sensors via a timestamp synchronization mechanism or a hardware synchronization mechanism. The multiple images can, for example, be image data from different perspectives of vehicle 10 captured by a camera, and / or point cloud data of various obstacles captured by a radar. It will be appreciated that in some embodiments of the present application, the multiple images can include images of actual obstacles surrounding vehicle 10, such as a first vehicle 01, a pedestrian 02, a building 03, a second vehicle 04, and a building 05.
[0086] After detecting the request to generate the first environment image, the vehicle-mounted device can obtain multiple images from different perspectives collected by multiple sensors on the vehicle.
[0087] S302 : Sampling features of multiple images according to first position data of each feature point in the first spatial data to obtain first features of N first obstacle models.
[0088] In some embodiments of the present application, after acquiring multiple images, the vehicle-mounted device may project the first spatial data of N first obstacle models into the coordinate space of the multiple images, and then collect the features of the first positions of the feature points of the N first obstacle models in the multiple images as the first features. It will be appreciated that the coordinate space of the multiple images can be determined based on the camera on the vehicle, and that different camera positions on the vehicle correspond to different coordinate spaces for the corresponding images.
[0089] In some embodiments of the present application, the images captured by the camera on the vehicle 10 may also include some distorted images. Therefore, in the process of projecting N first obstacle models onto the spatial data of multiple images based on the first spatial data, the effect of image distortion may be adjusted.
[0090] For example, the spatial size, position, and rotation angle of the first obstacle model stored in the vehicle-mounted device are mapped from the undistorted image to the distortion error of the distorted image. After the first spatial data of the first obstacle model is projected into the coordinate space of the distorted image, it needs to be adjusted for the distortion error.
[0091] For example, Figure 4A According to some embodiments of the present application, a schematic diagram of obtaining distortion error is shown.
[0092] like Figure 4A As shown, the vehicle-mounted device can first obtain the pixel coordinates of the undistorted image and then determine the pixel coordinates of the distorted image based on a distortion calculation formula. The distortion calculation formula is primarily based on the physical properties of the camera lens and the imaging principle. The vehicle-mounted device can determine the distortion calculation formula based on the camera lens of each camera on the vehicle.
[0093] The vehicle-mounted device can obtain a distortion offset table by subtracting the pixel coordinates of the undistorted image from the pixel coordinates of the distorted image. The vehicle-mounted device stores the distortion offset table and can then find the distortion error from the table, thereby reducing the calculation process and the computing power required by the vehicle-mounted device.
[0094] It is understandable that in other embodiments, the distortion offset table may also be determined by other devices and then stored in the vehicle-mounted device.
[0095] Figure 4B According to some embodiments of the present application, a schematic diagram of obtaining first features of N first obstacle models is shown.
[0096] like Figure 4B As shown, after obtaining the first spatial data of N first obstacle models, the on-board device can determine the three-dimensional position of the first obstacle model projected into the space of the undistorted multiple images. This three-dimensional position includes the spatial size, position, and rotation angle of the first obstacle model in the spatial data of the multiple images, that is, the position of each feature point of the first obstacle model in the space of the multiple images. The on-board device then queries the distortion offset table to obtain the distortion error. The on-board device can adjust the three-dimensional position of the first obstacle model based on the distortion error to obtain the three-dimensional position of the first obstacle in the space of the distorted multiple images. Based on the three-dimensional position of the first obstacle model in the space of the distorted multiple images, the on-board device then collects features from the distorted multiple images, thereby obtaining the first features of the N first obstacle models.
[0097] It can be understood that by using the distortion error in the distortion offset table, the vehicle-mounted device does not need to calculate the distortion error, thereby reducing the computing power required by the vehicle-mounted device in the process of generating the first environment image, so that the vehicle-mounted device can be better arranged on the vehicle-side platform.
[0098] It will be appreciated that, as described below, the process of projecting the target obstacle's target space data into the coordinate space of multiple images can also be performed by querying the distortion offset map to determine the target obstacle's three-dimensional position within the coordinate space of the multiple images. In other words, when projecting various obstacle models into the coordinate space of multiple images, the corresponding obstacle's three-dimensional position within the distorted image coordinate space can be determined by querying the distortion offset map, thereby reducing the computational effort of the onboard equipment.
[0099] In some embodiments of the present application, in order to increase the computing speed of the vehicle-mounted equipment, some vehicle-mounted equipment only supports low-precision 8-bit integer (int8) data input. It can be understood that since int8 quantization uses fewer bits to represent data, it has significant advantages in storage requirements. This storage space saving can significantly reduce deployment costs. In some embodiments of the present application, the distortion error in the distortion offset table can be stored in the vehicle-mounted equipment using int8 quantization. However, the accuracy of the distortion error quantized by int8 is low. In order to ensure the accuracy of the distortion error, in some embodiments of the application, a high-precision 16-bit integer (int16) quantized distortion error can be represented by two int8 data.
[0100] For example, Figure 5 A schematic diagram of a process for processing distortion errors is shown according to some embodiments of the present application.
[0101] like Figure 5 As shown, for a distortion error (int16), where int16 represents high-precision 16-bit integer data, it can be represented by a front (int8) data and a rear (int8) data, where int8 represents low-precision 16-bit integer data. Among them, front = distortion error (int16) / 2 8 Round down, rear = distortion error (int16) - front × 2 8 -2 7 It can be understood that since the first bit of the data in the int8 format is the sign bit, the value of the data in the int8 format is only the last 7 bits, so 2 needs to be subtracted. 7 During the calculation process, distortion error (int16) = front × 2 8 +rear+27 .
[0102] It can be understood that the data quantized by int16 is a 16-bit binary number. Since the first bit is the sign bit, the decimal value range that int16 can represent is from -32768 to 32767. In other words, the decimal value X that can be represented by the binary number of the distortion error after int16 quantization is between -32768 and 32767. By dividing X by 2 8 The range of the value obtained by rounding down is 0 to 127, which is exactly the value that can be represented by an 8-bit binary number (the highest bit in an 8-bit binary number is the sign bit, so the range of the decimal number represented by 8-bit binary number is -128 to 127), that is, front is in the form of int8.
[0103] However, since front is X divided by 2 8 Therefore, front×2 8 It may also be smaller than X, which can be expressed by rear as X-front×2 8 Part. And X-front×2 8 The maximum value is 255, which is larger than the value that can be represented by an int8, so rear needs to subtract 2. 7 That is, X-front×2 8 -2 7 The range is also -128 to 127, so rear can be represented by an int8. For example, for the integer 255, it is represented by 0000000001111111 in int16, so the decimal number of front is 0, that is, 255 / 256=0.996 is rounded down to 0, and the int8 form of front is 00000000. The decimal number of rear is 255-0×2 8 -2 7 =127. Then rear represented by int8 is 01111111.
[0104] Similarly, for the integer -255, it is represented by int16 as 1000000001111111, the decimal number of the front is -1, that is, -255 / 256=-0.996 is rounded down to -1, and the int8 form of the front is 10000001, so the decimal number of the rear is -255-(-1)×2 8 -2 7 =-127, rear is represented by int8 as 11111111.
[0105] In the above method, two int8 format data can be used to represent an int16 quantized distortion error data, which can not only improve the computing speed of the vehicle-mounted equipment, but also improve the computing accuracy of the vehicle-mounted equipment.
[0106] It is understood that since the number of first obstacle models is N, the actual number of obstacles may be greater or less than N. If the actual number of obstacles is less than N, the first spatial data of redundant first obstacle models (i.e., first obstacle models that are not the closest to any actual obstacle) do not need to be adjusted. The similarity between these redundant first obstacle models and the images of the actual obstacles in the forward-view image is 0, meaning that no actual obstacles correspond to these first obstacle models. If the actual number of obstacles is greater than N, the onboard device may select actual obstacles closest to the vehicle according to the neural network and associate them with the first obstacle models, adjusting the first spatial data of these corresponding first obstacle models to ensure that the image of the obstacle closest to the vehicle in the first environment image is generated first.
[0107] For example, in an embodiment of the present application, the vehicle-mounted device determines actual obstacles in the forward-view image only within a space between 200m and 100m from the vehicle 10. For example, the space in front of the vehicle 10 may be within 150m. The space to the left and right sides, as well as above and below the vehicle 10, may be within 30m to 80m from the vehicle 10, for example, preferably 50m. The space behind the vehicle 10, for example, may be within 80m to 120m from the vehicle 10, for example, preferably 100m.
[0108] S303 : Select N1 first obstacle models from the N first obstacle models according to the multiple images and the first feature.
[0109] For example, in some embodiments of the present application, the vehicle-mounted device may select N1 first obstacle models from N first obstacle models, where the similarity between the first features of these N1 first obstacle models and the features of the actual obstacle is greater than the similarity between the first features of the remaining first obstacle models and the features of the actual obstacle. Where 0 < N1 ≤ N. For example, in the embodiments of the present application, N1 is half of N, that is, when N is 384, N1 is 192. In other embodiments, N1 may also be other values, such as N1 = N / 3, N1 = N / 4, N1 = 200, N = 100, etc.
[0110] In some embodiments, N1 first obstacle models can be selected from N first obstacle models, each with the highest similarity to the corresponding actual obstacle's features. However, this requires sorting the N first obstacle models by similarity, which consumes considerable computing power. Therefore, in other embodiments, the first obstacle model with the highest similarity between its corresponding first feature and the actual obstacle's features can be selected between every two adjacent first obstacle models according to the preset arrangement of the N first obstacle models to serve as the obstacle model among the N1 first obstacle models.
[0111] For example, Figure 6 According to some embodiments of the present application, a process of selecting N1 first obstacle models is shown.
[0112] like Figure 6 As shown, a first obstacle model can be selected from every two adjacent first obstacle models, whose first features have a high degree of similarity to the features of the actual obstacles in the multiple images. It will be appreciated that since the positions of the N first obstacle models are initially arranged, there is no need to queue the N first obstacle models during the selection process, thereby improving the efficiency of the vehicle-mounted device in generating the first environment image. In other embodiments, the N first obstacle models can also be sorted based on their similarity to the features of the actual obstacles in the image, and then the top N1 first obstacle models with the highest similarity can be selected. However, this approach requires sorting the N first obstacle models based on their similarity to the actual obstacles in the image, and the sorting process consumes a high level of computing power of the vehicle-mounted device.
[0113] S304 : Optimize the first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models to obtain third spatial data of the N1 first obstacle models and the N2 historical obstacle models.
[0114] For example, in some embodiments of the present application, N = N1 + N2, meaning that there are always N obstacle models used to generate the first environment image (hereinafter, N target obstacle models represent the obstacle models in the N1 first obstacle models and the N2 historical obstacle models). It will be understood that the third spatial data includes spatial data obtained by multiple optimizations of the first spatial data of the N1 first obstacle models, and spatial data obtained by multiple optimizations of the second spatial data of the N2 historical obstacle models. After multiple optimizations of the first and second spatial data, the positions of the corresponding M feature points will also change. For example, each piece of third spatial data includes the second position data of the M feature points corresponding to the target obstacle model.
[0115] For example, in an embodiment of the present application, the on-board device may also obtain historical obstacle models for reference during the process of generating the first environment data for the current frame. For example, the on-board device may include third features of N third obstacle models corresponding to a historically generated second environment image, as well as fourth spatial data of the N third obstacle models, wherein the similarity between the third features and the image features of the actual obstacles corresponding to the second environment image is greater than a set threshold, and the third features are obtained by sampling the fourth spatial data of the N third obstacle models from the image features of the actual obstacles corresponding to the second environment image.
[0116] The in-vehicle device may select N2 third features having the highest similarity to the features of the image of the actual obstacle corresponding to the second environment image from the third features of the N third obstacle models.
[0117] The on-board device can then use the fourth spatial data of the third obstacle model corresponding to the N2 third features as the second spatial data for the N2 historical obstacle models. Furthermore, the first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models are hereinafter referred to as target spatial data. In other words, after multiple optimizations of the target spatial data of the N target obstacle models, the third spatial data of the target obstacle model can be obtained. This third spatial data includes the second position data corresponding to the M feature points of the target obstacle model.
[0118] In some embodiments of the present application, the second environment image may be the environment image generated by vehicle 10 in the frame preceding the current frame. After determining that each third obstacle model corresponds to the actual obstacle image in the previous frame (i.e., the actual obstacle image corresponding to the second environment image), the on-board device may obtain the velocity of each third obstacle model relative to vehicle 10 and calculate the position of each third obstacle model in the current frame based on the interval between each frame. In other words, the fourth spatial data may also be spatial data adjusted by calculating the velocity of each third obstacle model relative to vehicle 10 and the interval between each third obstacle model and the current frame.
[0119] In some embodiments of the present application, the vehicle-mounted device can, for example, adjust the positions of multiple feature points of N1 first obstacle models and the positions of multiple feature points of N2 historical obstacle models at least once based on the spatial data of the actual obstacle in multiple images, thereby obtaining third spatial data.
[0120] It is understandable that after obtaining the target space data of the target obstacle model, the on-board device can adjust the target space data of the target obstacle model multiple times (for example, three times) according to the spatial data of the actual obstacle model in the multiple images.
[0121] For example, taking the forward-looking camera as an example, the on-board equipment can project the target space data of N target obstacle models into the image captured by the forward-looking camera (hereinafter referred to as the forward-looking image), and collect the target features of the N target obstacle models in the forward-looking camera. It can be understood that since the N target obstacle models include N2 historical obstacle models, the N target obstacle models need to be resampled to obtain the target features.
[0122] The vehicle-mounted equipment can use the neural network model to determine the actual obstacles in the front view image, referring to Figure 1B Actual obstacles in the forward-view image may include pedestrian 02 and second vehicle 04. After the onboard device determines the actual obstacles, it can use the neural network model to associate the target obstacle model closest to pedestrian 02 with pedestrian 02. Similarly, it can associate the target obstacle model closest to second vehicle 04 with second vehicle 04.
[0123] In some embodiments, the on-board device can adjust the target spatial data of the target obstacle model corresponding to pedestrian 02 based on the spatial size, position, and rotation angle of pedestrian 02, thereby updating the target spatial data of the target obstacle. The adjustment process, for example, involves obtaining an offset between the spatial size, position, and rotation angle of pedestrian 02 and the first obstacle corresponding to pedestrian 02, then adding this offset to the spatial size, position, and rotation angle of the target obstacle model. The target spatial data of the target obstacle model corresponding to pedestrian 02 is then updated by adjusting the first positions of multiple feature points based on the added spatial size, position, and rotation angle of the target obstacle model. This means that the position information of the corresponding M feature points in the target spatial data is also updated.
[0124] Similarly, using the above method, each actual obstacle in multiple images can be mapped to N target obstacle models, and the target space data of the N target obstacle models can be updated. After multiple (e.g., two, three, four, five, etc.) rounds of the above optimization, the target space data of the N target obstacle models can be updated to the third space data.
[0125] S305 : Sampling features of the plurality of images according to the third spatial data to obtain a second feature of the target obstacle model, and generating a first environment image according to the second feature.
[0126] For example, after determining the third spatial data of the target obstacle model, features of multiple images can be sampled based on the second position data of each feature point in the third spatial data to obtain a second feature of the target obstacle model. The sampling process can be similar to the process of sampling the image features of multiple images based on the first spatial data to obtain the first feature of the first obstacle model in S302. It will be understood that the similarity between the second feature of the target obstacle model and the feature of the actual obstacle is greater than a set threshold. The set threshold can be, for example, any value between 60% and 95%.
[0127] For example, in some embodiments of the present application, after the vehicle-mounted device determines the second feature of the target obstacle, it can generate a first environment image based on the second feature.
[0128] In some embodiments, because some target obstacle models do not correspond to actual obstacle models, the similarity between the second features of these target obstacle models and the features of the corresponding actual obstacles in multiple images is relatively low, and may be less than a set threshold. However, the on-board device may still retain these target obstacle models. In the process of generating the first environment image based on the second features, the target obstacle models whose second features have a similarity between the features of the actual obstacles less than the set threshold may be discarded. In other words, only the second features of the target obstacle models whose second features have a similarity between the features of the actual obstacles greater than the set threshold are used to generate the first environment image.
[0129] It will be appreciated that since the second feature is obtained by sampling features from multiple images using M feature points of the target obstacle, the number M can be predetermined, and the value of M only needs to represent the spatial size, position, and rotation angle of the corresponding target obstacle. Therefore, the value of M can be set to a smaller value, for example, 13 in the embodiment of the present application, or 7 or 19 in other embodiments. This significantly reduces the amount of data required for the second feature, thereby reducing the computing power required by the onboard device. Furthermore, the first environment image generated by the onboard device based on the second feature has three-dimensional properties. For example, the spatial size, position, and rotation angle of each third obstacle can be determined, thus avoiding the loss of obstacle information (such as the loss of obstacle height information in the BEV space). Therefore, it has a better visual effect than the BEV image. Furthermore, when the vehicle 10 is driving on a slope, the first environment image can also display information about obstacles on both the uphill and downhill slopes, thereby improving the accuracy of obstacle detection by the onboard device.
[0130] The following describes a process in which the vehicle-mounted device generates the first environment image.
[0131] For example, Figure 7According to some embodiments of the present application, a schematic diagram of a vehicle-mounted device generating a first environment image is shown.
[0132] For example, after obtaining multiple images of the current frame, the on-board device can map the first spatial data of the N first obstacle models to the coordinate space of the multiple images, and sample the multiple images based on the first positions of the M feature points, thereby obtaining the first features of the N first obstacle models. The process of obtaining the first features of the N first obstacle models can refer to the process of S302.
[0133] After obtaining the first features of N first obstacle models, the on-board device can select N1 first obstacle models whose first features have the highest similarity to the actual obstacle. It is understood that the actual obstacle is an obstacle in multiple images. After selecting the N1 first obstacle models, the first spatial data and first features of the N1 first obstacle models can be determined. The process of determining the first spatial data and first features of the N1 first obstacle models can refer to the process of S303. The on-board device can then obtain the third features and second spatial data of N2 historical obstacle models, merge the second spatial data of the N2 historical obstacle models with the first spatial data of the N1 first obstacle model to form target spatial data for N target obstacle models, and merge the first features of the N1 first obstacle models with the third features of the N2 historical obstacle models to form target features for N target obstacle models, thereby obtaining the target spatial data and target features for the N target obstacle models. The process of obtaining the second spatial data and third features of the N2 historical obstacles can refer to the process of S304.
[0134] In this embodiment, the on-board device optimizes the target space data of the target obstacle model by, for example, sampling features from multiple images based on the target space data of N target obstacle models to update the target features of the N target obstacle models. It will be appreciated that because the target obstacle model includes the second space data and third features of N2 historical obstacle models, the features of the N target obstacle models need to be updated. Furthermore, the original target features of the N target obstacle models can be used by the self-attention module to determine the features of more important target obstacles.
[0135] For example, inter-frame interaction can be performed on the target features of N target obstacle models (for example, target features that have not been resampled through target spatial data) to obtain features after inter-frame interaction. For example, inter-frame interaction refers to the interaction between the first features of the N1 first obstacle models in the current frame of the target obstacle model and the third features of the N2 historical obstacle models in the previous frame.
[0136] For example, Figure 8AA schematic diagram of inter-frame interaction is shown according to some embodiments of the present application.
[0137] like Figure 8A As shown in the figure, a self-attention module is configured in the vehicle-mounted device, and the self-attention module can perform self-attention optimization through a transformer model.
[0138] For example, the process of inter-frame interaction can be to multiply the target features of N target obstacle models, i.e., the first features of N1 first obstacles and the third features of N2 historical obstacles, with the query weight matrix in the transformer model to obtain the query features.
[0139] Then the third feature of the N2 historical obstacle models is multiplied by the key weight matrix in the transformer model to obtain the key feature.
[0140] Similarly, the third feature of the N2 historical obstacle models is multiplied by the value weight matrix in the transformer model to obtain the value feature.
[0141] For example, in the transformer model, key features are used to match query features to determine which features are most important for generating the features of the first environment image. The value features contain the actual information corresponding to the key features, which are weighted and summed in the attention mechanism to generate the final features of the first environment image.
[0142] For example, the query weight matrix, key weight matrix, and value weight matrix are trainable parameters of the Transformer model, which are updated and optimized during model training using the backpropagation algorithm. Therefore, after training the corresponding weight matrix, simply multiplying the input feature data by the corresponding weight matrix will obtain the corresponding query features, key features, and value features.
[0143] After obtaining the query, key, and value features, the self-attention module can be used to optimize the input features to obtain the corresponding output. The optimization process of the self-attention module is described in detail below. For example, the output of the inter-frame interaction can be called the inter-frame interaction feature. It is understood that the N inter-frame interaction features also include the N1 inter-frame interaction features of the first obstacle model and the N2 inter-frame interaction features of the historical obstacle models.
[0144] After completing the inter-frame interaction, the on-board equipment can perform intra-frame interaction on the features after the inter-frame interaction of the N target obstacle models, thereby obtaining the features after the intra-frame interaction.
[0145] For example, Figure 8B A schematic diagram of inter-frame interaction is shown according to some embodiments of the present application.
[0146] like Figure 8B As shown, the input for intra-frame interaction can be the features of N target obstacle models after inter-frame interaction. For example, the interaction process involves inputting the inter-frame interaction features of the N target obstacle models into the query weight matrix, key weight matrix, and value weight matrix of the transformer model to obtain the corresponding query features, key features, and value features. It can be understood that the inter-frame interaction features of the N target obstacle models include the inter-frame interaction features of the N1 first obstacle models and the inter-frame interaction features of the N2 historical obstacle models. The on-board device then optimizes the query features, key features, and value features based on the self-attention module to obtain the corresponding output. The optimization process of the self-attention module is described in detail below.
[0147] It is understood that after the intra-frame interaction is completed, the intra-frame interaction features of the N target obstacle models can be obtained. The on-board device can then perform feature fusion on the target features resampled from the target obstacle models with the features of the target obstacle models after the intra-frame interaction, thereby obtaining updated target features. For example, in an embodiment of the present application, the target features resampled from the N target obstacle models can be added to the features of the corresponding N target obstacle models after the intra-frame interaction to obtain new features of the N target obstacle models, which can serve as the updated target features of the N target obstacle models. It is understood that since the target features of the N target obstacle models include the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models, after the target features of the N target obstacle models are updated, the corresponding first features of the N1 first obstacle models and the third features of the N2 historical obstacle models will also be updated.
[0148] Next, we introduce the optimization process of the self-attention module.
[0149] For example, Figure 8C According to some embodiments of the present application, a schematic diagram of an optimization process of a self-attention module is shown.
[0150] like Figure 8C As shown in the figure, after the vehicle-mounted device obtains the corresponding query features, key features, and value features, it can perform a matrix multiplication operation (such as a dot product operation) on each query feature and all key features to obtain an attention score. This score reflects the similarity or correlation between the query feature and each key feature.
[0151] Then, the onboard device normalizes the attention scores. For example, the attention scores can be normalized by a softmax function so that the sum of all attention scores is 1. In this way, each attention score represents the weight of the corresponding value feature.
[0152] Next, the normalized attention scores are assigned as weights to the corresponding value features. This means that the attention scores are multiplied by each value feature through matrix multiplication. This means that value features corresponding to key features with higher similarity to the query feature receive a higher weight. The matrix multiplication results of all value features and their corresponding weights are then summed. This weighted sum is then applied to the query features, using the self-attention scores as weights. This weighted sum is the output of the self-attention mechanism.
[0153] It can be understood that the optimization process of the self-attention module in the above-mentioned inter-frame interaction and intra-frame interaction process can refer to Figure 8C process.
[0154] After obtaining the updated target features of the N target obstacle models, the on-board device can further adjust the target space data of the N target obstacle models based on the spatial data of multiple images. The adjustment process refers to the process of S304, thereby updating the target space data of the N target obstacle models. It can be understood that at this point, the target space data and target feature data of the N target obstacle models have been updated. In order to ensure the accuracy of the generated first environment image, the on-board device can also update the target space data of the N target obstacle models multiple times based on the neural network to obtain the third space data of the N target obstacle models. Similarly, the target features of the N target obstacle models can also be optimized multiple times in a cycle. For example, in some embodiments of the present application, the third space data is obtained by performing two cycles of optimization and adjustment on the target space data.
[0155] Exemplarily, the on-board device may perform multiple cyclic optimizations on the target space data and target features after the N target obstacle models are updated. The cyclic optimization process may include, for example, the above-mentioned inter-frame interaction and intra-frame interaction of the target features, projecting the target space data into the coordinate space of multiple images to resample the N target obstacle models to obtain features of the N resampled target obstacle models, adding and fusing the features of the N resampled target obstacle models with the features after the intra-frame interaction to update the target features, and further optimizing the target space data based on the updated target features to update the target space data.
[0156] In some embodiments of the present application, in the last cycle of the vehicle-mounted device, the feature obtained after intra-frame interaction may be the fourth feature, and the feature obtained by projecting the target spatial data into the coordinate space of multiple images and resampling the target obstacle model may be the second feature. It can be understood that the target spatial data is the spatial data updated in the last cycle of the vehicle-mounted device, and the target spatial data can be used as the third spatial data. In other words, the second feature is obtained by sampling the third spatial data in the coordinate space of multiple images. In the last cycle, the feature obtained by adding and fusing the second and fourth features is used as the fifth feature, and the vehicle-mounted device can generate the first environment image based on the fifth feature.
[0157] Through the above-mentioned cyclic optimization process, the accuracy of generating the first environment image can be improved.
[0158] The vehicle-mounted devices involved in the above embodiments are introduced below.
[0159] For example, Figure 9 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device 100 is shown.
[0160] The vehicle-mounted device 100 can be used to implement the image processing methods provided in the aforementioned embodiments.
[0161] like Figure 9 As shown, the in-vehicle device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output device 105. Among them:
[0162] The processor 101 may include one or more processing units, such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor or a programmable logic device (FPGA), a neural network processing unit (NPU), or the like. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be used to optimize a neural network model to be run. For example, in some embodiments of the present application, the neural network model may optimize the spatial data of a first obstacle model, and the NPU may be used to run the neural network model to be run.
[0163] System memory 102 is a volatile memory, such as random-access memory (RAM) or double data rate synchronous dynamic random access memory (DDR SDRAM). System memory is used to temporarily store data and / or instructions. For example, in some embodiments, system memory 102 can be used to store data provided by the aforementioned different services, such as sensor data, image data, or video data, and can also be used to store instructions for the image processing methods provided in the aforementioned embodiments.
[0164] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card. In other embodiments, the non-volatile memory 103 may be used to store instructions for the image processing methods provided in the aforementioned embodiments.
[0165] In particular, the system memory 102 and the non-volatile memory 103 may respectively include a temporary copy and a permanent copy of the instruction 107. The instruction 107 may include instructions that, when executed by at least one of the processors 101, enable the vehicle-mounted device 100 to implement the image processing method provided in various embodiments of the present application.
[0166] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the in-vehicle device 100, thereby enabling communication with any other suitable device via one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the in-vehicle device 100, for example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the in-vehicle device 100 may communicate with other devices via the communication interface 104. For example, the in-vehicle device 100 may obtain corresponding data from other devices via the communication interface 104.
[0167] The input / output device 105 can be an input device such as a keyboard, a mouse, etc., and an output device such as a display, etc. The user can interact with the in-vehicle device 100 through the input / output device 105 .
[0168] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface with other modules of the vehicle-mounted device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide interfaces to the system memory 102 and the non-volatile memory 103.
[0169] In some embodiments, at least one of the processors 101 may be packaged together with the logic of one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).
[0170] I understand. Figure 9 The structure of the vehicle-mounted device 100 shown is merely an example. In other embodiments, the vehicle-mounted device 100 may include more or fewer components than shown, or may combine or separate certain components, or may have different component arrangements. The components shown may be implemented in hardware, software, or a combination of software and hardware.
[0171] It can be understood that the vehicle-mounted device 100 can be any device configured on the vehicle, including but not limited to mobile phones, car computers, terminals in self-driving, wireless terminals in transportation safety, terminals in smart cities, etc.
[0172] An embodiment of the present application also provides a program product, which, when executed on a vehicle-mounted device, can enable the vehicle-mounted device to implement the image processing methods provided in the aforementioned embodiments.
[0173] An embodiment of the present application also provides a readable storage medium, which stores one or more programs. When the one or more programs are executed by a vehicle-mounted device, the vehicle-mounted device implements the image processing methods provided by the aforementioned embodiments.
[0174] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of this application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0175] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.
[0176] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0177] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or via other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable storage for transmitting information via electrical, optical, acoustical, or other propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) via the Internet. Thus, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0178] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0179] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.
[0180] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0181] While the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.
Claims
1. An image processing method for a vehicle-mounted device, characterized in that: The vehicle-mounted device stores first spatial data of N first obstacle models, each of the first spatial data of the first obstacle model including first position data of M feature points of the first obstacle model; the method includes: In response to a request to generate a first environment image of the vehicle, acquiring a plurality of images, wherein the plurality of images are images acquired from different perspectives by a plurality of sensors of the vehicle, and the plurality of images include images of actual obstacles around the vehicle; Sampling features of the plurality of images according to first position data of each feature point in the first spatial data to obtain first features of the N first obstacle models; selecting N1 first obstacle models from the N first obstacle models based on the multiple images, wherein similarities between first features of the N1 first obstacle models and features of the actual obstacle are greater than similarities between first features of remaining first obstacle models and features of the actual obstacle; Adjusting positions of the M feature points of the N1 first obstacle models and the M feature points of the N2 historical obstacle models at least once based on the spatial data of the actual obstacle in the multiple images, thereby obtaining third spatial data; Wherein, N=N1+N2, and each of the third spatial data includes second position data of M feature points corresponding to the obstacle model, and a similarity between a second feature obtained by sampling features of the multiple images based on the second position data of each feature point in the third spatial data and a feature of the actual obstacle is greater than a set threshold; The first environment image is generated according to the second feature.
2. The image processing method according to claim 1, wherein: The M feature points include M points on three coordinate axes that are orthogonal to each other in the coordinate space of the first obstacle model; The first position data of the M feature points are used to represent the initial spatial size, position and rotation angle of the first obstacle model.
3. The image processing method according to claim 1, wherein: The selecting N1 first obstacle models from the N first obstacle models according to the plurality of images includes: According to the preset arrangement positions of the N first obstacle models, a first obstacle model corresponding to the first feature and having a high similarity with a feature of the actual obstacle is selected from between every two adjacent first obstacle models as an obstacle model among the N1 first obstacle models.
4. The image processing method according to claim 1, wherein: The in-vehicle device includes third features of N third obstacle models corresponding to a historically generated second environment image, and fourth spatial data of the N third obstacle models, wherein a similarity between the third features and image features of actual obstacles corresponding to the second environment image is greater than a set threshold, and the third features are obtained by sampling the fourth spatial data of the N third obstacle models from the image features of the actual obstacles corresponding to the second environment image; The second spatial data of the N2 historical obstacle models are determined in the following manner: Selecting, from the N third features of the third obstacle models, N2 third features having the highest similarity to features of an image of an actual obstacle corresponding to the second environment image; The fourth spatial data of the third obstacle model corresponding to the N2 third features is used as the second spatial data of the N2 historical obstacle models.
5. The image processing method according to claim 4, characterized in that Generating the first environment image according to the second feature includes: Performing at least one self-attention adjustment on the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models to obtain a fourth feature; using a fifth feature obtained by fusing the fourth feature with the second feature as a feature for generating the first environment image; The similarity between the fifth feature and the features of the multiple images is higher than the similarity between the second feature and the features of the multiple images.
6. The image processing method according to claim 1, wherein: The plurality of images include distorted images; The vehicle-mounted device stores the distortion error of the first obstacle model, which is obtained by mapping the spatial size, position and rotation angle of the first obstacle model from the undistorted image to the distorted image; The third spatial data includes data obtained by projecting the first spatial data of the first obstacle model into the coordinate space of the multiple images after adjusting the distortion error.
7. A vehicle-mounted device, characterized in that: include: a memory for storing instructions; At least one processor is configured to execute the instructions so that the vehicle-mounted device implements the image processing method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the image processing method according to any one of claims 1 to 6.
9. A computer program product, characterized in that When the computer program product is run on a device, the device is caused to execute the image processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Obstacle detection method and device
CN112712009A
Obstacle detection method and device, vehicle and storage medium
CN113537047A