Image processing method, readable storage medium, program product and vehicle-mounted equipment

By storing and optimizing the feature point data of the obstacle model in the on-board equipment, the problem of increasing computing power demand caused by the expansion of the surrounding environment detection range is solved, and efficient and accurate environmental image generation is achieved.

CN120148009AActive Publication Date: 2025-06-13CONTINENTAL SMART CORE TECH (SHANGHAI) CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510621997.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

As the demand for the surrounding environment detection range of vehicles increases, the size of BEV space increases in square order, resulting in an increase in data required to generate BEV space, increasing the computing power demand of the vehicle-side platform and making it difficult to deploy on the vehicle-side platform.

Method used

By storing the first spatial data of N first obstacle models in the vehicle-mounted device, each model includes position data of M feature points. The vehicle-mounted device acquires multiple images in response to the first environmental image request of the generated vehicle, samples the image features according to the first spatial data, selects a model with high similarity to the actual obstacle features, and optimizes the model data to generate the first environmental image.

Benefits of technology

The computing power demand of the on-board equipment is reduced, the calculation amount required to generate the first environmental image is reduced, and the efficiency and accuracy of the vehicle's detection of the surrounding environment is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148009A_ABST
    Figure CN120148009A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to an image processing method, a readable storage medium, a program product and vehicle-mounted equipment. The image processing method comprises the following steps: storing first spatial data of a plurality of first obstacle models in the vehicle-mounted equipment, wherein the first spatial data of each first obstacle model comprises first position data of a plurality of feature points of the first obstacle model; in the process that the vehicle-mounted equipment generates the first environment image based on the multiple images of the multiple visual angles, the size of the first obstacle model can be optimized through the actual obstacles in the multiple images, and then the first environment image is generated through the multiple feature points of the first obstacle model. According to the method, only the plurality of feature points of the first obstacle need to be calculated in the process of generating the first environment image, so that the computing power demand of the vehicle-mounted equipment for generating the first environment image is reduced, and the generated first environment image is richer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly relates to an image processing method, a readable storage medium, a program product, and a vehicle-mounted device. Background Art

[0002] Multi-view image three-dimensional object detection currently has extensive applications in many fields and application scenarios, such as autonomous driving scenarios, intelligent transportation scenarios, industrial automation, and virtual reality and augmented reality scenarios, etc.

[0003] For example, in the assisted driving technology, through the bird's eye view (BEV) technology, an image in which the multi-view image features of the vehicle's surrounding environment are projected into the BEV space can be provided, and then the global view of the positions, sizes, and attributes of dynamic obstacles can be estimated, thus facilitating users to perform path planning and decision-making.

[0004] However, as the vehicle's demand for the detection range of the surrounding environment increases, the size of the BEV space grows exponentially, and the data to be processed for generating the BEV space increases, thereby increasing the computing power requirements of the vehicle-end platform, making it difficult to deploy the corresponding technology on the vehicle-end platform. Summary of the Invention

[0005] Embodiments of this application provide an image processing method, a readable storage medium, a program product, and a vehicle-mounted device.

[0006] In a first aspect, an embodiment of the present application provides an image processing method for an in-vehicle device on a vehicle. The in-vehicle device stores first spatial data of N first obstacle models, and the first spatial data of each first obstacle model includes first position data of M feature points of the first obstacle model. The method includes: in response to a request for generating a first environmental image of the vehicle, the in-vehicle device acquires multiple images, which are images from different perspectives collected by multiple sensors of the vehicle, and the multiple images include images of actual obstacles around the vehicle. The in-vehicle device samples the features of the multiple images according to the first position data of each feature point in the first spatial data to obtain first features of N first obstacle models. The in-vehicle device selects N1 first obstacle models from the N first obstacle models according to the multiple images, where the similarity between the first features of the N1 first obstacle models and the features of the actual obstacles is greater than the similarity between the first features of the remaining first obstacle models and the features of the actual obstacles. The in-vehicle device optimizes the first spatial data of the N1 first obstacle models and the second spatial data of N2 historical obstacle models to obtain third spatial data of the N1 first obstacle models and the N2 historical obstacle models; where N = N1 + N2, and each third spatial data includes second position data of M feature points of the corresponding obstacle model, and the similarity between the second features obtained by sampling the features of the multiple images according to the second position data of each feature point in the third spatial data and the features of the actual obstacles is greater than a set threshold. Thus, the in-vehicle device can generate a first environmental image according to the second features.

[0007] In some embodiments of the present application, the in-vehicle device stores first spatial data of N first obstacle models, and the first spatial data of the first obstacle model includes position data of M feature points of the first obstacle model. During the process of the in-vehicle device generating the first environmental image, the spatial data of the first obstacle model can be optimized by the spatial data of the actual obstacles in the multiple collected images, so that the first obstacle model can replace the actual obstacle. Then the in-vehicle device generates a first environmental image according to the M feature points of the first obstacle model, thereby reducing the computing power requirement of the in-vehicle device.

[0008] In some embodiments of the present application, since the vehicle needs to continuously generate environmental images during driving, the in-vehicle device can also use N2 historical obstacle models that have been optimized historically to participate in the process of generating the first environmental image currently. It can be understood that the in-vehicle device can estimate the position of the historical obstacle model at the current moment based on the speed of the historical obstacle model relative to the vehicle, so as to participate in the generation of the first environmental image at the current moment. Since the sizes of the historical obstacle models have been optimized, adding the historical obstacle models in the process of generating the first environmental image can improve the accuracy of generating the first environmental image and the speed of generating the first environmental image.

[0009] In a possible implementation of the first aspect above, the M feature points include M points on three mutually orthogonal coordinate axes in the coordinate space of the first obstacle model.

[0010] The first position data of the M feature points is used to represent the initial spatial size, position, and rotation angle of the first obstacle model.

[0011] In some embodiments of the present application, the spatial size, position, and rotation angle and other information of the first obstacle model can be represented by the positions of the M feature points. Optimizing the first spatial data of the first obstacle model is actually to adjust the positions of the M feature points of the first obstacle model.

[0012] Exemplarily, in some embodiments of the present application, the value of M can be 13. For example, there are 5 feature points on each of the three coordinate axes in the coordinate space of the first obstacle model, and three of them pass through the coordinate origin, and the coordinate origin can be the geometric center of the first obstacle. In other embodiments, M can also take other values. It can be understood that since the value of M is small, the computing power requirement for generating the first environmental image according to the first obstacle model is small.

[0013] In a possible implementation of the first aspect above, the selecting N1 first obstacle models from the N first obstacle models according to multiple images includes: The in-vehicle device selects, from between every two adjacent first obstacle models, the first obstacle model with a high similarity between the corresponding first feature and the feature of the actual obstacle according to the preset arrangement positions of the N first obstacle models as the obstacle model among the N1 first obstacle models.

[0014] It can be understood that by this way of selecting the N1 first obstacle models with higher similarity, the in-vehicle device does not need to sort the N first obstacle models according to similarity, thereby further reducing the computing power requirement of the in-vehicle device.

[0015] In a possible implementation of the above first aspect, the vehicle-mounted device includes the third features of N third obstacle models corresponding to the second environmental image generated historically, and the fourth spatial data of the N third obstacle models, where the similarity between the third features and the image features of the actual obstacles corresponding to the second environmental image is greater than a set threshold, and the third features are sampled from the image features of the actual obstacles corresponding to the second environmental image by the fourth spatial data of the N third obstacle models; the second spatial data of N2 historical obstacle models is determined by the following method: selecting N2 third features with the highest similarity to the features of the image of the actual obstacle corresponding to the second environmental image from the third features of the N third obstacle models; using the fourth spatial data of the third obstacle models corresponding to the N2 third features as the second spatial data of the N2 historical obstacle models.

[0016] In some embodiments of the present application, during the process of the vehicle-mounted device obtaining N2 historical obstacle models, N2 third obstacle models corresponding to N2 third features with the highest similarity to the real obstacles in the previous frame can be selected from the N third obstacle models corresponding to the environmental image generated in the previous frame as the N2 historical obstacle models. Similarly, the fourth spatial data corresponding to the N2 historical obstacle models can be used as the second spatial data of the N2 historical obstacle models. Among them, the N2 third features with the highest similarity to the real obstacles are, for example, sorting the similarities between the N third features and the real obstacles according to the magnitudes, and then selecting the top N2 third features with the highest similarities.

[0017] In a possible implementation of the above first aspect, optimizing the first spatial data of N1 first obstacle models and the second spatial data of N2 historical obstacle models to obtain the third spatial data of the N1 first obstacle models and the N2 historical obstacle models includes: adjusting the positions of M feature points of the N1 first obstacle models and the positions of M feature points of the N2 historical obstacle models at least once according to the spatial data of the actual obstacles in multiple images, so as to obtain the third spatial data.

[0018] In some embodiments of the present application, the vehicle-mounted device can optimize the first spatial data of the determined N1 first obstacle models and the N2 spatial data of the N2 historical obstacle models according to the spatial data of multiple images for multiple times, so as to obtain the third spatial data of the N1 first obstacle models and the N2 historical obstacle models. It can be understood that the optimization process is, for example, a process of reducing the errors in dimensions, positions, rotation angles, etc. between the N1 first obstacle models and the N2 historical obstacle models and the actual obstacles in multiple images.

[0019] In a possible implementation of the first aspect described above, the N2 historical obstacle models further include corresponding third features; generating the first environmental image according to the second feature includes: performing at least one self-attention adjustment on the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models to obtain a fourth feature; using the fifth feature obtained by fusing the fourth feature and the second feature as the feature for generating the first environmental image; wherein, the similarity between the fifth feature and the features of multiple images is higher than the similarity between the second feature and the features of multiple images.

[0020] In some embodiments of the present application, the vehicle-mounted device can also optimize the feature data corresponding to the N1 first obstacle models and the N2 historical obstacle models through a self-attention mechanism, thereby improving the accuracy of generating the first environmental image based on the N1 first obstacle models and the N2 historical obstacle models.

[0021] In a possible implementation of the first aspect described above, the multiple images include distorted images; the spatial dimensions, positions, and rotation angles of the first obstacle models stored in the vehicle-mounted device are distorted errors for mapping from undistorted images to distorted images; the third spatial data includes data obtained by projecting the first spatial data of the first obstacle models after being adjusted by the distorted errors into the coordinate spaces of the multiple images.

[0022] In some embodiments of the present application, since the images captured by the cameras on the vehicle may be distorted images. Therefore, in some embodiments, it is necessary for the vehicle-mounted device to perform distortion processing on the first obstacle models or historical obstacle models mapped into the coordinate spaces of the multiple images. However, the process of calculating the distortion offset through the distortion formula will increase the computing power requirements of the vehicle-mounted device. Therefore, by storing the distortion errors of each camera in the vehicle-mounted device, the vehicle-mounted device only needs to query to obtain the distortion errors, so as to adjust the positions of the N1 first obstacle models and the N2 historical obstacle models mapped into the coordinate spaces of the multiple images, thereby further reducing the computing power requirements of the vehicle-mounted device.

[0023] In a second aspect, the present application provides a vehicle-mounted device, including: a memory for storing instructions; at least one processor for executing the instructions to enable the device to implement the image processing method provided in the first aspect and any possible implementation of the first aspect described above. The beneficial effects achievable by the second aspect can refer to the beneficial effects of the image processing method provided in any implementation manner of the first aspect, and will not be elaborated here.

[0024] In a third aspect, the present application provides a computer-readable storage medium storing instructions that, when executed by a device, cause a computer to implement the image processing method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable by the third aspect can refer to the beneficial effects of the image processing method provided in any embodiment of the first aspect, which will not be elaborated here.

[0025] In a fourth aspect, the present application provides a computer program product that, when running on a device, causes the device to implement the image processing method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable by the fourth aspect can refer to the beneficial effects of the image processing method provided in any embodiment of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1A Shows a flowchart for generating a BEV space; Figure 1B Shows a schematic diagram during vehicle driving; Figure 2 Shows a schematic diagram of feature points according to some embodiments of the present application; Figure 3 Shows an implementation flowchart for generating a vehicle environment image according to an embodiment of the present application; Figure 4A According to some embodiments of the present application, shows a schematic diagram of obtaining a distortion error; Figure 4B According to some embodiments of the present application, shows a schematic diagram of obtaining the first features of N first obstacle models; Figure 5 According to some embodiments of the present application, shows a schematic diagram of a process for processing a distortion error; Figure 6 According to some embodiments of the present application, shows a process of selecting N1 first obstacle models; Figure 7 According to some embodiments of the present application, shows a schematic diagram of an in-vehicle device generating a first environment image; Figure 8A According to some embodiments of the present application, shows a schematic diagram of inter-frame interaction; Figure 8B According to some embodiments of the present application, shows a schematic diagram of inter-frame interaction; Figure 8C According to some embodiments of the present application, shows a schematic diagram of an optimization process of a self-attention module; Figure 9According to some embodiments of the present application, a schematic structural diagram of an in-vehicle device 100 is shown. Detailed implementation manners

[0027] Illustrative embodiments of the present application include, but are not limited to, an image processing method, a readable storage medium, a program product, and an in-vehicle device.

[0028] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification and specific implementation manners.

[0029] For ease of understanding, some terms involved in the present application and related technologies are explained below.

[0030] BEV technology: BEV technology is a technology that converts image information from the image space to the BEV space by a neural network. Through BEV technology, a complex three-dimensional environment can be simplified into a two-dimensional image. For example, in the field of autonomous driving, through BEV technology, a panoramic view looking down from above the vehicle can be generated based on the spatial information around the vehicle, so as to comprehensively display the environment around the vehicle, including the front, rear, left, right, and top situations, which enables the autonomous driving system to better understand the surrounding environment and improve the accuracy of perception and decision-making.

[0031] Next, the process of generating the BEV space by the in-vehicle device of the vehicle is introduced.

[0032] For example, Figure 1A A flowchart for generating the BEV space is shown.

[0033] It can be understood that the following processes can be executed by the in-vehicle device. The in-vehicle device in the embodiments of the present application can also be referred to as an in-vehicle terminal. The in-vehicle terminal can be a mobile phone, a car machine, a terminal in self-driving, a wireless terminal in transportation safety, a terminal in a smart city, etc. Hereinafter, a car machine will be used as an example for illustration. However, it can be understood that the technical solutions described in the present application are applicable to various in-vehicle electronic devices for three-dimensional object detection as described above, and are not limited to car machines.

[0034] S101, obtain sensor data collected by multiple sensors.

[0035] In some embodiments of the present application, the car machine deployed on the vehicle 10 can obtain sensor data collected by multiple sensors.

[0036] For example, Figure 1B A schematic diagram during vehicle driving is shown.

[0037] It can be understood that the vehicle 10 is usually equipped with multiple sensors, such as cameras (including, for example, front view, side view, rear view cameras, etc.), radars, lidars, etc. As Figure 1B shown, during the driving process of the vehicle 10, multiple sensor data can be collected at the same moment through a timestamp synchronization mechanism or a hardware synchronization mechanism. The sensor data is, for example, image data of different perspectives of the vehicle 10 collected by the camera, and / or point cloud data of various obstacles collected by the radar, etc.

[0038] For example, the vehicle 10 can collect environmental data near the vehicle 10 through sensors. Referring to Figure 1B , the environmental data collected by the sensors can, for example, include sensor data of the first vehicle 01, the pedestrian 02, the first building 03, the second vehicle 04, and the second building 05. The vehicle computer on the vehicle 10 can obtain the environmental data of each perspective collected by the sensors on the vehicle 10 at the same moment.

[0039] S102, preprocess the sensor data.

[0040] Exemplarily, after the vehicle computer on the vehicle 10 obtains the sensor data, it can preprocess the sensor data. For example, during the image processing, the images collected by the camera can be denoised, distortion corrected, etc. During the point cloud processing, the point cloud data collected by the lidar and / or radar can be filtered, denoised, and segmented. Then, the data of different sensors are aligned to the same time point to ensure data consistency.

[0041] S103, perform coordinate transformation on the sensor data.

[0042] Exemplarily, the vehicle computer can also perform coordinate transformation on the sensor data collected by the sensors, so as to convert the data collected by the sensors from the sensor coordinate system to the vehicle coordinate system. For example, the data of the camera, radar, and lidar are converted to the vehicle coordinate system centered on the vehicle 10.

[0043] Then, the sensor data converted to the vehicle coordinate system is further converted from the vehicle coordinate system to the BEV coordinate system. For example, the sensor data in the vehicle coordinate system is converted to the bird's-eye view coordinate system (top view).

[0044] Among them, the data collected by the camera can be transformed from the perspective view to the top view through inverse perspective mapping (IPM). The radar and lidar data can map the 3D point cloud data to the 2D top view through projection.

[0045] It can be understood that when processing sensor data, the in-vehicle computer on vehicle 10 can only process the sensor data within a preset range of vehicle 10. The preset range can be, for example, a circular range centered on vehicle 10 with a radius of R.

[0046] It can be understood that in some other embodiments, the range for the sensors of vehicle 10 to collect data can also be of other shapes, and the embodiments of the present application do not limit the shape of the range for the sensors of vehicle 10 to collect data.

[0047] S104, fuse the sensor data.

[0048] Exemplarily, during the process of fusing sensor data, the in-vehicle computer can perform feature extraction on the sensor data. For example, the in-vehicle computer can extract features such as lane lines and obstacles from the images captured by the camera, or extract information such as the position and speed of objects from the point clouds of the radar and lidar. Then the in-vehicle computer performs data fusion to fuse the features of different sensors.

[0049] S105, generate a BEV map based on the fused sensor data.

[0050] Exemplarily, in some embodiments, the in-vehicle computer can map the fused data into a two-dimensional grid map, and each grid represents a certain range of space (such as 0.1m×0.1m). In some embodiments, the in-vehicle computer can also perform semantic information addition (for example, annotating semantic information such as lane lines, obstacles, pedestrians, and vehicles in the BEV map) and dynamic update (for example, dynamically adjusting the BEV map according to the real-time update of the sensor data), etc. Exemplarily, the BEV map generated by the in-vehicle computer can also refer to Figure 1B the view.

[0051] It can be understood that during the above process of generating the BEV map, only the sensor data of the space within the radius R of vehicle 10 is processed. Therefore, if it is necessary to increase the detection range of vehicle 10 (R becomes larger), then more sensor data in the space to be processed will be required, thereby increasing the computing power requirement of the in-vehicle computer of vehicle 10. For example, if the detection range of vehicle 10 increases, the data processed in the above processes of S102 to S105 will increase, thereby increasing the computing power requirement of the in-vehicle computer of vehicle 10.

[0052] For example, in the process of S104, in order to increase the detection range of vehicle 10, the sampling points in the sensor data by the in-vehicle computer need to be denser to obtain more accurate features of the sensor data. Therefore, the in-vehicle computer has a relatively high computing power requirement for the sensors. In the process of S105, since the in-vehicle computer extracts more features of the sensor data, the in-vehicle computer also requires relatively high computing power to project the features of the fused sensor data into the BEV space.

[0053] As described above, as the vehicle's demand for the detection range of the surrounding environment increases, more data needs to be processed during the vehicle's target detection process in the three-dimensional space, resulting in a high computing power requirement for the in-vehicle computer, making it difficult to deploy the corresponding technology on the vehicle platform.

[0054] To solve the problem of the high computing power required for the in-vehicle computer to generate three-dimensional space data, this application proposes an image processing method. The in-vehicle device stores the first space data of N first obstacle models, and the first space data of each first obstacle model includes the first position data of M feature points of the first obstacle model. The method includes: In response to a request to generate a first environmental image of the vehicle, the in-vehicle device acquires multiple images from different perspectives collected by multiple sensors of the vehicle, and the multiple images include images of actual obstacles around the vehicle.

[0055] The in-vehicle device projects the first space data of N first obstacle models into the coordinate space of the multiple images, samples the features of the multiple images according to the first space data, and obtains the first features of N first obstacle models. The in-vehicle device selects N1 first obstacle models from the N first obstacle models according to the multiple images, where the similarity between the first features of the N1 first obstacle models and the features of the actual obstacles is greater than the similarity between the first features of the remaining first obstacle models and the features of the actual obstacles.

[0056] The in-vehicle device optimizes the first space data of the N1 first obstacle models and the second space data of the N2 historical obstacle models to obtain the third space data of the N1 first obstacle models and the N2 historical obstacle models; where N = N1 + N2, and each third space data includes the second position data of M feature points of the corresponding obstacle model, and the similarity between the second features obtained by sampling the features of the multiple images according to the second position data of each feature point in the third space data and the features of the actual obstacles is greater than a set threshold. The in-vehicle device generates a first environmental image according to the second features.

[0057] Through the above solution, the in-vehicle device does not need to generate a first environmental image based on the features corresponding to a large amount of sensor data, but generates a first environmental image based on the optimized M feature points in the N1 preset first obstacle models and the second features collected from the optimized M feature points in the N2 historical obstacle models in the multiple images. Since each obstacle model (including the first obstacle model and the historical obstacle model) can represent the spatial information of the complete obstacle model (such as the spatial size, position, and rotation angle of the obstacle model) through fewer feature points, therefore, the in-vehicle computer does not need to occupy much computing power during the process of obtaining the second features, thus reducing the computing power requirement of the in-vehicle computer.

[0058] In some embodiments of the present application, the multiple first obstacle models stored in the vehicle-mounted device may include common objects during traffic driving, such as buildings, people, animals, traffic lights, vehicles (including sedans, motorcycles, bicycles, etc.), trees, rivers, etc. The first spatial data of the first obstacle may be the first positions of M feature points of the first obstacle model. Among them, the M feature points of the first obstacle model are, for example, M points on three mutually orthogonal coordinate axes in the coordinate space of the first obstacle model, and the first position data of the M feature points is used to represent the initial spatial dimensions (such as length, width, and height data), position, and rotation angle, etc. of the first obstacle model. Among them, the spatial dimensions of the first obstacle model may be the average dimensions of common objects during traffic driving. For example, the first obstacle model may be a human. Among humans, the average height of an adult male is 1.75 m, and the average height of an adult female is 1.62 m. The first obstacle model may be a small sedan, with a length of about 4.0 meters to 4.5 meters, a width of about 1.7 meters to 1.8 meters, and a height of about 1.4 meters to 1.5 meters.

[0059] For example, Figure 2 Some embodiments according to the present application show a schematic diagram of feature points.

[0060] As Figure 2 shown, in the coordinate space of the first obstacle model, the first obstacle model can be regarded as a rectangular block, and the geometric center of the rectangular block can be the coordinate center of the first obstacle model. In some embodiments of the present application, taking M = 13 as an example, 13 feature points can be established through the dimensions of the first obstacle model in three-dimensional directions. The 13 feature points may include 5 feature points evenly distributed along the coordinate axes in the length direction, width direction, and height direction of the first obstacle model respectively. Among them, for the 5 feature points in each coordinate axis direction, the distance between the two end feature points is the dimension of the first obstacle model in that coordinate direction. It can be understood that since the feature point at the geometric center of the first obstacle passes through the three coordinate axes, the feature point at the geometric center of the first obstacle is calculated twice, and the first obstacle model includes a total of 13 feature points.

[0061] It can be understood that in other embodiments, the number of feature points of the first obstacle model may also be other numbers, such as M may be 7, 19, etc. Or the feature points of the first obstacle model are distributed in other ways, as long as they can represent the spatial data such as the dimensions and rotation angle of the first obstacle model. The embodiments of the present application do not limit the number and positions of the feature points of the first obstacle model.

[0062] Next, the process of the in-vehicle terminal generating the first environmental image based on the data collected by the sensors on the vehicle in the embodiments of the present application will be introduced.

[0063] Figure 3 An embodiment flowchart of generating a vehicle environmental image is shown according to an embodiment of the present application.

[0064] Exemplarily, in some embodiments of the present application, the in-vehicle device stores the first spatial data of N first obstacles, and the first spatial data of each first obstacle model includes the first position data of M feature points of the first obstacle model. For example, in the embodiments of the present application, N is greater than 384, and the first position data is used to identify the initial spatial dimensions (such as length, width, and height data), position, and rotation angle, etc. of the first obstacle model. The spatial dimensions of the first obstacle model can be the average dimensions of common objects in the corresponding traffic driving process. In the initial state, the positions of the N first obstacles can be evenly distributed, for example, and the rotation angles of the N first obstacles can be preset values, such as 0°.

[0065] It can be understood that if N is relatively large, the in-vehicle device needs to spend a relatively long time generating the environmental data of the vehicle, but the accuracy of the generated environmental data of the vehicle is relatively high. Therefore, when specifically setting the number of N, the number of N can be determined with reference to the computing power and accuracy requirements of the in-vehicle device. The embodiments of the present application do not limit the number of N. For example, N can also be 400, 500, or 600, etc.

[0066] It can be understood that the execution subject of each of the following processes can all be the in-vehicle device arranged on the vehicle. During the introduction of each of the following processes, the execution subject of each process is not limited.

[0067] As Figure 3 shown, this process includes: S301, in response to a request for generating the first environmental image of the vehicle, obtain multiple images.

[0068] Exemplarily, in some embodiments of the present application, the multiple images are images of different perspectives collected by multiple sensors of the vehicle, and the multiple images include images of actual obstacles around the vehicle.

[0069] For example, referring to Figure 1B, during the driving of the vehicle 10, multiple images can be collected by multiple sensors at the same moment through a timestamp synchronization mechanism or a hardware synchronization mechanism. The multiple images can be, for example, image data of different perspectives of the vehicle 10 collected by a camera, and / or point cloud data of each obstacle collected by a radar, etc. It can be understood that in some embodiments of the present application, the multiple images can include images of actual obstacles around the vehicle 10. For example, the actual obstacles can be the first vehicle 01, the pedestrian 02, the building 03, the second vehicle 04, and the building 05, etc.

[0070] After detecting a request to generate a first environmental image, the in-vehicle device can obtain multiple images of different perspectives collected by multiple sensors on the vehicle.

[0071] S302, sample the features of the multiple images according to the first position data of each feature point in the first spatial data to obtain the first features of N first obstacle models.

[0072] In some embodiments of the present application, after the in-vehicle device obtains multiple images, it can project the first spatial data of the N first obstacle models into the coordinate space of the multiple images, and then collect the features of the first positions of each feature point of the N first obstacle models in the multiple images as the first features. It can be understood that the coordinate space of the multiple images can be determined according to the camera on the vehicle. Different positions of the camera on the vehicle correspond to different coordinate spaces of the corresponding images.

[0073] In some embodiments of the present application, the images captured by the camera on the vehicle 10 may also include some distorted images. Therefore, during the process of projecting the N first obstacle models based on the first spatial data into the spatial data of the multiple images, the influence of image distortion can be adjusted.

[0074] Exemplarily, the spatial dimensions, positions, and rotation angles of the first obstacle model stored in the in-vehicle device are the distortion errors from the undistorted image mapped to the distorted image. After the first spatial data of the first obstacle model is projected into the coordinate space of the distorted image, distortion error adjustment is required.

[0075] For example, Figure 4A According to some embodiments of the present application, a schematic diagram of obtaining the distortion error is shown.

[0076] As Figure 4A shown, the in-vehicle device can first obtain the pixel coordinates of the undistorted image, and then determine the pixel coordinates of the distorted image according to the distortion calculation formula, where the distortion calculation formula is mainly determined based on the physical characteristics and imaging principle of the camera lens. The in-vehicle device can determine the distortion calculation formula according to the camera lenses of each camera on the vehicle.

[0077] The in-vehicle device can obtain a distortion offset table by subtracting the pixel coordinates of the distorted image from the pixel coordinates of the undistorted image. The in-vehicle device stores the distortion offset table, and then the in-vehicle device can look up the distortion error from the distortion offset table, thereby reducing the calculation process of the in-vehicle device and reducing the computing power requirement of the in-vehicle device.

[0078] It can be understood that in some other embodiments, the distortion offset table can also be determined by other devices and then stored in the in-vehicle device.

[0079] Figure 4B According to some embodiments of the present application, a schematic diagram of obtaining the first features of N first obstacle models is shown.

[0080] As Figure 4B shown, after the in-vehicle device obtains the first spatial data of N first obstacle models, it can determine the three-dimensional position of the first obstacle model projected into the space of multiple undistorted images. The three-dimensional position includes the spatial dimensions, positions, and rotation angles of the first obstacle model in the spatial data of multiple images, that is, the positions of the respective feature points of the first obstacle model in the space of multiple images. Then the in-vehicle device queries the distortion offset table to obtain the distortion error, and the in-vehicle device can adjust the three-dimensional position of the first obstacle model through the distortion error to obtain the three-dimensional position of the first obstacle in the space of multiple distorted images. Then, based on the three-dimensional position of the first obstacle model in the space of multiple distorted images, the features of the multiple distorted images are collected, thereby obtaining the first features of N first obstacle models.

[0081] It can be understood that through the distortion error in the distortion offset table, the in-vehicle device does not need to calculate the distortion error, thereby reducing the computing power requirement in the process of the in-vehicle device generating the first environmental image. So that the in-vehicle device can be better arranged on the vehicle end platform.

[0082] It can be understood that hereinafter, the process of projecting the target spatial data of the target obstacle into the coordinate space of multiple images can also determine the three-dimensional position of the target obstacle projected into the coordinate space of multiple images by querying the distortion offset map. That is to say, in the process of projecting various obstacle models into the coordinate space of multiple images, the three-dimensional position of the corresponding obstacle in the distorted image coordinate space can be determined by querying the distortion offset map, thereby reducing the calculation amount of the in-vehicle device.

[0083] In some embodiments of the present application, in order to improve the computing speed of in-vehicle devices, some in-vehicle devices only support the input of data in the form of low-precision 8-bit integers (int8). It can be understood that since int8 quantization uses fewer bits to represent data, it has significant advantages in terms of storage requirements, and this saving in storage space can significantly reduce the deployment cost. In some embodiments of the present application, the distortion error in the distortion offset table can be stored in the in-vehicle device in the form of int8 quantization. However, the accuracy of the distortion error quantized by int8 is relatively low. In order to ensure the accuracy of the distortion error, in some embodiments of the application, a distortion error quantized by a high-precision 16-bit integer (int16) can be represented by two int8-form data.

[0084] For example, Figure 5 According to some embodiments of the present application, a schematic diagram of a process for processing distortion errors is shown.

[0085] As Figure 5 shown, for a distortion error (int16), where int16 represents data in the form of a high-precision 16-bit integer, it can be represented by a front (int8) data and a rear (int8) data, and int8 represents data in the form of a low-precision 8-bit integer. Among them, front = distortion error (int16) / 2 8 rounded down, rear = distortion error (int16) - front × 2 8 -2 7 . It can be understood that since the first bit of the data in the int8 form is the sign bit, the value of the data in the int8 form is only the last 7 bits, so 2 needs to be subtracted 7 . During the calculation, distortion error (int16) = front × 2 8 + rear + 2 7 .

[0086] It can be understood that the data quantized by int16 is a 16-bit binary number. Since the first bit is the sign bit, the decimal value range that int16 can represent is -32768 to 32767. That is to say, the decimal value X that the binary number of the distortion error after int16 quantization can represent is in the range of -32768 to 32767. By dividing X by 2 8 and rounding down, the value range obtained is 0 to 127, which is exactly the value that an 8-bit binary number can represent (the highest bit in the 8-bit binary number is the sign bit, so the decimal number range represented by the 8-bit binary number is -128 to 127). That is to say, front is in the int8 form.

[0087] However, since front is X divided by 28 is obtained by rounding down. Therefore, front × 2 8 may also be less than X. The part of X - front × 2 can be represented by rear 8 And X - front × 2 8 has a maximum value of 255, which is greater than the value that an int8 can represent. Therefore, rear also needs to subtract a 2 7 . That is to say, the range of X - front × 2 8 - 2 7 is also from - 128 to 127. Therefore, rear can be represented by an int8. For example, for the integer 255, when represented by int16, it is 0000000001111111. The decimal number of front is 0, that is, 255 / 256 = 0.996, and after rounding down, it is 0. The int8 form of front is represented as 00000000. Then the decimal number of rear is 255 - 0 × 2 8 - 2 7 = 127. The int8 representation of rear is 01111111.

[0088] Similarly, for the integer - 255, when represented by int16, it is 1000000001111111. The decimal number of front is - 1, that is, - 255 / 256 = - 0.996, and after rounding down, it is - 1. The int8 form of front is represented as 10000001. Then the decimal number of rear is - 255 - (- 1) × 2 8 - 2 7 = - 127. The int8 form of rear is represented as 11111111.

[0089] In the above way, a int16 - quantized distortion error data can be represented by two int8 - form data, which can not only improve the operation speed of in - vehicle devices, but also improve the operation accuracy of in - vehicle devices.

[0090] It can be understood that since the number of the first obstacle models is N, while the actual number of obstacles may be larger or smaller than N. If the actual number of obstacles is smaller than N, the first spatial data of the redundant first obstacle models (i.e., the first obstacle models that are not the closest to each actual obstacle) may not need to be adjusted, and the similarity between the redundant first obstacle models and the images of the actual obstacles in the front view image is 0, that is to say, there is no actual obstacle corresponding to the first obstacle model. If the actual number of obstacles is larger than N, the vehicle-mounted device can select the actual obstacles that are close to the vehicle from the first obstacle models according to the neural network and adjust the first spatial data of the corresponding first obstacle models to ensure that the images of the obstacles closest to the vehicle in the first environmental image are preferentially generated.

[0091] For example, in the embodiments of the present application, the vehicle-mounted device determines the actual obstacles only in the space where the distance from the vehicle 10 is within 200m to 100m in the front view image. For example, the space within 150m in front of the vehicle 10 can be selected. The spaces for determining the actual obstacles on the left and right sides and above and below the vehicle 10 can be within a distance of 30m to 80m from the vehicle 10. For example, 50m can be preferably selected. The space for determining the actual obstacles behind the vehicle 10 can be within a distance of 80m to 120m from the vehicle 10. For example, 100m can be preferably selected.

[0092] S303, select N1 first obstacle models from the N first obstacle models according to multiple images and the first feature.

[0093] Exemplarily, in some embodiments of the present application, the vehicle-mounted device can select N1 first obstacle models from the N first obstacle models, and the similarity between the first features of the N1 first obstacle models and the features of the actual obstacles is greater than the similarity between the first features of the remaining first obstacle models and the features of the actual obstacles. Wherein, 0 < N1 ≤ N. For example, in the embodiments of the present application, N1 is half of N. That is to say, when N is 384, N1 is 192. In other embodiments, N1 can also be other values. For example, N1 = N / 3, or N1 = N / 4, N1 = 200, N = 100, etc.

[0094] In some embodiments, N1 first obstacle models with the highest similarity to the features of the corresponding actual obstacles can be selected from the N first obstacle models. However, this requires sorting the similarities of the N first obstacle models, and the sorting process consumes a lot of computing power. Therefore, in other embodiments, the first obstacle models with high similarity between the corresponding first features and the features of the actual obstacles can be selected from between every two adjacent first obstacle models according to the preset arrangement positions of the N first obstacle models as the obstacle models among the N1 first obstacle models.

[0095] For example, Figure 6 Some embodiments according to the present application illustrate a process of selecting N1 first obstacle models.

[0096] As Figure 6 shown, a first obstacle model with a relatively high similarity between the first feature and the features of the actual obstacles in multiple images can be selected from every two adjacent first obstacle models. It can be understood that since the positions of the N first obstacle models are initially arranged, there is no need to queue the N first obstacle models during the process of selecting N1 first obstacle models, thereby improving the efficiency of the vehicle-mounted device in generating the first environmental image. In some other embodiments, the similarities between the N first obstacle models and the features of the actual obstacles in the image can also be sorted, and then the top N1 first obstacle models with the highest similarities can be selected. However, this method requires sorting the similarities between the N first obstacle models and the actual obstacles in the image, and the sorting process will consume a relatively high computing power of the vehicle-mounted device.

[0097] S304. Optimize the first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models to obtain the third spatial data of the N1 first obstacle models and the N2 historical obstacle models.

[0098] Exemplarily, in some embodiments of the present application, N = N1 + N2, that is to say, the obstacle models for generating the first environmental image are always N (hereinafter, the N target obstacle models are used to represent the obstacle models in the N1 first obstacle models and the N2 historical obstacle models). It can be understood that the third spatial data includes the spatial data obtained by optimizing the first spatial data of the N1 first obstacle models multiple times, and the spatial data obtained by optimizing the second spatial data of the N2 historical obstacle models multiple times. After the first spatial data and the second spatial data are optimized multiple times, the positions of the corresponding M feature points will also change. For example, each third spatial data includes the second position data of the M feature points corresponding to the target obstacle model.

[0099] Exemplarily, in the embodiments of the present application, the vehicle-mounted device can also obtain historical obstacle models as a reference during the process of generating the first environmental data of the current frame. For example, the vehicle-mounted device includes the third features of the N third obstacle models corresponding to the historically generated second environmental image, and the fourth spatial data of the N third obstacle models, where the similarity between the third features and the image features of the actual obstacles corresponding to the second environmental image is greater than a set threshold, and the third features are sampled from the image features of the actual obstacles corresponding to the second environmental image by the fourth spatial data of the N third obstacle models.

[0100] The in-vehicle device can select N2 third features with the highest feature similarity to the image of the actual obstacle corresponding to the second environmental image from the third features of the N third obstacle models.

[0101] Then, the in-vehicle device can use the fourth spatial data of the third obstacle models corresponding to the N2 third features as the second spatial data of the N2 historical obstacle models. And hereinafter, the first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models are referred to as target spatial data. That is to say, the third spatial data of the target obstacle models can be obtained after optimizing the target spatial data of the N target obstacle models multiple times. The third spatial data includes the second position data corresponding to M feature points of the target obstacle models.

[0102] In some embodiments of the present application, the second environmental image can be the environmental image generated by the vehicle 10 in the previous frame of the current frame. After the in-vehicle device determines that each third obstacle model corresponds to the actual obstacle image in the previous frame (i.e., the actual obstacle image corresponding to the second environmental image), it can obtain the speed of each third obstacle model relative to the vehicle 10 and calculate the position of each third obstacle model in the current frame according to the time interval between each frame. That is to say, the fourth spatial data can also be the spatial data adjusted by calculating the speed of each third obstacle model relative to the vehicle 10 and the time interval relative to the current frame.

[0103] In some embodiments of the present application, the in-vehicle device can, for example, adjust the positions of multiple feature points of the N1 first obstacle models and the positions of multiple feature points of the N2 historical obstacle models at least once according to the spatial data of the actual obstacle in multiple images, so as to obtain the third spatial data.

[0104] It can be understood that after the in-vehicle device obtains the target spatial data of the target obstacle models, it can adjust the target spatial data of the target obstacle models multiple times (for example, adjust three times) according to the spatial data of the actual obstacle models in multiple images.

[0105] For example, taking the front view camera as an example, the in-vehicle device can project the target spatial data of the N target obstacle models onto the image collected by the front view camera (hereinafter referred to as the front view image), and collect the target features of the N target obstacle models in the front view camera. It can be understood that since the N target obstacle models include N2 historical obstacle models, it is necessary to resample the N target obstacle models to obtain the target features.

[0106] The in-vehicle device can use the neural network model to determine each actual obstacle in the front view image, refer to Figure 1BThe actual obstacles in the front view image may include a pedestrian 02, a second vehicle 04, etc. After the vehicle-mounted device determines the actual obstacles, it may correspond the target obstacle model closest to the pedestrian 02 to the pedestrian 02 through a neural network model. Similarly, the target obstacle model closest to the second vehicle 04 may be corresponded to the second vehicle 04.

[0107] In some embodiments, the vehicle-mounted device may adjust the target space data of the target obstacle model corresponding to the pedestrian 02 through the spatial dimension, position, rotation angle, etc. of the pedestrian 02, so as to update the target space data of the target obstacle. The adjustment process is, for example, to obtain the offset of the spatial dimension, position, and rotation angle between the pedestrian 02 and the first obstacle corresponding to it, then add the offset to the spatial dimension, position, and rotation angle of the target obstacle model, and then adjust the first positions of multiple feature points of the target obstacle model according to the added spatial dimension, position, and rotation angle of the target obstacle model, so as to update the target space data of the target obstacle model corresponding to the pedestrian 02, that is, the position information of the corresponding M feature points in the target space data will also be updated.

[0108] Similarly, through the above method, each actual obstacle in multiple pictures can be corresponded to N target obstacle models, and the target space data of the N target obstacle models can be updated. After multiple (such as 2 times, 3 times, 4 times, 5 times, etc.) optimizations as above, the target space data of the N target obstacle models can be updated to the third space data.

[0109] S305, sample the features of multiple images according to the third space data to obtain the second features of the target obstacle model, and generate a first environmental image according to the second features.

[0110] Exemplarily, after determining the third space data of the target obstacle model, the features of multiple images may be sampled according to the second position data of each feature point in the third space data to obtain the second features of the target obstacle model. The sampling process may refer to the process in S302 where the first obstacle model samples the image features of multiple images according to the first space data to obtain the first features. It can be understood that the similarity between the second features of the target obstacle model and the features of the actual obstacle is greater than a set threshold. The set threshold may be any value between 60% - 95%, for example.

[0111] Exemplarily, in some embodiments of the present application, after the vehicle-mounted device determines the second features of the target obstacle, it may generate a first environmental image according to the second features.

[0112] In some embodiments, since some target obstacle models do not correspond to the actual obstacle models, the similarity between the second features of these target obstacle models and the features of the corresponding actual obstacles in multiple images is small and may be less than a set threshold. However, the vehicle-mounted device can still retain these target obstacle models. During the process of generating the first environmental image based on the second features, the target obstacle models with the similarity between the second features and the features of the actual obstacles less than the set threshold can be discarded. That is to say, only the second features of the target obstacle models with the similarity between the second features and the features of the actual obstacles exceeding the set threshold are used to generate the first environmental image.

[0113] It can be understood that since the second features are obtained by sampling the features of multiple images through M feature points of the target obstacle, the number of M can be determined in advance, and as long as the value of M can represent the spatial size, position, and rotation angle of the corresponding target obstacle, the value of M can be set to be small. For example, in the embodiments of the present application, it can be 13, and in some other embodiments, it can be 7 or 19, etc. In this way, the data volume of the second features can be greatly reduced, thereby reducing the computing power requirements of the vehicle-mounted device. Moreover, the first environmental image generated by the vehicle-mounted device based on the second features has three-dimensional attributes. For example, the spatial size, position, and rotation angle of each third obstacle can be determined, thus avoiding the loss of obstacle information (such as the loss of the height information of the obstacle in the BEV space). Therefore, it has a better visual effect compared with the BEV map. Also, when the vehicle 10 is driving on a slope, the information of the obstacles on the slope can also be displayed in the first environmental image, thereby improving the accuracy of the vehicle-mounted device in detecting obstacles.

[0114] Next, the process of the vehicle-mounted device generating the first environmental image will be introduced.

[0115] For example, Figure 7 According to some embodiments of the present application, a schematic diagram of a vehicle-mounted device generating a first environmental image is shown.

[0116] Exemplarily, after the vehicle-mounted device obtains multiple images of the current frame, it can map the first spatial data of N first obstacle models to the coordinate space of the multiple images and sample in the multiple images based on the first positions of M feature points, so as to obtain the first features of the N first obstacle models. The process of obtaining the first features of the N first obstacle models can refer to the process of S302.

[0117] After the vehicle-mounted device obtains the first features of N first obstacle models, it can select N1 first obstacle models with the highest similarity between the first features and the actual obstacles. It can be understood that the actual obstacles are the obstacles in multiple images. After selecting the N1 first obstacle models, the first spatial data and the first features of the N1 first obstacle models can be determined. The process of determining the first spatial data and the first features of the N1 first obstacle models can refer to the process of S303. Then, the vehicle-mounted device can also obtain the third features and the second spatial data of N2 historical obstacle models, and merge the second spatial data of the N2 historical obstacle models and the first spatial data of the N1 first obstacle models into the target spatial data of N target obstacle models, and merge the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models into the target features of the N target obstacle models, so as to obtain the target spatial data and the target features of the N target obstacle models. Among them, the process of obtaining the second spatial data and the third features of the N2 historical obstacles can refer to the process of S304.

[0118] In this embodiment, the process of the vehicle-mounted device optimizing the target spatial data of the target obstacle model is, for example, sampling the features of multiple images based on the target spatial data of the N target obstacle models to update the target features of the N target obstacle models. It can be understood that since the target obstacle model includes the second spatial data and the third features of the N2 historical obstacle models, therefore, it is necessary to update the features of the N target obstacle models, and the original target features of the N target obstacle models can also determine the features of more important target obstacles through the self-attention module.

[0119] For example, the target features of the N target obstacle models (for example, the target features that have not been resampled by the target spatial data) can be subjected to inter-frame interaction to obtain the features after inter-frame interaction. The inter-frame interaction is, for example, the interaction between the first features of the N1 first obstacle models in the current frame and the third features of the N2 historical obstacle models in the previous frame in the target obstacle model.

[0120] For example, Figure 8A Some embodiments according to the present application show a schematic diagram of inter-frame interaction.

[0121] As Figure 8A shown, a self-attention module is configured in the vehicle-mounted device, and the self-attention module can perform self-attention optimization through a transformer.

[0122] For example, the process of inter-frame interaction can be to multiply the target features of N target obstacle models, that is, the first features of N1 first obstacles and the third features of N2 historical obstacles, with the query weight matrix in the transformer model to obtain query features.

[0123] Then multiply the third features of N2 historical obstacle models with the key weight matrix in the transformer model to obtain key features.

[0124] Similarly, multiply the third features of N2 historical obstacle models with the value weight matrix in the transformer model to obtain value features.

[0125] Exemplarily, in the transformer model, key features are used to match with query features to determine which feature pairs are most important for generating the features of the first environmental image. And value features contain the actual information corresponding to the key features, and these information are weighted and summed in the attention mechanism to generate the final features of the first environmental image.

[0126] Exemplarily, the query weight matrix, key weight matrix, and value weight matrix are trainable parameters of the transformer model, and they are updated and optimized through the backpropagation algorithm during the model training process. Therefore, after training the corresponding weight matrices, only need to multiply the input feature data with the corresponding weight matrices to obtain the corresponding query features, key features, and value features.

[0127] After obtaining the query features, key features, and value features, the input features can be optimized based on the self-attention module to obtain the corresponding output. The optimization process of the self-attention module is described in detail below. For example, the output result of inter-frame interaction can be called inter-frame interaction features. It can be understood that the N inter-frame interaction features also include the inter-frame interaction features of N1 first obstacle models and the inter-frame interaction features of N2 historical obstacle models.

[0128] After the in-vehicle device completes the inter-frame interaction, it can perform intra-frame interaction on the features after inter-frame interaction of N target obstacle models to obtain the features after intra-frame interaction.

[0129] For example, Figure 8B Some embodiments according to the present application show a schematic diagram of inter-frame interaction.

[0130] As Figure 8BAs shown, the input of the intra-frame interaction can be the features after the inter-frame interaction of N target obstacle models. The interaction process is, for example, to input the features after the inter-frame interaction of N target obstacle models into the query weight matrix, key weight matrix, and value weight matrix of the transformer model respectively to obtain the corresponding query features, key features, and value features. It can be understood that the features after the inter-frame interaction of N target obstacle models include the features after the inter-frame interaction of N1 first obstacle models and the features after the inter-frame interaction of N2 historical obstacle models. Then, the vehicle-mounted device optimizes the query features, key features, and value features based on the self-attention module to obtain the corresponding output. The optimization process of the self-attention module is described in detail below.

[0131] It can be understood that after the intra-frame interaction is completed, the intra-frame interaction features of N target obstacle models can be obtained. Then, the vehicle-mounted device can perform feature fusion on the target features of the resampled target obstacle models and the features after the intra-frame interaction of the target obstacle models to obtain the updated target features. For example, in the embodiments of the present application, the target features of the resampled N target obstacle models can be added to the features after the intra-frame interaction of the corresponding N target obstacle models to obtain the new features of the N target obstacle models, and the new features can be used as the updated target features of the N target obstacle models. It can be understood that since the target features of N target obstacle models include the first features of N1 first obstacle models and the third features of N2 historical obstacle models, after the target features of N target obstacle models are updated, the corresponding first features of N1 first obstacle models and the third features of N2 historical obstacle models will also be updated.

[0132] Next, the optimization process of the self-attention module is introduced.

[0133] For example, Figure 8C According to some embodiments of the present application, a schematic diagram of the optimization process of a self-attention module is shown.

[0134] As Figure 8C shown, after the vehicle-mounted device obtains the corresponding query features, key features, and value features, matrix multiplication operations (such as dot product operations) can be performed on each query feature and all key features to obtain attention scores. These scores reflect the similarity or correlation degree between the query features and each key feature.

[0135] Then, the in-vehicle device normalizes the attention scores. For example, the attention scores can be normalized through the softmax function so that the sum of all attention scores is 1. In this way, each attention score represents the weight of the corresponding value feature.

[0136] Next, the normalized attention scores are used as weights and assigned to the corresponding value features, that is, the attention scores are multiplied with each value feature as weights through matrix multiplication. This means that the value features corresponding to the key features with higher similarity to the query feature will obtain greater weights. Then, the results of the matrix multiplication of all value features and their corresponding weights are summed up, that is, the self-attention scores are used as weights to perform weighted summation on the query feature. This weighted sum is the output of the self-attention mechanism.

[0137] It can be understood that the optimization process of the self-attention module in the above inter-frame interaction and intra-frame interaction processes can refer to Figure 8C the process.

[0138] After obtaining the updated target features of the N target obstacle models, the in-vehicle device can also readjust the target space data of the N target obstacle models based on the space data of multiple images. The adjustment process refers to the process of S304, so as to update the target space data of the N target obstacle models. It can be understood that at this point, both the target space data and the target feature data of the N target obstacle models have been updated. To ensure the accuracy of the generated first environmental image, the in-vehicle device can also update the target space data of the N target obstacle models based on the neural network multiple times to obtain the third space data of the N target obstacle models. Similarly, the target features of the N target obstacle models can also be optimized through multiple loops. For example, in some embodiments of the present application, the third space data is obtained by performing 2-loop optimization and adjustment on the target space data.

[0139] Exemplarily, the in-vehicle device can perform multiple-loop optimization on the updated target space data and target features of the N target obstacle models. The process of the multiple-loop optimization is, for example, the above process of performing inter-frame interaction and intra-frame interaction on the target features, projecting the target space data into the coordinate space of multiple images to resample the N target obstacle models to obtain the resampled features of the N target obstacle models, adding and fusing the resampled features of the N target obstacle models with the features after intra-frame interaction to update the target features, and further optimizing the target space data based on the updated target features to update the target space data.

[0140] In some embodiments of the present application, in the last cycle of the vehicle-mounted device, the feature obtained after intra-frame interaction may be the fourth feature, and the feature obtained by resampling the target obstacle model by projecting the target spatial data into the coordinate spaces of multiple images may be the second feature. It can be understood that the target spatial data is the updated spatial data in the previous cycle of the vehicle-mounted device, and this target spatial data can be used as the third spatial data. That is to say, the second feature is obtained by sampling the third spatial data in the coordinate spaces of multiple images. In the last cycle, the feature obtained by adding and fusing the second feature and the fourth feature is used as the fifth feature, and the vehicle-mounted device can generate the first environmental image based on the fifth feature.

[0141] Through the above process of loop optimization, the accuracy of generating the first environmental image can be improved.

[0142] The vehicle-mounted device involved in each of the above embodiments will be introduced below.

[0143] For example, Figure 9 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device 100 is shown.

[0144] The vehicle-mounted device 100 can be used to implement the image processing method provided in the foregoing embodiments.

[0145] As Figure 9 shown, the vehicle-mounted device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output device 105. Among them: The processor 101 may include one or more processing units. For example, it may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA), a neural-network processing unit (NPU), etc. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be used to optimize the neural network model to be run. For example, in some embodiments of the present application, the neural network model may optimize the spatial data of the first obstacle model, and the NPU may be used to run the neural network model to be run.

[0146] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 may be used to store the data provided by the foregoing different services, such as sensor data, image data, or video data, etc., and may also be used to store the instructions of the image processing method provided by the foregoing embodiments.

[0147] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 103 may be used to store instructions of the image processing method provided in the foregoing embodiments, etc.

[0148] Specifically, the system memory 102 and the non-volatile memory 103 may respectively include: a temporary copy and a permanent copy of the instructions 107. The instructions 107 may include: when executed by at least one of the processors 101, enabling the vehicle-mounted device 100 to implement the image processing method provided in the various embodiments of the present application.

[0149] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the vehicle-mounted device 100, and then communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the vehicle-mounted device 100. For example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the vehicle-mounted device 100 may communicate with other devices through the communication interface 104. For example, the vehicle-mounted device 100 may obtain corresponding data from other devices through the communication interface 104.

[0150] The input / output device 105 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. The user may interact with the vehicle-mounted device 100 through the input / output device 105.

[0151] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface for other modules of the vehicle-mounted device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide an interface connected to the system memory 102 and the non-volatile memory 103.

[0152] In some embodiments, at least one of the processors 101 may be logically packaged with one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).

[0153] It can be understood that Figure 9 The structure of the in-vehicle device 100 shown is only an example. In other embodiments, the in-vehicle device 100 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0154] It can be understood that the in-vehicle device 100 can be any device configured on a vehicle, including but not limited to mobile phones, in-vehicle computers, terminals in self-driving, wireless terminals in transportation safety, terminals in a smart city, and so on.

[0155] The embodiments of the present application also provide a program product. When executed on an in-vehicle device, the program product can enable the in-vehicle device to implement the image processing methods provided in the foregoing embodiments.

[0156] The embodiments of the present application also provide a readable storage medium. One or more programs are stored in the readable storage medium. When the one or more programs are executed by the in-vehicle device, the in-vehicle device is enabled to implement the image processing methods provided in the foregoing embodiments.

[0157] The embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0158] The program code can be applied to the input instructions to execute the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.

[0159] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.

[0160] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to, floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical discs, read only memory (ROM), random-access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms using the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a machine (e.g., computer) readable form.

[0161] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a different manner and / or order than shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0162] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application. This does not mean that there are no other units / modules in the above device embodiments.

[0163] It should be noted that in the examples and the description of the present patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one" does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0164] Although the present application has been illustrated and described by reference to certain preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the scope of the present application.

Claims

1. An image processing method, used for an on-board device on a vehicle, characterized in that: The vehicle-mounted device stores first spatial data of N first obstacle models, each of the first spatial data of the first obstacle model includes first position data of M feature points of the first obstacle model; the method includes: In response to a request to generate a first environment image of the vehicle, a plurality of images are acquired, where the plurality of images are images of different perspectives acquired by a plurality of sensors of the vehicle, and the plurality of images include images of actual obstacles around the vehicle; Sampling the features of the plurality of images according to the first position data of each feature point in the first spatial data to obtain first features of the N first obstacle models; selecting N1 first obstacle models from the N first obstacle models according to the multiple images, wherein similarities between first features of the N1 first obstacle models and features of the actual obstacle are greater than similarities between first features of remaining first obstacle models and features of the actual obstacle; The first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models are optimized to obtain third spatial data of the N1 first obstacle models and the N2 historical obstacle models; wherein N=N1+N2, and each of the third spatial data includes second position data of M feature points of the corresponding obstacle model, and a similarity between second features obtained by sampling features of the multiple images according to the second position data of each feature point in the third spatial data and features of the actual obstacle is greater than a set threshold; The first environment image is generated according to the second feature.

2. The image processing method according to claim 1, characterized in that: The M feature points include M points on three coordinate axes that are orthogonal to each other in the coordinate space of the first obstacle model; The first position data of the M feature points are used to represent the initial spatial size, position and rotation angle of the first obstacle model.

3. The image processing method according to claim 1, characterized in that: The selecting N1 first obstacle models from the N first obstacle models according to the multiple images includes: According to the preset arrangement positions of the N first obstacle models, a first obstacle model having a high similarity between the first feature and the feature of the actual obstacle is selected from between every two adjacent first obstacle models as an obstacle model among the N1 first obstacle models.

4. The image processing method according to claim 1, characterized in that: The vehicle-mounted device includes third features of N third obstacle models corresponding to the second environment image generated historically, and fourth spatial data of the N third obstacle models, wherein the similarity between the third features and the image features of the actual obstacles corresponding to the second environment image is greater than the set threshold, and the third features are obtained by sampling the fourth spatial data of the N third obstacle models from the image features of the actual obstacles corresponding to the second environment image; The second spatial data of the N2 historical obstacle models are determined in the following manner: Selecting N2 third features having the highest feature similarity to the image of the actual obstacle corresponding to the second environment image from the third features of the N third obstacle models; The fourth spatial data of the third obstacle model corresponding to the N2 third features is used as the second spatial data of the N2 historical obstacle models.

5. The image processing method according to claim 1, characterized in that: The optimizing the first spatial data of the N1 first obstacle models and the second spatial data of the N2 historical obstacle models to obtain the third spatial data of the N1 first obstacle models and the N2 historical obstacle models includes: According to the spatial data of the actual obstacle in the multiple images, the positions of the M feature points of the N1 first obstacle models and the positions of the M feature points of the N2 historical obstacle models are adjusted at least once to obtain third spatial data.

6. The image processing method according to claim 4, characterized in that: The step of generating the first environment image according to the second feature includes: Performing at least one self-attention adjustment on the first features of the N1 first obstacle models and the third features of the N2 historical obstacle models to obtain a fourth feature; using a fifth feature obtained by fusing the fourth feature with the second feature as a feature for generating the first environment image; Among them, the similarity between the fifth feature and the features of the multiple images is higher than the similarity between the second feature and the features of the multiple images.

7. The image processing method according to claim 1, characterized in that: The multiple images include distorted images; The vehicle-mounted device stores the distortion error of the spatial size, position and rotation angle of the first obstacle model mapped from the undistorted image to the distorted image; The third spatial data includes data obtained by projecting the first spatial data of the first obstacle model into the coordinate space of the plurality of images after the distortion error is adjusted.

8. A vehicle-mounted device, characterized in that: include: A memory for storing instructions; At least one processor is used to execute the instructions so that the vehicle-mounted device implements the image processing method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the image processing method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product is executed on a device, the device is enabled to execute the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Obstacle detection method and device

    CN112712009A

  • Obstacle detection method and device, vehicle and storage medium

    CN113537047A

  • Obstacle detection method and device, vehicle-mounted terminal and storage medium

    CN118736523A

  • Image fusion segmentation method and device, equipment and storage medium

    CN119229108A

  • Autonomous platform guidance systems with auxiliary sensors and obstacle avoidance

    US10571926B1