Mirror data processing method and apparatus
By using mirror data processing methods and 3D scene reconstruction technology to project mirror image information onto non-mirror areas, the perception problem of intelligent driving in blind spots is solved, and more accurate path planning and safety improvement are achieved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- YINWANG INTELLIGENT TECHNOLOGIES CO LTD
- Filing Date
- 2025-01-24
- Publication Date
- 2026-07-30
AI Technical Summary
Existing intelligent driving technologies struggle to accurately perceive the remote environment and predict potential traffic risks in scenarios with blind spots, especially in mountainous areas with curves or narrow intersections, leading to an increased risk of accidents.
By acquiring the perception data of the mirror area, the information in the mirror image is projected onto the corresponding 3D scene of the non-mirror using 3D scene reconstruction technology, thereby reconstructing 3D scene information with a larger visual range and rendering it to improve perception capabilities.
It enables accurate representation of the relative positional relationship between objects and the vehicle within the blind spot, improves the vehicle's perception capability within the blind spot, and allows for more accurate driving path planning, thereby enhancing driving safety.
Smart Images

Figure CN2025074762_30072026_PF_FP_ABST
Abstract
Description
A method and apparatus for processing mirror data Technical Field
[0001] This application relates to the field of intelligent vehicles, and more particularly to a mirror data processing method and apparatus. Background Technology
[0002] With the development of intelligent driving functions in smart cars, these vehicles are gradually acquiring the ability to drive intelligently in various complex environments. However, current smart cars still face significant challenges in scenarios with blind spots, such as mountain curves or narrow intersections where visual blind spots exist. In these areas, intelligent driving functions still require considerable adjustments. For instance, in mountainous road conditions with curves or narrow intersections, the vehicle's intelligent driving system often cannot clearly perceive the road conditions after a turn, lacks prediction of the remote environment, and cannot accurately predict potential traffic risks, leading to an increased risk of accidents. Although existing intelligent driving solutions provide some obstacle detection and avoidance capabilities through high-precision maps, radar perception, and visual recognition, the perception capabilities they rely on have not completely overcome the limitations of perspective and environment, making it difficult to achieve accurate dynamic prediction and long-term path planning in curves or narrow intersections.
[0003] Convex mirrors placed in the environment are an important source of visual data in scenarios with blind spots. However, some existing solutions have limited perception capabilities of convex mirrors, making it difficult to utilize them to assist intelligent driving functions. Therefore, how to more effectively utilize information from convex mirrors has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a mirror data processing method and apparatus for reconstructing a three-dimensional scene from information contained in a mirror image and projecting it onto a non-mirror corresponding three-dimensional scene, thereby obtaining three-dimensional scene information with a larger visual range.
[0005] In view of the above, in a first aspect, this application provides a mirror data processing method, comprising: acquiring sensing data, the sensing data including data collected by at least one sensor, which may include a first image collected by an image sensor; if the first image includes a mirror region, acquiring first spatial information from the mirror region, the first spatial information including information of elements included in the mirror region, the mirror region including areas with mirror reflection, such as mirror objects such as plane mirrors or convex mirrors in the scene where the sensing data is collected; and subsequently projecting the first spatial information into a first space according to the mirror coordinate system corresponding to the mirror region to obtain second spatial information, the first space being the coordinate system corresponding to the image sensor, the second spatial information including information of elements in the mirror region in the first space.
[0006] In this embodiment, for images captured in a scene containing mirrors, features of the mirrored areas are extracted and projected onto the space corresponding to the vehicle. This allows the vehicle to perceive obstacles in the environment through the mirrors, reconstruct a 3D scene that may include blind spots, and project it onto the space corresponding to the non-mirrored areas. This enables objects in the mirrored areas and objects in the non-mirrored areas to be reconstructed in the same space, which can more accurately express the relative positional relationships between elements in the space. This can improve the perception range of the device, especially the perception capability for blind spot areas.
[0007] In one possible implementation, the aforementioned method further includes: rendering based on the second spatial information to obtain a rendered image, which includes a scene image after the scene information in the mirror is projected onto the space of the non-mirror area; and displaying the rendered image on a display interface. In this embodiment of the application, after obtaining the second spatial information including the mirror image, rendering can be performed to obtain a rendered image, and the rendered image can be displayed to the user, so that the user can observe the scene information in the blind spot through the displayed interface.
[0008] In one possible implementation, the aforementioned method is applied to a vehicle, further comprising: planning a driving path for the vehicle based on second spatial information, such as determining the spatial geometric positional relationship between other objects in the vehicle's scene based on the second spatial information, and planning a drivable path that can avoid obstacles based on this spatial geometric positional relationship. In this embodiment, for scenes with blind spots, the scene information contained within a mirror can be projected onto the space corresponding to the non-mirror area, thereby obtaining a representation of the relative positional relationship between objects within the blind spot and the vehicle in the same space. This improves the vehicle's perception ability of elements within the blind spot, enabling more accurate driving path planning and enhancing vehicle driving safety.
[0009] In one possible implementation, the aforementioned method further includes: identifying whether the first image includes a mirror region based on the perceptual data. In this application embodiment, the method can also identify whether the first image includes a mirror region based on the perceptual data, so that if the first image includes a mirror region, first spatial information can be obtained from the mirror region, allowing for more accurate processing of the mirror region.
[0010] In one possible implementation, the aforementioned process of identifying whether a first image includes a mirrored region based on perceptual data may include: extracting perceptual features from the perceptual data; identifying whether a first image includes a mirror based on the perceptual features; and, if a mirrored region is included in the first image, detecting and identifying the mirrored region in the first image.
[0011] In this embodiment of the application, features can be extracted from the perceptual data, and the presence of a mirror region in the first image can be identified based on the perceptual features. When a mirror region exists in the first image, the range of the mirror region can be identified so that the mirror region can be processed in a targeted manner in the future.
[0012] In one possible implementation, the aforementioned acquisition of first spatial information from the mirror region may include: if the first image includes a mirror region, performing depth estimation on the mirror region to obtain a depth map; projecting the pixels of the mirror region onto a three-dimensional space based on the depth map to obtain a mirror region point cloud; and acquiring first spatial information based on the mirror region point cloud. In this embodiment, if the first image includes a mirror region, depth estimation can be performed on each pixel within the mirror region to obtain a depth map representing the depth of the mirror region. Based on this depth, each pixel within the mirror region can be projected onto a three-dimensional space to form a mirror region point cloud. The first spatial information characterizing the spatial features within the mirror region can then be extracted from the mirror region point cloud, facilitating subsequent identification of the relative spatial positions of objects and devices within the mirror region.
[0013] In one possible implementation, when the first image includes a mirror region, obtaining the first spatial information from the mirror region includes: obtaining depth information corresponding to the first image based on perception data; and then constructing a three-dimensional Gaussian sphere based on the depth information. The three-dimensional Gaussian sphere may include information obtained by constructing a three-dimensional scene from the information in the first image, which may include the scene of the mirror region reconstructed from the perspective of the mirror. The aforementioned first spatial information may include the three-dimensional Gaussian sphere.
[0014] In this embodiment of the application, depth information corresponding to all regions in the first image can be extracted based on perception data, and a three-dimensional Gaussian sphere can be constructed based on the depth information. The three-dimensional Gaussian sphere can include the scene of the mirror region reconstructed from the perspective of the mirror, so the three-dimensional scene of the mirror region can be directly constructed.
[0015] In one possible implementation, the aforementioned sensing data further includes first point cloud data, which comprises data acquired by radar. The aforementioned acquisition of depth information corresponding to the first image based on the sensing data includes: extracting a first feature from the first image; extracting a second feature from the first point cloud data; fusing the first and second features to obtain a third feature; and encoding and decoding the third feature using a codec to obtain depth information. In this embodiment, when sensing data is also acquired using radar, features representing depth can be extracted from the point cloud data acquired by radar, and a very accurate depth representing each pixel in the first image can be identified using a codec.
[0016] In one possible implementation, the aforementioned construction of a three-dimensional Gaussian sphere based on depth information includes: determining Gaussian parameters based on the depth information; and constructing a three-dimensional Gaussian sphere based on the Gaussian parameters. In this embodiment, when constructing the three-dimensional Gaussian sphere, the Gaussian parameters can be determined based on depth information to achieve scene reconstruction in the mirror region based on depth.
[0017] In one possible implementation, the aforementioned determination of Gaussian parameters based on depth information includes: projecting a first image into a three-dimensional space based on depth information to obtain second point cloud data, the second point cloud data being used to determine the center position of a three-dimensional Gaussian sphere; extracting color features from the first image and determining color information in the three-dimensional Gaussian sphere based on the color features; estimating ambient light information based on the first image and determining the illumination distribution of the three-dimensional Gaussian sphere based on the ambient light information, the illumination distribution being used to determine the spherical harmonic coefficients; acquiring information about elements in the first image and mapping the element information to the space corresponding to the three-dimensional Gaussian sphere to obtain element information in the three-dimensional Gaussian sphere; determining specularity based on the first image, specularity being used to represent the reflection intensity of a region in the three-dimensional Gaussian sphere, and specularity being used to indicate whether a mirror exists in the first image; the aforementioned Gaussian parameters include at least one of center position, color information, spherical harmonic coefficients, element information, or reflection intensity. Therefore, in this embodiment, a three-dimensional Gaussian sphere can be constructed from multiple dimensions, and each pixel in the three-dimensional Gaussian sphere can represent the attributes of each object in the space corresponding to the actual scene, thus enabling a more accurate and comprehensive construction of the three-dimensional scene.
[0018] In one possible implementation, the aforementioned process of projecting the first spatial information into the first space based on the mirror coordinate system corresponding to the mirror region to obtain the second spatial information includes: obtaining the information of the mirror region in a three-dimensional Gaussian sphere based on the mirror degree; performing clustering based on the information of the mirror region in the three-dimensional Gaussian sphere to obtain a mirror Gaussian cluster; and projecting the information of the mirror region in the three-dimensional Gaussian sphere into the first space based on the relative relationship between the mirror coordinate system corresponding to the mirror region and the first space to obtain the second spatial information.
[0019] In this embodiment, the specularity in a three-dimensional Gaussian sphere can be used to identify mirror regions, thereby enabling the construction of a scene in the mirror region using the three-dimensional Gaussian sphere, and obtaining three-dimensional spatial information that can represent the scene in the mirror.
[0020] In one possible implementation, the aforementioned method further includes: acquiring historical data, which includes data collected over historical time periods, and the historical data overlaps with the scene contained in the perceived data; subsequently determining the motion speed of each element in the first image based on the historical data, wherein the Gaussian parameter also includes the motion speed of each element. In this embodiment, historical data can also be used to determine the motion speed of each element in the scene, thereby enabling a more accurate 3D reconstruction of each object in the real scene and obtaining spatial information that can more accurately represent the real scene.
[0021] In one possible implementation, the aforementioned method further includes: determining the historical path of each element based on historical data; and obtaining the movement path of each element in a future time period based on the first spatial information and the historical path of each element.
[0022] In this embodiment, the historical paths of each element can be used to predict the future movement paths of each element, thereby predicting the changes of each element in the real scene and obtaining spatial information that is closer to the real scene.
[0023] Secondly, this application provides a mirror data processing apparatus, comprising:
[0024] The acquisition module is used to acquire sensing data, which includes a first image captured by an image sensor;
[0025] The first scene construction module is used to obtain first spatial information from the mirror region when the first image includes a mirror region. The first spatial information includes information about the elements included in the mirror region, and the mirror region includes a region where there is mirror reflection.
[0026] The second scene construction module is used to project the first spatial information into the first space according to the mirror coordinate system corresponding to the mirror area to obtain the second spatial information. The first space is the coordinate system corresponding to the image sensor, and the second spatial information includes the information of the elements in the mirror area in the first space.
[0027] The effects achieved by the second aspect or any optional implementation of the second aspect can be referred to the description of the first aspect or any optional implementation of the first aspect, and will not be repeated hereafter.
[0028] In one possible implementation, the aforementioned apparatus further includes a display module, configured to: render based on second spatial information to obtain a rendered image; and display the rendered image.
[0029] In one possible implementation, the aforementioned device is applied to a vehicle, and the device further includes a planning and control module for planning a driving path for the vehicle based on the second spatial information.
[0030] In one possible implementation, the aforementioned first scene construction module is further configured to: identify whether the first image includes a mirror area based on the perception data.
[0031] In one possible implementation, the aforementioned first scene construction module is specifically used for: extracting perceptual features from perceptual data; identifying whether a first image includes a mirror based on the perceptual features; and, if the first image includes a mirror, identifying the mirror region in the first image.
[0032] In one possible implementation, the aforementioned first scene construction module is specifically used for: if the first image includes a mirror region, estimating the depth of the mirror region to obtain a depth map; projecting the pixels of the mirror region into a three-dimensional space based on the depth map to obtain a mirror region point cloud; and obtaining first spatial information based on the mirror region point cloud.
[0033] In one possible implementation, when the first image includes a mirror region, the first scene construction module is specifically used to: obtain depth information corresponding to the first image based on perception data; construct a three-dimensional Gaussian sphere based on the depth information, wherein the three-dimensional Gaussian sphere includes a scene of the mirror region reconstructed from the perspective of the mirror, and the first spatial information includes the three-dimensional Gaussian sphere.
[0034] In one possible implementation, the aforementioned sensing data further includes first point cloud data, which includes data collected by radar;
[0035] The first scene construction module is specifically used for: extracting a first feature from the first image and extracting a second feature from the first point cloud data; fusing the first feature and the second feature to obtain a third feature; and using a codec to encode and decode the third feature to obtain depth information.
[0036] In one possible implementation, the aforementioned first scene construction module is specifically used for: determining Gaussian parameters based on depth information; and constructing a three-dimensional Gaussian sphere based on the Gaussian parameters.
[0037] In one possible implementation, the aforementioned first scene construction module is specifically used for: projecting a first image into a three-dimensional space based on depth information to obtain second point cloud data, the second point cloud data being used to determine the center position of a three-dimensional Gaussian sphere; extracting color features from the first image and determining color information in the three-dimensional Gaussian sphere based on the color features; estimating ambient light information based on the first image and determining the illumination distribution of the three-dimensional Gaussian sphere based on the ambient light information, the illumination distribution being used to determine the spherical harmonic coefficients; acquiring information about elements in the first image and mapping the element information to the space corresponding to the three-dimensional Gaussian sphere to obtain element information in the three-dimensional Gaussian sphere; determining specularity based on the first image, specularity being used to represent the reflection intensity of a region in the three-dimensional Gaussian sphere, and specularity being used to indicate whether a mirror exists in the first image; the Gaussian parameters include at least one of the center position, color information, spherical harmonic coefficients, element information, or reflection intensity.
[0038] In one possible implementation, the aforementioned second scene construction module is specifically used to: obtain information about the mirror region in a three-dimensional Gaussian sphere based on the mirrorness; perform clustering based on the information about the mirror region in the three-dimensional Gaussian sphere to obtain a mirror Gaussian cluster; and project the information about the mirror region in the three-dimensional Gaussian sphere onto the first space based on the relative relationship between the mirror coordinate system corresponding to the mirror region and the first space to obtain second space information.
[0039] In one possible implementation, the aforementioned second scene construction module is specifically used for: acquiring historical data, which includes data collected during historical time periods, and the historical data overlaps with the scene contained in the perceived data; determining the motion speed of each element in the first image based on the historical data, and the Gaussian parameter also includes the motion speed of each element.
[0040] In one possible implementation, the aforementioned apparatus further includes a path prediction module, configured to: determine the historical path of each element based on historical data; and obtain the movement path of each element in a future time period based on the first spatial information and the historical path of each element.
[0041] Thirdly, embodiments of this application provide a computing device including a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor calls program code in the memory to perform processing-related functions in the method shown in any of the first aspects above.
[0042] Fourthly, embodiments of this application provide a vehicle including a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor calls program code in the memory to perform processing-related functions in the method shown in any of the first aspects above.
[0043] Fifthly, embodiments of this application provide a digital processing chip or chip, the chip including a processing unit and a communication interface, the processing unit obtaining program instructions through the communication interface, the program instructions being executed by the processing unit, the processing unit being used to perform processing-related functions as described in the first aspect or any optional embodiment of the first aspect.
[0044] In a sixth aspect, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any optional implementation thereof.
[0045] In a seventh aspect, embodiments of this application provide a computer program product comprising a computer program / instructions, which, when executed by a processor, causes the processor to perform the method described in the first aspect or any optional implementation thereof. Attached Figure Description
[0046] Figure 1 is a schematic diagram of a vehicle structure provided in an embodiment of this application;
[0047] Figure 2 is a flowchart illustrating a mirror data processing method provided in an embodiment of this application;
[0048] Figure 3 is a flowchart illustrating another mirror data processing method provided in an embodiment of this application;
[0049] Figure 4 is a flowchart illustrating another mirror data processing method provided in an embodiment of this application;
[0050] Figure 5 is a schematic diagram of a display interface provided in an embodiment of this application;
[0051] Figure 6 is a schematic diagram of another display interface provided in an embodiment of this application;
[0052] Figure 7 is a flowchart illustrating another mirror data processing method provided in an embodiment of this application;
[0053] Figure 8 is a schematic diagram of a mirror data processing device provided in an embodiment of this application;
[0054] Figure 9 is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0055] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0056] To facilitate understanding, some concepts or categories involved in the embodiments of this application will be explained below.
[0057] (1) World Model
[0058] It refers to an internal representation or simulation structure used in fields such as artificial intelligence (AI), robotics, and computer vision to describe, understand, and predict the environment and its dynamic changes.
[0059] (2) transformer
[0060] A transformer architecture is a feature extraction network that includes both an encoder and a decoder. Of course, in some cases, a transformer architecture may not include an encoder but may include a decoder.
[0061] Encoder: Learns features, such as pixel features, in the global receptive field through self-attention.
[0062] Decoder: Learns the features of the desired module, such as the features of the output box, through self-attention and cross-attention.
[0063] For example, the structure of a Transformer layer in an existing scheme may include a multi-head attention network and a feedforward network module. Taking natural language processing as an example, the multi-head attention network obtains corresponding weight values by calculating the relevance between words, thus obtaining context-related word representations, which is the core part of the Transformer structure. The feedforward network further transforms the obtained representations to obtain the final output of the Transformer layer. In addition to these two important components, residual layers (ADD) and linear normalization (Norm) are also stacked on these two components to optimize the output of the Transformer layer.
[0064] (3) Multi-resolution encoder
[0065] In deep learning, especially in the fields of computer vision and image processing, a neural network architecture is used to extract multi-level, multi-scale features from input data.
[0066] (4) 3D Gaussian Splatting (3DGS)
[0067] By using Gaussian functions to represent points or volumes in three-dimensional space, efficient and accurate representation of three-dimensional scenes can be achieved.
[0068] 3DGS uses Gaussian functions to represent points or volumes in three-dimensional space. Each point can be described by a Gaussian function, which defines the point's spatial location and its degree of diffusion along various axes. In 3DGS, the probability of a point's existence is determined by the value of the Gaussian function, allowing for the modeling of blurred boundaries on surfaces rather than strict geometric boundaries. 3DGS can represent three-dimensional shapes simultaneously at multiple scales, thus capturing details at different levels, from microscopic to macroscopic.
[0069] (5) Occupancy (OCC) modeling
[0070] OCC (Optical Character Class) modeling refers to voxelizing a space, representing a 3D scene using voxel meshes, and modeling the scene by utilizing whether the voxel meshes are occupied and the semantics of the voxels (such as object category, color, or other attribute information). For example, if a voxel grid within a certain range is occupied, it means that an object exists within that range. Input data for OCC modeling can include point cloud data or images. By using the geometry and semantics of objects within the scene contained in OCC, a 3D scene can be constructed very accurately and vividly.
[0071] (6) Convex mirror distortion correction
[0072] In the fields of image processing and computer vision, convex mirrors, due to their curvature, cause specific geometric distortions in the images reflected through them. Distortion correction refers to restoring the geometric accuracy of an image, ensuring that the lines and shapes in the image are closer to the proportions and forms of the real world.
[0073] (7) Time of Flight (ToF) camera
[0074] A Time-of-Flight (ToF) camera is a type of camera that uses time-of-flight technology for depth perception. ToF technology calculates the distance between an object and the camera by measuring the time it takes for light to travel from the camera to the object and back, thus enabling 3D imaging and depth measurement.
[0075] (8) Point cloud downsampling
[0076] This refers to a method in 3D point cloud data processing that aims to preserve the original geometric features and structural information of the data as much as possible by reducing the number of points.
[0077] (9) Autoregressive model
[0078] By utilizing historical information in the sequence, future outputs are generated step by step, so that each step of generation depends on the results generated previously.
[0079] The method provided in this application can be applied to various vision task scenarios, such as intelligent driving scenarios for robots, flight scenarios for drones, or driving scenarios for intelligent vehicles, and other application scenarios of intelligent devices.
[0080] In the application scenarios of the aforementioned intelligent devices, such as intelligent driving of vehicles or scanning and mapping of robots, mirrors are often used to reflect blind spots. For example, in mountainous areas with curves or narrow intersections, autonomous driving systems often cannot clearly perceive the road conditions after a turn, lack prediction of the remote environment, and cannot accurately predict potential traffic risks, leading to an increased risk of accidents. Although existing intelligent driving technologies provide certain obstacle detection and avoidance functions through high-precision maps, radar perception, and visual recognition, the perception capabilities relied upon by traditional solutions have not completely overcome the limitations of perspective and environment, making it difficult to achieve accurate dynamic prediction and long-term path planning in curves or narrow intersections. Therefore, to reduce blind spots, using mirrors to reflect areas that are not visible to the user is a very important driver assistance method, widely used in traffic environments.
[0081] Taking convex mirrors as an example, their unique design provides a wider field of view, helping drivers to anticipate potential obstacles or pedestrians in curves or areas with limited visibility, thus avoiding traffic accidents. However, in intelligent driving scenarios, how to effectively utilize the image information provided by convex mirrors to enhance the vehicle's environmental perception capabilities and achieve more accurate decision-making and path planning remains a problem to be solved.
[0082] In existing solutions for intelligent driving around curves and narrow intersections, current solutions fall into two categories: high-risk detection based on traffic areas and convex mirror-based target detection-assisted perception and alerting schemes. The high-risk detection-based scheme analyzes the road geometry and surrounding obstacles. When it detects a narrowing traffic area, it determines that the vehicle has entered a potentially high-risk zone, automatically issuing a warning and implementing a deceleration or slow-moving strategy. However, this scheme can only assess risk based on the current environment and lacks the ability to predict the scene after a turn, failing to identify potential obstacles ahead of the curve. The convex mirror-based target detection-assisted perception scheme analyzes image information obtained from a convex mirror to detect the presence of obstacles or potential hazards. If a vehicle or other obstacle appears in the field of vision after a turn, the system can issue a warning, reminding the driver to take over control or take deceleration measures. However, this scheme can only detect obstacles in the convex mirror from 2D images and cannot effectively reconstruct the spatial relationship between objects and the vehicle, thus hindering accurate and long-term planning and control. To achieve autonomous driving perception and control capabilities in blind spots such as curves or parking lots, it is necessary to understand convex mirrors and generate 3D spatial information based on the image information in the convex mirrors, thereby enabling longer-term planning and improving active safety.
[0083] Therefore, this application provides a mirror data processing method that can project the elements contained in the mirror onto the space corresponding to the sensor based on the perception data containing the mirror area, thereby reconstructing the spatial relationship between the elements in the blind spot and the device, and providing users with a more effective blind spot field of view.
[0084] The method provided in this application can be applied to various vision task scenarios, such as intelligent driving scenarios for robots, flight scenarios for drones, or driving scenarios for intelligent vehicles, and other application scenarios of intelligent devices. Accordingly, the method provided in this application can be deployed in electronic devices such as robots, intelligent vehicles, or drones.
[0085] The following description uses a vehicle as an example of the electronic device. The vehicle mentioned below can also be replaced with a robot, drone, or other electronic device, which will not be elaborated further.
[0086] Referring to Figure 1, which is a schematic diagram of a vehicle structure provided in an embodiment of this application, the vehicle 100 can be configured in an intelligent driving mode. For example, the vehicle 100 can control itself while in intelligent driving mode, determine whether there are obstacles in the surrounding environment, and control the vehicle 100 based on the obstacle information. When the vehicle 100 is in intelligent driving mode, it can also be set to operate without human interaction.
[0087] Figure 1 is a functional block diagram of a vehicle 100 provided in an embodiment of this application. The vehicle 100 can be configured to a full or partial intelligent driving mode. For example, the vehicle 100 can obtain environmental information about its surroundings through the perception system 120, and obtain an intelligent driving strategy based on the analysis of the surrounding environmental information, or present the analysis results to the user.
[0088] Vehicle 100 may include various subsystems, such as an infotainment system 110, a perception system 120, a decision control system 130, a drive system 140, and a computing platform 160. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and component of vehicle 100 may be interconnected via wired or wireless means.
[0089] In some embodiments, the infotainment system 110 may include a communication system 111, an entertainment system 112, and a navigation system 113.
[0090] Communication system 111 may include wireless communication system 111, which can communicate wirelessly with one or more devices directly or via a communication network. For example, wireless communication system 111 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. Wireless communication system 111 may communicate using WiFi and wireless local area network (WLAN). In some embodiments, wireless communication system 146 may communicate directly with devices using an infrared link, Bluetooth, or ZigBee. Wireless communication system 111 may include one or more dedicated short range communications (DSRC) devices, which may include public and / or private data communications between vehicles and / or roadside stations.
[0091] The entertainment system 112 may include a central control screen, a microphone, and speakers. Users can listen to the radio and play music within the vehicle using the entertainment system 112; or connect their mobile phones to the vehicle and project their screens onto the central control screen, which may be touch-sensitive, allowing users to operate the system by touching the screen. In some cases, the microphone can acquire the user's voice signal, and analysis of the voice signal can enable the user to control certain aspects of the vehicle 100, such as adjusting the interior temperature. In other cases, music can be played to the user through the speakers. In this embodiment, the rendered image mentioned below can be displayed on a screen installed in the vehicle, such as a central control screen, rearview mirror, HUD (head-up display), or AR HUD, or other displays that can display information.
[0092] The navigation system 113 may include map services to provide navigation for the vehicle 100, and the navigation system 113 may be used in conjunction with the vehicle's global positioning system 121 and inertial measurement unit 122. The map may be a two-dimensional map, a high-precision map, or a map constructed based on data collected during the vehicle's operation.
[0093] The perception system 120 may include several sensors for sensing information about the environment surrounding the vehicle 100. For example, the perception system 120 may include a global positioning system 121 (which may be a GPS system, a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 122, a lidar 123, a millimeter-wave radar 124, an ultrasonic radar 125, and a camera device 126. The perception system 120 may also include sensors from the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, a fuel gauge, an oil temperature gauge, etc.). Sensor data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the vehicle 100. The perception data collected by sensors in the vehicle mentioned below in this application may include information collected by the various units in the perception system 120.
[0094] The inertial measurement unit 122 is used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration. In some embodiments, the inertial measurement unit 122 may be a combination of an accelerometer and a gyroscope.
[0095] The lidar 123 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the lidar 123 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components.
[0096] The millimeter-wave radar 124 can use radio signals to sense objects in the surrounding environment of the vehicle 100. In some embodiments, in addition to sensing objects, the millimeter-wave radar 124 can also be used to sense the speed and / or direction of travel of objects.
[0097] The ultrasonic radar 125 can use ultrasonic signals to sense objects around the vehicle 100.
[0098] The camera device 126 can be used to capture image information of the surrounding environment of the vehicle 100. The camera device 126 may include a monocular camera, a binocular camera, a structured light camera, and a panoramic camera, etc. The image information acquired by the camera device 126 may include still image information or video stream information.
[0099] The decision control system 130 includes a computing system 131 that analyzes and makes decisions based on information acquired by the sensing system 120. The decision control system 130 also includes a vehicle controller 132 that controls the power system of the vehicle 100, and a steering system 133, a throttle 134, and a braking system 135 for controlling the vehicle 100.
[0100] The computing system 131 can process and analyze various information acquired by the perception system 120 to identify targets, objects, and / or features in the environment surrounding the vehicle 100. The targets may include pedestrians or animals, and the objects and / or features may include traffic signals, road boundaries, and obstacles. The computing system 131 may use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other techniques. In some embodiments, the computing system 131 may be used to map the environment, track objects, estimate object speeds, etc. The computing system 131 can analyze the acquired information and derive a control strategy for the vehicle.
[0101] The vehicle controller 132 can be used to coordinate the control of the vehicle's power battery and engine 141 to improve the power performance of the vehicle 100.
[0102] The steering system 133 can be used to adjust the forward direction of the vehicle 100. For example, in one embodiment, it can be a steering wheel system.
[0103] The throttle 134 is used to control the operating speed of the engine 141 and thus the speed of the vehicle 100.
[0104] Braking system 135 is used to control the deceleration of vehicle 100. Braking system 135 can use friction to slow down the rotational speed of wheel 144. In some embodiments, braking system 135 can convert the kinetic energy of wheel 144 into electric current. Braking system 135 may also take other forms to slow down the rotational speed of wheel 144 to control the speed of vehicle 100.
[0105] The drive system 140 includes components that provide powered motion to the vehicle 100. In one embodiment, the drive system 140 may include an engine 141, an energy source 142, a transmission system 143, and wheels 144. The engine 141 may be an internal combustion engine, an electric motor, an air-compressed engine, or other types of engine combinations, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. The engine 141 converts the energy source 142 into mechanical energy.
[0106] Examples of energy sources 142 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy source 142 may also provide energy to other systems of vehicle 100.
[0107] The drivetrain 143 transmits mechanical power from the engine 141 to the wheels 144. The drivetrain 143 may include a gearbox, a differential, and a drive shaft. In one embodiment, the drivetrain 143 may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels 144.
[0108] Some or all of the functions of vehicle 100 are controlled by computing platform 160. Computing platform 160 may include at least one processor 151, which can execute instructions 153 stored in a non-transitory computer-readable medium such as memory 152. In some embodiments, computing platform 160 may also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner.
[0109] Processor 151 can be any conventional processor, such as a commercially available CPU. Alternatively, processor 151 may also include a graphics processing unit (GPU), a field-programmable gate array (FPGA), a system-on-chip (SoC), an application-specific integrated circuit (ASIC), or a combination thereof. Processor 151 can be located on a device remote from the vehicle and can communicate wirelessly with the vehicle.
[0110] In some embodiments, memory 152 may contain instructions 153 (e.g., program logic) that can be executed by processor 151 to perform various functions of vehicle 100. Memory 152 may also contain additional instructions, including instructions for sending data to, receiving data from, interacting with, and / or controlling one or more of the infotainment system 110, perception system 120, decision control system 130, and drive system 140.
[0111] In addition to instruction 153, memory 152 may also store data such as road maps, route information, vehicle position, direction, speed, and other similar vehicle data, as well as other information. This information can be used by vehicle 100 and computing platform 160 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes.
[0112] The computing platform 160 can control the functions of the vehicle 100 based on inputs received from various subsystems, such as the drive system 140, the perception system 120, and the decision control system 130. For example, the computing platform 160 can utilize inputs from the decision control system 130 to control the steering system 133 to avoid obstacles detected by the perception system 120. In some embodiments, the computing platform 160 is operable to provide control over many aspects of the vehicle 100 and its subsystems.
[0113] Alternatively, one or more of these components may be installed separately from or associated with vehicle 100. For example, memory 152 may exist partially or completely separately from vehicle 100. The components may be communicatively coupled together in a wired and / or wireless manner.
[0114] Optionally, the above components are just an example. In actual applications, the components in the above modules may be added or deleted according to actual needs. Figure 1 should not be construed as a limitation on the embodiments of this application.
[0115] The aforementioned vehicle 100 can be any vehicle or vehicle-mounted terminal capable of intelligent driving, such as a car, truck, motorcycle, bus, ship, airplane, helicopter, recreational vehicle, amusement park vehicle, construction equipment, tram, golf cart, or train. This application embodiment does not impose any special limitations on this type of vehicle.
[0116] The method flow provided in the embodiments of this application will be described below in conjunction with the aforementioned system architecture and vehicle architecture.
[0117] Referring to Figure 2, a flowchart of a mirror data processing method provided in an embodiment of this application is shown below.
[0118] 201. Acquire sensing data, which includes the first image captured by the image sensor.
[0119] The perception data includes data collected by at least one sensor, and the perception data set may include multiple frames of perception data. For example, if the method provided in this application embodiment is applied to a vehicle, the perception data set may include data collected by at least one sensor in the vehicle regarding the current scene of the vehicle.
[0120] The at least one sensor may specifically include an image sensor or radar, for example, it may include the sensors included in the perception system 120 in FIG1, see the description of the perception system 120.
[0121] Accordingly, the data in the sensing dataset includes point cloud data acquired by lidar, image data acquired by image sensors, or data acquired by ultrasonic radar, etc. When multiple types of data exist simultaneously, the calibration parameters of each sensor can be used to map each frame of sensing data into an aligned matrix, so as to facilitate subsequent processing of the aligned data.
[0122] In the embodiments of this application, it can be applied to the perception of scenes containing mirrors. For visual processing tasks involving mirrors, radar may not be able to capture the content reflected by the mirrors, but an image sensor can be used to capture images containing mirrors.
[0123] For ease of distinction, the image captured by the image sensor in the current scene will be referred to as the first image.
[0124] 202. If the first image includes a mirror region, obtain first spatial information from the mirror region.
[0125] Specifically, a pre-trained neural network can be used to extract the spatial features of the mirror region to obtain first spatial information, which includes information about the elements included in the mirror region, and the mirror region includes areas where mirror reflection exists.
[0126] The mirror can be a plane mirror, a convex mirror, or a concave mirror, etc., which can reflect all or part of the light. The specific form can be determined according to the actual application scenario. This application does not limit the specific form of the mirror.
[0127] In one possible implementation, whether a mirrored region is included in the first image can be identified based on perceptual data. If a mirrored region is included in the first image, targeted spatial information extraction of the mirrored region can be performed. Therefore, in this embodiment, a recognition process is set up for mirrored scenes, thereby enabling the extraction of more useful information in mirrored scenes.
[0128] In one possible implementation, a pre-trained neural network can be used to extract perceptual features from the perceptual data; based on the perceptual features, it can be identified whether the first image includes a mirror surface; if the first image includes a mirror surface, the mirror region in the first image can be identified, specifically through a pre-trained segmentation network or edge detection network, etc. In this embodiment, the presence of mirror information in the perceptual data can be identified by extracting features from the perceptual data.
[0129] In one possible implementation, if the first image includes a mirror region, depth estimation can be performed on the mirror region to obtain a depth map, which can represent the depth value corresponding to each pixel in the mirror region. Based on the depth map, the pixels of the mirror region in the first image are projected into a three-dimensional space to obtain a mirror region point cloud. This is equivalent to projecting the mirror region in the first image into the same space based on the depth values corresponding to each pixel in the mirror region, resulting in a pixel-level mirror region point cloud. Subsequently, first spatial information is obtained from the mirror region point cloud. For example, point cloud downsampling can be performed on the mirror region point cloud to obtain a spatial representation, or a pre-trained model can be used to obtain a representation that can represent the space corresponding to the mirror, specifically information about each element in the mirror space, such as the position, category, or color of each element in the mirror space.
[0130] For example, in one specific implementation, a pre-trained semantic segmentation network can be used to perform semantic segmentation on the perceived data. Based on the semantic segmentation results, it can be determined whether there is semantically related information about the mirror. If there is no mirror, the scene reconstruction of the vehicle space can be performed directly. If there is a mirror, a depth map can be extracted. Based on the depth map, the pixels of the mirror area in the image and the corresponding semantics are projected into the same space to obtain a pixel-level mirror area point cloud. Then, the mirror area point cloud is downsampled to obtain the spatial representation of the scene, which is the first spatial information.
[0131] In one possible implementation, depth information corresponding to the first image can be obtained based on the perception data, and a three-dimensional Gaussian sphere can be constructed based on the depth information. The three-dimensional Gaussian sphere includes the scene of the mirror region reconstructed from the viewpoint of the mirror, and the aforementioned first spatial information can include the three-dimensional Gaussian sphere. In this embodiment of the application, the scene reconstruction of the space corresponding to the mirror can be performed by constructing a three-dimensional Gaussian sphere to reconstruct the specific scene information in the space corresponding to the mirror.
[0132] In one possible implementation, the aforementioned sensing data may further include point cloud data acquired by radar, referred to as first point cloud data for ease of distinction. Specifically, a first feature can be extracted from the first image, and a second feature can be extracted from the first point cloud data; then, the first feature and the second feature are fused to obtain a third feature; the third feature is encoded and decoded using a codec to obtain depth information. In this embodiment, radar point cloud data can be used to assist in obtaining the depth corresponding to the image, thereby obtaining more accurate depth information through the radar's detection capabilities.
[0133] In one possible implementation, Gaussian parameters can be determined based on depth information; and a three-dimensional Gaussian sphere can be constructed based on the Gaussian parameters. This three-dimensional Gaussian sphere can serve as the first spatial information and can be used to represent the specific information of each element in the image.
[0134] In one possible implementation, the Gaussian parameters may include, but are not limited to, one or more of the following: Gaussian sphere position, color, spherical harmonic coefficient, transparency, category, velocity, or specularity. The Gaussian sphere position is used to represent the center point of the three-dimensional Gaussian sphere, the color is used to represent the color of each point in the three-dimensional Gaussian sphere, the spherical harmonic coefficient is used to represent the lighting or reflection characteristics of the three-dimensional Gaussian sphere under different viewing angles, the category represents the category of each element contained in the three-dimensional Gaussian sphere, the velocity represents the movement velocity of the elements that may be included in the three-dimensional Gaussian sphere, and the specularity represents the reflectivity of the elements that may be included in the three-dimensional Gaussian sphere, that is, whether the corresponding element is mirror-like.
[0135] Accordingly, the process of determining Gaussian parameters based on depth information may include one or more of the following:
[0136] The first image is projected into three-dimensional space based on the depth information to obtain the second point cloud data, which is used to determine the center position of the three-dimensional Gaussian sphere; color features are extracted from the first image, and color information in the three-dimensional Gaussian sphere is determined based on the color features.
[0137] Ambient light information is estimated based on the first image pair, and the illumination distribution of the three-dimensional Gaussian sphere is determined based on the ambient light information. This illumination distribution is used to determine the spherical harmonic coefficients.
[0138] Information about the elements in the first image is obtained and mapped to the space corresponding to the three-dimensional Gaussian sphere to obtain the element information in the three-dimensional Gaussian sphere, which includes the category of the element.
[0139] The specularity is determined based on the first image. The specularity is used to represent the reflection intensity of a region in a three-dimensional Gaussian sphere. The specularity is used to indicate whether there is a mirror in the first image.
[0140] Therefore, in the embodiments of this application, a three-dimensional Gaussian sphere can be constructed to reconstruct the scene in the mirror, thereby making full use of the information reflected in the mirror.
[0141] In one possible implementation, historical data can also be acquired, including data acquired during historical time periods. The historical data overlaps with the scene contained in the perceived data. For example, the historical data may include scene information reconstructed using data from the previous frame, including information on elements that are the same as those in the current scene. The motion speed of each element in the first image can be determined based on the historical data, and the aforementioned Gaussian parameters also include the motion speed of each element.
[0142] 203. Based on the mirror coordinate system corresponding to the mirror area, project the first spatial information into the first space to obtain the second spatial information.
[0143] The first space is the coordinate system corresponding to the image sensor, and the second space information includes the information of the elements in the mirror region projected into the first space.
[0144] Typically, the coordinate system corresponding to the mirror space differs from the space of the real scene. Therefore, the coordinate system corresponding to the image sensor can be represented as the coordinate system of the space corresponding to the scene. Based on the geometric relationship between the mirror coordinate system and the first space, the information of the first space is projected into the first space to obtain the second space information. This second space information can include the information of each element in the space corresponding to the actual scene after the information of each element in the mirror space is projected into the space corresponding to the actual scene, such as the position, color, or category of each element.
[0145] Therefore, in this embodiment, the information contained in the mirror in the image acquired by the image sensor can be used to reconstruct the scene and project it into the space corresponding to the visible area, so as to project the elements within the visual blind spot into the space, enabling the vehicle to perceive the elements within the visual blind spot and improve the vehicle's environmental perception capability.
[0146] In one possible implementation, for cases where the first spatial information includes a three-dimensional Gaussian sphere, the information of the mirror region within the three-dimensional Gaussian sphere can be obtained based on the specularity. Clustering is then performed on the information of the mirror region within the three-dimensional Gaussian sphere to obtain a mirror Gaussian cluster. Based on the relative relationship between the mirror coordinate system corresponding to the mirror region and the first space, the information of the mirror region within the three-dimensional Gaussian sphere is projected onto the first space to obtain the second spatial information. In this embodiment, Gaussian clusters for mirrors can be obtained by clustering the information within the Gaussian sphere, and based on the geometric relationship between the mirror coordinate system and the first space, these Gaussian clusters are projected onto the space within the visually visible area, thereby increasing the vehicle's perception range.
[0147] In another possible implementation, if a semantic segmentation network is used to identify whether a mirror exists in the scene, and when reconstructing the scene from the point cloud of the mirror area when a mirror exists, the spatial representation of the scene can be fused with the representation within the visually visible area based on the geometric relationship between the mirror coordinate system and the first space, thereby obtaining a perceptual representation that includes the visual blind spot. Therefore, in this embodiment, by utilizing the semantics of each element to reconstruct the scene, a spatial representation including both the visually visible area and the visual blind spot area can be obtained, which can increase the device's perception range of the environment.
[0148] In one possible implementation, after obtaining the second spatial information, rendering can be performed based on the second spatial information to obtain a rendered image; and the rendered image can be displayed on the display interface, so that the user can view the information within the blind spot of the field of view in a timely manner through the display interface, such as displaying the rendered image on the central control screen or HUD.
[0149] In one possible implementation, the method provided in this application embodiment can be applied to the intelligent driving function of a vehicle. After obtaining the second spatial information, the vehicle can also be planned according to the second spatial information. That is, the vehicle is controlled based on the reconstructed spatial information containing blind spots, thereby improving the vehicle's perception of blind spots and improving the vehicle's driving safety.
[0150] In one possible implementation, after obtaining the second spatial information, the historical paths of each element can be determined based on historical data; and based on the first spatial information and the historical paths of each element, the movement paths of each element in the future time period can be obtained. This is equivalent to combining the historical movement paths of each element with their current positions in the scene to predict their future movement paths. For example, for intelligent driving functions of vehicles, the possible movement paths of each element can be predicted in real time, thereby determining the vehicle's driving decisions promptly and improving driving safety.
[0151] The foregoing has described the method flow provided by the embodiments of this application. The following is a more detailed description of the method flow provided by the embodiments of this application in conjunction with specific application scenarios.
[0152] For example, taking the intelligent driving function of a vehicle as an example, after the intelligent driving function of the vehicle is activated, the sensors installed in the vehicle can be used to collect perception data. Regarding mirrors that may exist on the road, these can be plane mirrors or convex mirrors. Convex mirrors are typically installed to expand the field of vision. For example, in complex scenarios such as narrow roads or curves, using a convex mirror to reflect objects in the blind spot can improve the driver's visual range.
[0153] Referring to Figure 3, this is a schematic diagram of the architecture of the method provided in the embodiment of this application.
[0154] The method provided in this application embodiment can include multiple modules, such as a perception module 301, a mirror recognition module 302, a mirror scene generation module 303, a scene fusion module 304, a GUI module 305, and a planning and control module 306.
[0155] The perception module 301 can be used to collect environmental data, including but not limited to data collected by sensors such as cameras, LiDAR, millimeter-wave radar, and Time-of-Flight (ToF) cameras. Optionally, it can also perform time synchronization and spatial alignment between sensor data such as LiDAR and millimeter-wave radar and image data to generate multimodal perception data; it can also preprocess data from various sensors, such as desensitization or noise reduction, and fuse the preprocessed data to obtain more usable multimodal perception data. Especially for sensors set in different locations in the vehicle, the fused perception data can represent the environment further ahead of the vehicle, such as a forward-looking camera capturing convex mirror reflections in scenes such as curves, narrow intersections, or parking lots, and radar being used to detect objects that may exist in the environment.
[0156] The mirror recognition module 302 can be used to identify the presence of mirrors in the environment based on multimodal perception data. If a mirror is found, subsequent scene generation modules 303 and scene fusion modules 304 are activated to construct a scene representing the visual blind spot for the mirrored scene. Specifically, it can detect convex mirrors by acquiring images from a forward-facing camera. Specifically, the mirror recognition module 302 first preprocesses the image, such as denoising and image enhancement, and then uses deep learning algorithms (such as convolutional neural networks CNN) for tasks such as target detection and object recognition. If a convex mirror is detected, the next step of convex mirror scene generation is performed; if no convex mirror is detected, the normal vehicle perception process can be executed.
[0157] The mirror scene generation module 303 can analyze images acquired by image sensors, such as the aforementioned first image, extract and understand depth information in the image reflected by the convex mirror, and convert the 2D information of the front view into spatial representation information such as bird's-eye view, occupancy grid, or 3DGS. This spatial representation information includes information within the visual blind spot range reflected in the convex mirror. This enables the vehicle's intelligent driving system not only to identify the type and location of obstacles, but also to determine the distance and relative motion state between objects and the vehicle through depth reasoning. Specifically, the steps performed by the mirror scene generation module 303 may include: first, extracting mirror pixels based on the detection module; second, establishing a distortion perception layer to estimate radial distortion parameters and perform distortion correction; and finally, reconstructing the mirror scene from the mirror perspective based on 3D representations such as bird's-eye view, occupancy grid, or 3DGS.
[0158] The scene fusion module 304 can be used to transform the scene in the mirror to the space corresponding to the vehicle coordinate system according to the transformation relationship from the mirror coordinate system to the vehicle coordinate system. This allows the information perceived from the mirror image to be in the same space as the information in the non-mirror area, making it easier for the vehicle's intelligent driving function to perceive information within the visual range and the visual blind spot.
[0159] The GUI module 305 can be used to render the complete scene information output by the scene fusion module 304 and display it through the display interface, thereby showing the user information within the visual blind spot range.
[0160] The planning and control module 306 can determine appropriate driving decisions for the vehicle based on input information. If there are obstacles in front of the vehicle, it can determine driving decisions such as braking, lane changing, or deceleration to ensure safe driving. In scenarios with mirrored surfaces, it performs path planning and dynamic control based on fused 3D spatial scene information to ensure the intelligent driving vehicle can drive safely and smoothly in complex environments such as curves or narrow intersections. The path planning algorithm deployed in the planning and control module 306 dynamically generates a safe and smooth driving trajectory based on fused scene information (including the vehicle's current position, target position, and obstacle positions). The control algorithm deployed in the planning and control module 306 generates precise vehicle control signals according to the trajectory instructions output by the path planning module, enabling the vehicle to drive safely along the planned path. The core of the control algorithm lies in tracking control and dynamic response, ensuring that the vehicle closely follows the planned trajectory and reacts promptly in complex dynamic environments.
[0161] Furthermore, embodiments of this application provide intelligent driving scenarios with blind spots, such as curves, parking lots, and narrow intersections. Some possible application scenarios are described below.
[0162] Scene 1: Curve Scene
[0163] In mountainous or urban bends, a vehicle's field of vision is limited, especially the environment behind the turn, which often cannot be fully understood using conventional sensors (such as forward-facing cameras and radar). In these situations, the vehicle uses a forward-facing camera to capture images reflected by a convex mirror. The intelligent driving system can extract obstacle information from these reflected images, such as whether there are other vehicles, pedestrians, or obstacles in the field of vision after the turn. By generating three-dimensional spatial data, the system can accurately determine the distance and position of these objects, thereby making corresponding deceleration, avoidance, or steering adjustments.
[0164] Scenario 2: Parking Lot
[0165] In complex parking lots or narrow lanes, traditional sensors may not be able to fully detect all obstacles or moving objects, especially between parking spaces or in areas with limited visibility. By installing convex mirrors at specific locations in the parking lot, vehicles can use reflected images to detect the presence of obstacles in advance, such as pedestrians, other vehicles, or parked objects. The system analyzes this image information, combined with data from other sensors, to generate real-time 3D spatial data and determine the dynamic situation of objects, thereby precisely controlling the vehicle's parking position or performing obstacle avoidance maneuvers.
[0166] Scene 3: Narrow intersection
[0167] At narrow intersections, autonomous driving systems often face challenges such as insufficient visibility and complex environments. Convex mirrors can provide crucial reflected images in these scenarios, helping the system anticipate potential obstacles after a turn or at the intersection. By combining depth information from sensors such as LiDAR, the system can construct a 3D spatial model and assess the environment after a turn, thereby making decisions in advance regarding whether to slow down or adjust the path.
[0168] The method provided in the embodiments of this application will be further described below in conjunction with the aforementioned application scenarios.
[0169] The method provided in this application embodiment can be divided into two methods for reconstructing the scene in the mirror by extracting the spatial representation of the mirror space: OCC modeling or 3DGS modeling. The different modeling methods are described below.
[0170] Implementation Method 1: Scene Reconstruction via 3DGS
[0171] Referring to Figure 4, a flowchart of another mirror data processing method provided in this application embodiment is shown below.
[0172] 401. Acquire multimodal sensing data.
[0173] First, the multimodal perception data can include, but is not limited to, radar point cloud data or images. Specifically, images can include those collected by image sensors located at different positions within the vehicle, such as those deployed at the front, body, or rear of the vehicle. Radar point cloud data can include point cloud data collected by radars such as LiDAR or millimeter-wave radar. Furthermore, radar point cloud data and images can cover the same scene; for example, both the radar point cloud data and the images can contain information about a convex mirror positioned on the road.
[0174] 402. Depth estimation.
[0175] Specifically, depth estimation can be performed using multimodal sensing data, such as estimating the depth values corresponding to a set of pixels in an image, where a set of pixels may include one or more pixels.
[0176] Specifically, dense depth maps can be calculated by fusing image and point cloud information. First, 3x3 convolutions are performed on the input image and input sparse point cloud to extract features, and the resulting image features and point cloud features are concatenated together. Next, a multi-resolution encoder is used to extract multi-resolution features from the concatenated image and point cloud features. Then, transposed convolutions are used to upsample the spatial resolution of the multi-resolution features to form a multi-resolution decoder, and the features output from each encoding layer are passed to the corresponding decoding layer through skip connections. Finally, the output of the last layer of the decoding layer is input into a 1x1 convolution to calculate a dense depth map with the same resolution as the input image.
[0177] In this embodiment, high-precision and high-resolution depth estimation can be achieved, improving the understanding of space and enhancing the final rendering accuracy.
[0178] 403. Calculate the Gaussian parameters.
[0179] The Gaussian parameters can include, but are not limited to, one or more of the following: Gaussian sphere position, color, spherical harmonics, transparency, category, velocity, or specularity. Gaussian sphere position represents the coordinates of each pixel in the 3D Gaussian sphere; color represents the color value of each pixel in the 3D Gaussian sphere; spherical harmonics represent the lighting and reflection characteristics of each pixel in the 3D Gaussian sphere under different viewing angles; transparency represents the transparency of each pixel in the 3D Gaussian sphere; category represents the category of the elements composed of each pixel; velocity represents the movement speed of each pixel in the 3D Gaussian sphere; and specularity can represent the reflection intensity of each pixel.
[0180] By using Gaussian parameters, a three-dimensional Gaussian sphere can be constructed to simulate real-world scenes. Each pixel has corresponding parameters such as Gaussian sphere position, color, spherical harmonic coefficient, transparency, category, velocity, or specularity. Multiple pixels can form a three-dimensional Gaussian sphere to simulate the environment.
[0181] For example, the dense depth map generated by the depth estimation module can be used to convert pixels into point cloud data in 3D space through back projection, thereby determining the center position of the Gaussian sphere.
[0182] Extract color information at the corresponding position from the input image. By projecting the position of the Gaussian sphere onto the image plane, sample the RGB values of the corresponding pixels as the color of the Gaussian sphere.
[0183] Spherical harmonics are used to describe the illumination and reflection characteristics of a Gaussian sphere under different viewpoints. By estimating ambient light, the illumination distribution in the scene is analyzed, and the spherical harmonics are extracted as the spherical harmonics values of the Gaussian sphere.
[0184] The transparency parameter describes the transparency of the Gaussian sphere and is initialized to 1 during the initialization phase.
[0185] The category parameter is used to identify the object category represented by the Gaussian sphere, such as vehicle, pedestrian, bicycle, etc. Through the semantic segmentation network, the pixel-level category information in the image is mapped to the Gaussian sphere to obtain the initial category value of the Gaussian sphere.
[0186] The velocity parameter describes the motion state of the Gaussian sphere in three-dimensional space and is represented by a velocity vector. The initial velocity of the Gaussian sphere is calculated by the displacement between consecutive frames.
[0187] The specularity parameter describes the reflection characteristics of the Gaussian sphere surface, especially the intensity of specular reflection. Edge detection and highlight area analysis are used to calculate the initial specular value of the Gaussian sphere.
[0188] Therefore, in this embodiment, a three-dimensional Gaussian sphere can be constructed using perceived data to simulate the specific digital information of a real scene, enabling the vehicle to perceive the real environment. For scenes containing convex mirrors, the specularity of the convex mirrors in the environment can be used to represent them, so that the constructed three-dimensional Gaussian sphere not only includes scene information within the visually visible range but also information about the scene reflected in the convex mirrors.
[0189] 404. Convex mirror processing for scene reconstruction.
[0190] After obtaining the three-dimensional Gaussian sphere, the mirror points are extracted based on the mirror properties of the Gaussian sphere and a 3D space is generated from them.
[0191] Specifically, firstly, mirror pixels are extracted based on specularity, and a pixel is identified as a mirror pixel when its specularity is greater than a fixed value. Secondly, the mirror pixels are clustered to obtain a mirror Gaussian cluster, which is then input into the distortion perception layer to estimate radial distortion parameters. The estimated radial distortion parameters are then applied to the original image for distortion correction. Finally, the geometric relationship between the mirror coordinate system and the vehicle coordinate system is determined, including position, rotation, and translation parameters. A coordinate transformation matrix is then applied to convert the position, orientation, and other parameters of the reconstructed 3D Gaussian sphere in the mirror to the vehicle coordinate system, thereby projecting the scene in the mirror onto the space corresponding to the real scene.
[0192] 405. Scene generation: Outputs scene spatial information and predicted object motion paths.
[0193] Specifically, spatial information of the current scene and path prediction of each element in the scene can be generated based on the fused Gaussian representation of the entire scene and the historical paths of each element.
[0194] First, based on the aforementioned Gaussian parameters and historical paths, the parameters are normalized and input into the pre-trained encoder to obtain latent space features, which are then quantized based on the codebook. Second, the quantized latent space features are input into the autoregressive module for spatial and path prediction in the next frame. Finally, the latent space tokens are decoded into Gaussian parameters and path information for the next frame by the decoder, which are used to make driving decisions or control vehicle movement.
[0195] Therefore, in this application, a scene reconstruction scheme based on 3DGS that maps 2D images of convex mirrors to 3D space and the motion path prediction of each element is provided. This can achieve end-to-end prediction of future prediction paths and control signals, and improve the prediction and control accuracy in occluded space environments such as curves or intersections.
[0196] In addition, the scene space information and predicted object motion paths output in step 405 can be displayed to the user on the display interface after rendering, so that the user can know the surrounding environment in a timely manner through the content displayed on the display interface.
[0197] For example, as shown in Figures 5 and 6, for blind spot scenarios such as curves or narrow intersections (Figure 5 shows a curve scenario and Figure 6 shows a narrow intersection scenario), more comprehensive scene information can be displayed to the user on the display interface, including scene information of visual blind spots, so that the user can be informed of the surrounding environment in a timely manner, which can improve the driving safety of the vehicle.
[0198] In this embodiment, to address potential blind spot issues, information from a 2D convex mirror image is projected into 3D space, thereby representing the blind spot space in three dimensions. This improves environmental perception in scenarios with blind spots, such as curves or narrow intersections, enhancing the performance of vehicle intelligent driving functions. This significantly improves driving safety in blind spot scenarios. Furthermore, the introduction of a Gaussian sphere position based on joint depth estimation using images and point clouds in the 3DGS initialization process improves spatial rendering accuracy. Moreover, by predicting future scene spatial information and the movement paths of various elements based on historical perception data and the historical movement paths of each element, an end-to-end intelligent driving solution can be achieved.
[0199] Implementation Method 2: Scene Reconstruction via OCC
[0200] Referring to Figure 7, a flowchart of another mirror data processing method provided in this application embodiment is shown below.
[0201] 701. Acquire multimodal sensing data.
[0202] Step 701 is similar to step 401 mentioned above, and will not be described again here.
[0203] 702. Perform semantic segmentation on multimodal sensing data.
[0204] Pre-trained semantic segmentation networks can be used to perform semantic segmentation on input images or point cloud data, and identify the categories of each element included in the perceptual data.
[0205] 703. Determine whether a convex mirror exists based on the semantic segmentation result. If yes, proceed to step 704; otherwise, proceed to step 709.
[0206] Based on the semantic segmentation results, it is determined whether there is a convex mirror in the scene corresponding to the perception data. If so, subsequent data processing is performed on the convex mirror. If not, the scene information of the space where the vehicle is located can be directly reconstructed based on the perception data.
[0207] Specifically, the semantic segmentation result can include category labels corresponding to each element. The category labels contained in the semantic segmentation can be traversed to identify whether there are category labels related to mirrors, thereby identifying whether there are mirrors in the scene.
[0208] 704. Distortion removal.
[0209] Specifically, based on the semantic segmentation results, the information of the pixels corresponding to the convex mirror region can be extracted from the image, a distortion perception layer can be established to estimate the radial distortion parameter, and the region corresponding to the convex mirror in the image can be dedistorted according to the parameter, thereby obtaining a convex mirror region with less distortion.
[0210] 705. Depth estimation.
[0211] Depth estimation is performed on the distorted image to obtain a depth map of the scene in the mirror.
[0212] Specifically, depth estimation can be performed using multimodal sensing data, such as estimating the depth value corresponding to a set of pixels in an image. A set of pixels can include one or more pixels.
[0213] Specifically, dense depth maps can be calculated by fusing image and point cloud information. First, 3x3 convolutions are performed on the input image and input sparse point cloud to extract features, and the resulting image features and point cloud features are concatenated together. Next, a multi-resolution encoder is used to extract multi-resolution features from the concatenated image and point cloud features. Then, transposed convolutions are used to upsample the spatial resolution of the multi-resolution features to form a multi-resolution decoder, and the features output from each encoding layer are passed to the corresponding decoding layer through skip connections. Finally, the output of the last layer of the decoding layer is input into a 1x1 convolution to calculate a dense depth map with the same resolution as the input image.
[0214] 706. Point cloud generation.
[0215] After obtaining the depth map, based on the depth values of each pixel contained in the depth map, the pixels in the convex mirror region of the image are projected into the same three-dimensional space to obtain pixel-level point cloud data.
[0216] 707. Create a mirror space OCC.
[0217] After obtaining pixel-level point cloud data, the pixel-level point cloud data can be processed, such as using the OCC tool to downsample the point cloud and obtain the OCC spatial representation, which is a representation of the space of the scene contained in the mirror area.
[0218] 708. Mirror space OCC coordinate transformation.
[0219] For the mirror space OCC, its corresponding coordinate system is the mirror space corresponding to the mirror area in the image. Based on the positional relationship between the mirror area and the non-mirror area in the point cloud, the positional relationship between the mirror coordinate system and the vehicle coordinate system can be determined. Based on this positional relationship, the mirror space OCC can be projected into the vehicle coordinate system, that is, the space corresponding to the scene where the vehicle is located.
[0220] 709. Construct the vehicle's OCC.
[0221] In the absence of convex mirrors in the current scene, the OCC tool can be used directly to construct the vehicle's OCC using perception data, resulting in a full-scene OCC that does not include convex mirrors.
[0222] 710. Integration of Mirror Space OCC and Vehicle Space OCC.
[0223] After projecting the mirror space OCC onto the vehicle coordinate system, the projected mirror OCC can be merged with the vehicle space OCC to obtain the full-scene OCC. This allows each element in the mirror area to be in the same space as each element in the non-mirror area, thus better representing the relative positional relationship between each element in the real scene.
[0224] 711. Path planning.
[0225] After obtaining the full-scene OCC, target detection can be performed based on the full-scene OCC, and the vehicle can be planned to drive according to the detected target or controlled to drive according to the planned driving path.
[0226] In this embodiment, addressing the blind spot problem caused by convex mirrors, depth estimation is used to generate OCC-based spatial information from the mirror image. This allows for a three-dimensional understanding of the blind spot space, enabling accurate dynamic prediction and long-term path planning within blind spots at curves or narrow intersections. This significantly improves vehicle safety in scenarios with blind spots, such as curves or narrow intersections. Furthermore, the OCC-based blind spot convex mirror scene fusion requires less computation, offers high real-time performance, and enables more efficient scene perception.
[0227] The foregoing has described the method flow provided in the embodiments of this application. The following describes the structure of the apparatus for executing the foregoing method flow.
[0228] Referring to Figure 8, a schematic diagram of a mirror data processing device provided in an embodiment of this application includes:
[0229] Acquisition module 801 is used to acquire sensing data, the sensing data including a first image acquired by an image sensor;
[0230] The first scene construction module 802 is used to obtain first spatial information from the mirror region when the first image includes a mirror region. The first spatial information includes information about the elements included in the mirror region, and the mirror region includes a region where there is mirror reflection.
[0231] The second scene construction module 803 is used to project the first spatial information into the first space according to the mirror coordinate system corresponding to the mirror area to obtain the second spatial information. The first space is the coordinate system corresponding to the image sensor, and the second spatial information includes the information of the elements in the mirror area in the first space.
[0232] In one possible implementation, the aforementioned apparatus further includes a display module 804, configured to: render based on second spatial information to obtain a rendered image; and display the rendered image.
[0233] In one possible implementation, the aforementioned device is applied to a vehicle, and the device further includes a planning and control module 805 for planning a driving path for the vehicle based on the second spatial information.
[0234] In one possible implementation, the aforementioned first scene construction module 802 is further configured to: identify whether the first image includes a mirror area based on the perception data.
[0235] In one possible implementation, the aforementioned first scene construction module 802 is specifically used for: extracting perceptual features from perceptual data; identifying whether a first image includes a mirror surface based on the perceptual features; and, if the first image includes a mirror surface, identifying the mirror area in the first image.
[0236] In one possible implementation, the aforementioned first scene construction module 802 is specifically used to: when the first image includes a mirror region, perform depth estimation on the mirror region to obtain a depth map; project the pixels of the mirror region onto a three-dimensional space based on the depth map to obtain a mirror region point cloud; and obtain first spatial information based on the mirror region point cloud.
[0237] In one possible implementation, when the first image includes a mirror region, the first scene construction module 802 is specifically used to: obtain depth information corresponding to the first image based on perception data; construct a three-dimensional Gaussian sphere based on the depth information, wherein the three-dimensional Gaussian sphere includes a scene of the mirror region reconstructed from the perspective of the mirror, and the first spatial information includes the three-dimensional Gaussian sphere.
[0238] In one possible implementation, the aforementioned sensing data further includes first point cloud data, which includes data collected by radar;
[0239] The first scene construction module 802 is specifically used for: extracting a first feature from the first image and extracting a second feature from the first point cloud data; fusing the first feature and the second feature to obtain a third feature; and using a codec to encode and decode the third feature to obtain depth information.
[0240] In one possible implementation, the aforementioned first scene construction module 802 is specifically used for: determining Gaussian parameters based on depth information; and constructing a three-dimensional Gaussian sphere based on the Gaussian parameters.
[0241] In one possible implementation, the aforementioned first scene construction module 802 is specifically used for: projecting a first image into a three-dimensional space based on depth information to obtain second point cloud data, the second point cloud data being used to determine the center position of a three-dimensional Gaussian sphere; extracting color features from the first image and determining color information in the three-dimensional Gaussian sphere based on the color features; estimating ambient light information based on the first image and determining the illumination distribution of the three-dimensional Gaussian sphere based on the ambient light information, the illumination distribution being used to determine the spherical harmonic coefficients; acquiring information about elements in the first image and mapping the element information to the space corresponding to the three-dimensional Gaussian sphere to obtain element information in the three-dimensional Gaussian sphere; determining specularity based on the first image, specularity being used to represent the reflection intensity of a region in the three-dimensional Gaussian sphere, and specularity being used to indicate whether a mirror exists in the first image; the Gaussian parameters include at least one of the center position, color information, spherical harmonic coefficients, element information, or reflection intensity.
[0242] In one possible implementation, the aforementioned second scene construction module 803 is specifically used to: obtain information about the mirror region in a three-dimensional Gaussian sphere based on the mirrorness; perform clustering based on the information about the mirror region in the three-dimensional Gaussian sphere to obtain a mirror Gaussian cluster; and project the information about the mirror region in the three-dimensional Gaussian sphere into the first space based on the relative relationship between the mirror coordinate system corresponding to the mirror region and the first space to obtain second space information.
[0243] In one possible implementation, the aforementioned second scene construction module 803 is specifically used to: acquire historical data, which includes data collected during historical periods, and the historical data overlaps with the scene contained in the perception data; determine the motion speed of each element in the first image based on the historical data, and the Gaussian parameter also includes the motion speed of each element.
[0244] In one possible implementation, the aforementioned apparatus further includes a path prediction module 806, configured to: determine the historical path of each element based on historical data; and obtain the movement path of each element in a future time period based on the first spatial information and the historical path of each element.
[0245] Figure 9 shows a schematic diagram of the hardware structure of a computing device 90 provided in an embodiment of this application. This computing device 90 can be used to implement the steps of the methods shown in Figures 2 to 7, and may specifically include the aforementioned vehicle or devices deployed on a vehicle.
[0246] The computing device 90 shown in Figure 9 may include a processor 901, a memory 902, a communication interface 903, and a bus 904. The processor 901, the memory 902, and the communication interface 903 can be connected via the bus 904.
[0247] The processor 901 is the control center of the computing device 90. It can be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor, such as a GPU or NPU, which can be adapted to the actual application scenario.
[0248] As an example, processor 901 may include one or more CPUs, and may also include other processors, such as the CPU, NPU or GPU shown in Figure 9.
[0249] The memory 902 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0250] In one possible implementation, the memory 902 can exist independently of the processor 901. The memory 902 can be connected to the processor 901 via a bus 904 and is used to store data, instructions, or program code. When the processor 901 calls and executes the instructions or program code stored in the memory 902, it can implement the methods provided in the embodiments of this application, such as the methods shown in Figures 2 to 7.
[0251] In another possible implementation, the memory 902 can also be integrated with the processor 901.
[0252] The communication interface 903 is used for connecting the computing device 90 to other devices via a communication network, which may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The communication interface 903 may include a receiving unit for receiving data and a transmitting unit for sending data.
[0253] Bus 904 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 9, but this does not indicate that there is only one bus or one type of bus.
[0254] It should be noted that the structure shown in Figure 9 does not constitute a limitation on the computing device 90. In addition to the components shown in Figure 9, the computing device 90 may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0255] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0256] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0257] This application also provides a computer-readable storage medium storing a program for training a model or performing inference tasks, which, when run on a computer, causes the computer to perform all or part of the steps in the methods described in the embodiments shown in Figures 2 to 7 above.
[0258] This application also provides a digital processing chip. This digital processing chip integrates circuitry for implementing the aforementioned processor or processor functions, and one or more interfaces. When the digital processing chip integrates a memory, it can perform the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the method steps of any one or more of the foregoing embodiments based on the program code stored in the external memory.
[0259] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0260] The apparatus provided in this application embodiment can be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip to execute the methods described in the embodiments shown in Figures 2 to 7. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0261] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0262] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0263] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0264] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0265] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0266] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. The term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps in this application does not imply that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved. The division of modules in this application is a logical division. In actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed may be through some ports, and the indirect coupling or communication connection between modules may be electrical or other similar forms, which are not limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules can be selected to achieve the purpose of the solution in this application according to actual needs.
Claims
1. A mirror data processing method, characterized by, include: Acquire sensing data, the sensing data including a first image acquired by an image sensor; If the first image includes a mirror region, first spatial information is obtained from the mirror region. The first spatial information includes information about the elements included in the mirror region, and the mirror region includes a region where there is specular reflection. Based on the mirror coordinate system corresponding to the mirror area, the first spatial information is projected into the first space to obtain the second spatial information. The first space is the coordinate system corresponding to the image sensor, and the second spatial information includes the information of the elements in the mirror area in the first space.
2. The method of claim 1, wherein, The method further includes: Rendering is performed based on the second spatial information to obtain a rendered image; The rendered image is displayed.
3. The method according to claim 1 or 2, characterized in that, The method is applied to a vehicle, and the method further includes: The vehicle's driving path is planned based on the second spatial information.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: The first image is identified based on the perceived data to determine whether it includes the mirror area.
5. The method of claim 4, wherein, The step of identifying whether the first image includes the mirror region based on the perceived data includes: Extract perceptual features from the perceptual data; Based on the perceived features, identify whether the first image includes a mirror surface; If the mirror surface is included in the first image, the mirror surface region in the first image is identified.
6. The method according to claim 4 or 5, characterized in that, The step of obtaining the first spatial information from the mirror region includes: If the first image includes a mirrored region, depth estimation is performed on the mirrored region to obtain a depth map; The pixels of the mirror region are projected into three-dimensional space based on the depth map to obtain the point cloud of the mirror region; The first spatial information is obtained from the point cloud of the mirror area.
7. The method according to any one of claims 1-3, characterized in that, If the first image includes a mirrored region, obtaining first spatial information from the mirrored region includes: The depth information corresponding to the first image is obtained based on the perceived data; A three-dimensional Gaussian sphere is constructed based on the depth information. The three-dimensional Gaussian sphere includes a scene of the mirror region reconstructed from the perspective of the mirror. The first spatial information includes the three-dimensional Gaussian sphere.
8. The method of claim 7, wherein, The sensing data also includes first point cloud data, which includes data collected by radar. The step of obtaining the depth information corresponding to the first image based on the perceived data includes: Extract a first feature from the first image and extract a second feature from the first point cloud data; By combining the first feature and the second feature, a third feature is obtained; The third feature is encoded and decoded using a codec to obtain the depth information.
9. The method according to claim 7 or 8, characterized in that, The construction of a three-dimensional Gaussian sphere based on the depth information includes: Determine the Gaussian parameters based on the depth information; The three-dimensional Gaussian sphere is constructed based on the Gaussian parameters.
10. The method of claim 9, wherein, Determining the Gaussian parameters based on the depth information includes: The first image is projected into a three-dimensional space based on the depth information to obtain second point cloud data, which is used to determine the center position of the three-dimensional Gaussian sphere. Alternatively, color features can be extracted from the first image, and color information in the three-dimensional Gaussian sphere can be determined based on the color features; Alternatively, ambient light information can be estimated based on the first image pair, and the illumination distribution of the three-dimensional Gaussian sphere can be determined based on the ambient light information, wherein the illumination distribution is used to determine the spherical harmonic coefficients; Alternatively, information about the elements in the first image can be obtained, and the information about the elements can be mapped to the space corresponding to the three-dimensional Gaussian sphere to obtain the element information in the three-dimensional Gaussian sphere; Alternatively, the specularity can be determined based on the first image, whereby the specularity represents the reflection intensity of a region within the three-dimensional Gaussian sphere, and the specularity also indicates whether a mirror exists in the first image. The Gaussian parameters include at least one of the center position, the color information, the spherical harmonic coefficient, the element information, or the reflection intensity.
11. The method of claim 10, wherein, The step of projecting the first spatial information into the first space according to the mirror coordinate system corresponding to the mirror area to obtain the second spatial information includes: Information about the mirror region in the three-dimensional Gaussian sphere is obtained based on the specularity. Clustering is performed on the information in the three-dimensional Gaussian sphere of the mirror region to obtain a mirror Gaussian cluster; Based on the mirror coordinate system corresponding to the mirror region and its relative relationship with the first space, the information of the mirror region in the three-dimensional Gaussian sphere is projected onto the first space to obtain the second space information.
12. The method according to claim 10 or 11, characterized in that, The method further includes: Acquire historical data, which includes data collected during historical periods, and the historical data overlaps with the scenes contained in the perceived data; The motion speed of each element in the first image is determined based on the historical data, and the Gaussian parameter also includes the motion speed of each element.
13. The method of claim 12, wherein, The method further includes: The historical path of each element is determined based on the historical data; Based on the first spatial information and the historical paths of each element, the movement paths of each element in future time periods are obtained.
14. A mirror data processing apparatus, characterized by include: The acquisition module is used to acquire sensing data, which includes a first image captured by an image sensor; A first scene construction module is used to obtain first spatial information from the mirror region when the first image includes a mirror region. The first spatial information includes information about the elements included in the mirror region, and the mirror region includes a region where there is mirror reflection. The second scene construction module is used to project the first spatial information into the first space according to the mirror coordinate system corresponding to the mirror area to obtain the second spatial information. The first space is the coordinate system corresponding to the image sensor, and the second spatial information includes the information of the elements in the mirror area in the first space.
15. The apparatus of claim 14, wherein, The device further includes a display module, used for: Rendering is performed based on the second spatial information to obtain a rendered image; The rendered image is displayed.
16. The apparatus of claim 14 or 15, wherein, The device is applied to a vehicle, and the device further includes: The planning and control module is used to plan a driving path for the vehicle based on the second spatial information.
17. The apparatus of any one of claims 14-16, wherein, The first scene construction module is also used for: The first image is identified based on the perceived data to determine whether it includes the mirror area.
18. The apparatus of claim 17, wherein, The first scene construction module is specifically used for: Extract perceptual features from the perceptual data; Based on the perceived features, identify whether the first image includes a mirror surface; If the mirror surface is included in the first image, the mirror surface region in the first image is identified.
19. The apparatus of claim 17 or 18, wherein, The first scene construction module is specifically used for: If the first image includes a mirrored region, depth estimation is performed on the mirrored region to obtain a depth map; The pixels of the mirror region are projected into three-dimensional space based on the depth map to obtain the point cloud of the mirror region; The first spatial information is obtained from the point cloud of the mirror area.
20. The apparatus of any one of claims 14-16, wherein, In the case that the first image includes a mirrored region, the first scene construction module is specifically used for: The depth information corresponding to the first image is obtained based on the perceived data; A three-dimensional Gaussian sphere is constructed based on the depth information. The three-dimensional Gaussian sphere includes a scene of the mirror region reconstructed from the perspective of the mirror. The first spatial information includes the three-dimensional Gaussian sphere.
21. The apparatus of claim 20, wherein, The sensing data also includes first point cloud data, which includes data collected by radar. The first scene construction module is specifically used for: Extract a first feature from the first image and extract a second feature from the first point cloud data; By combining the first feature and the second feature, a third feature is obtained; The third feature is encoded and decoded using a codec to obtain the depth information.
22. The apparatus of claim 20 or 21, wherein, The first scene construction module is specifically used for: Determine the Gaussian parameters based on the depth information; The three-dimensional Gaussian sphere is constructed based on the Gaussian parameters.
23. The apparatus of claim 22, wherein, The first scene construction module is specifically used for: The first image is projected into a three-dimensional space based on the depth information to obtain second point cloud data, which is used to determine the center position of the three-dimensional Gaussian sphere. Alternatively, color features can be extracted from the first image, and color information in the three-dimensional Gaussian sphere can be determined based on the color features; Alternatively, ambient light information can be estimated based on the first image pair, and the illumination distribution of the three-dimensional Gaussian sphere can be determined based on the ambient light information, wherein the illumination distribution is used to determine the spherical harmonic coefficients; Alternatively, information about the elements in the first image can be obtained, and the information about the elements can be mapped to the space corresponding to the three-dimensional Gaussian sphere to obtain the element information in the three-dimensional Gaussian sphere; Alternatively, the specularity can be determined based on the first image, whereby the specularity represents the reflection intensity of a region within the three-dimensional Gaussian sphere, and the specularity also indicates whether a mirror exists in the first image. The Gaussian parameters include at least one of the center position, the color information, the spherical harmonic coefficient, the element information, or the reflection intensity.
24. The apparatus of claim 23, wherein, The second scene construction module is specifically used for: Information about the mirror region in the three-dimensional Gaussian sphere is obtained based on the specularity. Clustering is performed on the information in the three-dimensional Gaussian sphere of the mirror region to obtain a mirror Gaussian cluster; Based on the mirror coordinate system corresponding to the mirror region and its relative relationship with the first space, the information of the mirror region in the three-dimensional Gaussian sphere is projected onto the first space to obtain the second space information.
25. The apparatus of claim 23 or 24, wherein, The second scene construction module is specifically used for: Acquire historical data, which includes data collected during historical periods, and the historical data overlaps with the scenes contained in the perceived data; The motion speed of each element in the first image is determined based on the historical data, and the Gaussian parameter also includes the motion speed of each element.
26. The apparatus of claim 25, wherein, The device further includes: a path prediction module, used for: The historical path of each element is determined based on the historical data; Based on the first spatial information and the historical paths of each element, the movement paths of each element in future time periods are obtained.
27. A computing device, comprising: The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, wherein when the code is executed, the computing device performs the method as described in any one of claims 1-13.
28. A vehicle characterized by The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, wherein when the code is executed, the device performs the steps of the method as described in any one of claims 1-13.
29. A computer storage medium, comprising, The computer storage medium stores instructions that, when executed by the computer, cause the computer to perform the method according to any one of claims 1 to 13.
30. A computer program product, characterised in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 13.