Method and device for generating image, electronic equipment and computer program product

By generating target mesh models and adjusting the environment in a 3D virtual world, and combining the advantages of point clouds and mesh models, the problems of low image acquisition efficiency and insufficient accuracy in existing technologies are solved. This enables the generation of accurate and diverse images in a 3D virtual world, thereby improving the image acquisition capabilities of autonomous driving systems.

CN121982138APending Publication Date: 2026-05-05ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ROBERT BOSCH GMBH
Filing Date
2024-10-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to generate diverse and accurate images in a 3D virtual world, especially when simulating different weather and lighting conditions, resulting in low image acquisition efficiency and insufficient accuracy for autonomous driving systems.

Method used

By generating a target mesh model for the first point cloud, adjusting the environment within the target mesh model, and combining the initial image with the adjusted environmental information to generate a second image, the advantages of both are integrated to achieve accurate and diverse image acquisition in a 3D virtual world.

Benefits of technology

The generated images possess both the high accuracy of point cloud acquisition and the environmental adjustment capability of mesh models, enabling the generation of accurate and diverse images in a 3D virtual world, thus improving the efficiency and accuracy of image acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982138A_ABST
    Figure CN121982138A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for generating an image, electronic equipment and a computer program product. The method includes generating a target mesh model for a first point cloud and a first image corresponding to the first point cloud. The method further includes obtaining adjusted environmental information by adjusting the environment in the target mesh model. The method further includes generating a second image based on the first image and the adjusted environmental information. According to the method disclosed by the invention, the generated image not only has the advantage of high accuracy of the image acquired in the point cloud, but also has the capability of adjusting the environment, so that accurate and diversified images can be generated in the three-dimensional virtual world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing, and more specifically, to methods, apparatus, electronic devices, and computer program products for generating images. Background Technology

[0002] The field of autonomous driving is developing rapidly. Autonomous driving systems can identify and classify various objects on the road, such as other vehicles, pedestrians, road signs, and traffic lights. They can also predict the behavior of these objects, for example, determining whether a pedestrian will suddenly cross the road or whether a vehicle in front is preparing to change lanes. In addition, autonomous driving systems need to understand complex driving situations, such as making the correct turning decisions at intersections or maintaining lane stability on highways.

[0003] Autonomous driving systems perceive the external environment through various data sources, such as images. To enable autonomous vehicles to drive safely under diverse conditions, these systems need to handle images that encompass different weather conditions, road conditions, and traffic patterns. In a sense, the performance of an autonomous driving system depends on the diversity of the images provided. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, electronic device, and computer program product for generating images. In a first aspect of this disclosure, a method for generating an image is provided. The method includes generating a target mesh model for a first point cloud and a first image corresponding to the first point cloud. The method further includes obtaining adjusted environment information by adjusting the environment in the target mesh model. The method further includes generating a second image based on the first image and the adjusted environment information.

[0005] In a second aspect of this disclosure, a method for driving a vehicle is provided. The method includes acquiring a target image, wherein the target image is generated based on the initial image and environmental information, the initial image being generated based on a first point cloud, and the environmental information being obtained by adjusting an environment for a target mesh model of the first point cloud. The method also includes determining driving parameters of a simulated vehicle using a planning network based on a second image. The method further includes driving the vehicle by a driving system based on the driving parameters.

[0006] In a third aspect of this disclosure, a method for training a planning network is provided, the method comprising acquiring a target image, wherein the target image is generated based on an initial image and environment information, the initial image being generated based on a first point cloud, and the environment information being obtained by adjusting the environment for a target mesh model of the first point cloud. The method further comprises training the planning network using training samples including the target image.

[0007] In a fourth aspect of this disclosure, an apparatus is provided. The apparatus includes a first generation unit configured to generate a target mesh model for a first point cloud and a first image corresponding to the first point cloud. The apparatus further includes an information adjustment unit configured to obtain adjusted environmental information by adjusting the environment in the target mesh model. The apparatus also includes a second generation unit configured to generate a second image based on the first image and the adjusted environmental information.

[0008] In a fifth aspect of this disclosure, an electronic device is provided, comprising at least one processor and instructions coupled to the at least one processor and having instructions stored thereon, which, when executed by the at least one processor, cause the electronic device to perform the methods described in the first to third aspects of this disclosure.

[0009] In a sixth aspect of this disclosure, a computer program product is provided, which is tangibly stored on a non-transient computer-readable medium and includes machine-executable instructions that, when executed, cause a machine to perform the methods described according to the first to third aspects of this disclosure.

[0010] In a seventh aspect of this disclosure, a computer-readable storage medium is provided that stores machine-executable instructions thereon, wherein the machine-executable instructions are executed by a processor to implement the methods described according to the first to third aspects of this disclosure.

[0011] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0013] Figure 1 A schematic diagram illustrating an example application scenario of a method for rendering video according to an embodiment of this application is shown;

[0014] Figure 2 A flowchart of a method for generating an image according to an embodiment of the present disclosure is shown;

[0015] Figure 3 A schematic diagram illustrating the determination of a Gaussian point cloud according to an embodiment of the present disclosure is shown;

[0016] Figure 4AA schematic diagram of a method for generating an image according to an embodiment of the present disclosure is shown;

[0017] Figure 4B A schematic diagram of a method for driving a vehicle according to an embodiment of the present disclosure is shown;

[0018] Figure 5A An architecture diagram of generated and applied images according to embodiments of the present disclosure is shown;

[0019] Figure 5B A schematic diagram of a method for training a planning network according to an embodiment of the present disclosure is shown;

[0020] Figure 6 A schematic diagram of an apparatus for generating an image according to an embodiment of the present disclosure is shown; and

[0021] Figure 7 A schematic block diagram of an example device suitable for implementing embodiments of the present disclosure is shown.

[0022] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0024] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0025] To increase the diversity of images, it is usually necessary to dispatch personnel to use image sensors during actual driving to capture various situations that the vehicle may encounter. Sometimes, to improve driving capabilities in special weather conditions (such as rainy weather), it is necessary to acquire images specifically for those conditions. However, this method of image acquisition is very inefficient.

[0026] In related technologies, virtual 3D worlds can be constructed to improve efficiency. These virtual 3D worlds simulate different environments, such as varying weather and lighting, to collect images encompassing diverse conditions, serving as training samples. However, in 3D point cloud worlds built using methods like Gaussian sputtering, some environmental information cannot be edited, resulting in the inability to collect diverse images covering various conditions. Furthermore, the simulation level in constructed 3D mesh worlds is relatively low, leading to inaccurate image acquisition.

[0027] To address this, this disclosure proposes a method for generating images. The method includes generating a target mesh model for a first point cloud and a first image corresponding to the first point cloud. The method further includes obtaining adjusted environment information by adjusting the environment within the target mesh model. The method also includes generating a second image based on the first image and the adjusted environment information. According to embodiments of this disclosure, by establishing corresponding point clouds and mesh models, images can be acquired from the same perspective in two 3D virtual worlds. Furthermore, by adjusting the environment within the mesh model and then fusing the adjusted environment information into the image acquired from the point cloud, the generated image combines the high accuracy of images acquired from point clouds with the advantages of adjusting the environment, thereby enabling the generation of accurate and diverse images in a 3D virtual world.

[0028] Figure 1 A schematic diagram illustrating an example application scenario of a method for generating images according to embodiments of this application is shown. For example... Figure 1 As shown, the example environment 100 includes a server 102, a Gaussian point cloud 104 (i.e., a first point cloud), an initial image 106 (i.e., a first image), a target mesh model 108, adjusted environmental information 110, and a fused image 112 (i.e., a second image). In this embodiment, the method for generating the image is performed by the server 102.

[0029] Server 102 generates a target mesh model 108 for the Gaussian point cloud 104 and an initial image 106 corresponding to the Gaussian point cloud. A rendering engine can be deployed on server 102 and used to render the Gaussian point cloud to obtain a three-dimensional virtual world represented by the Gaussian point cloud 104. The initial image 106 can be acquired from the three-dimensional virtual world represented by the Gaussian point cloud 104. Figure 1 The initial image 106 shown is an image captured by the simulation image sensor of the simulated vehicle in the 3D virtual world. The target mesh model 108 and the Gaussian point cloud 104 have a corresponding relationship; for example, they can be generated based on the same material and indicate the same 3D virtual world.

[0030] Server 102 obtains adjusted environment information by adjusting the environment within the target mesh model 108. Server 102 can use a rendering engine to render the target mesh model 108, thereby obtaining a 3D virtual world represented by the target mesh model 108. Server 102 can then introduce modules (such as light source modules) into the rendering engine to adjust, or edit, this 3D virtual world. After adjustment, the changed information can be directly obtained from the rendering engine as adjusted environment information. For example, a lightmap can be baked using the rendering engine. Furthermore, server 102 generates a blended image 112 based on the initial image 106 and the adjusted environment information 110. This blended image 112 can be generated through image fusion.

[0031] As understood by those skilled in the art, instances of server 102 can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited herein. Server 102 can constitute part of a distributed system. It is understood that a distributed system is a system composed of multiple nodes, which can be computers, servers, or other processing nodes. They are interconnected through a network and work collaboratively. In a distributed system, users typically face a unified service entry point, while multiple nodes behind it jointly provide this service. These nodes can be located in different physical locations and communicate and coordinate through message passing. Distributed systems can process and store data and share this data among different nodes to achieve higher availability, reliability, and performance. Furthermore, distributed systems can be used to perform various tasks, including but not limited to data processing, storage management, and scientific computing.

[0032] The above combination Figure 1 An example environment 100 in which embodiments of this disclosure can be implemented and an application scenario for a method for generating images are described below. The following is in conjunction with... Figure 2 A flowchart describing a method 200 for generating an image according to an embodiment of the present disclosure is provided. This method can be performed by a server.

[0033] At box 202, a target mesh model for the first point cloud and a corresponding first image are generated. A point cloud is a collection of numerous independent points, each with precise 3D coordinates. Furthermore, each point may possess other information, such as transparency parameters and other parameters. This first point cloud is the underlying representation of a 3D virtual world composed of many point data points; rendering this first point cloud yields a visually realistic 3D virtual world. Point clouds typically capture complex shapes and high-precision surface details, but due to a lack of connectivity information between points, they are unsuitable for real-time rendering, making it difficult to edit some environmental information. The target mesh model is a 3D virtual world composed of many meshes. A mesh model is a structure composed of vertices, edges, and faces, with faces typically being triangles or quadrilaterals. Mesh models are suitable for real-time rendering because their structured data facilitates rapid processing and rendering; however, the rendered world is often inaccurate and unrealistic.

[0034] To effectively combine the advantages of point clouds and mesh models, embodiments of this disclosure establish a point cloud and mesh model with a corresponding relationship as the fusion basis, wherein the first point cloud and the target mesh model represent the same 3D virtual world. The first image is an image captured from a specific location in the 3D virtual world obtained by rendering the first point cloud from a specific perspective. For example, the 3D virtual world obtained by rendering the first point cloud can be captured using a simulated image sensor, and the resulting image is the first image.

[0035] At box 204, adjusted environment information is obtained by adjusting the environment in the target mesh model. This environment indicates at least a portion of the attribute information of the 3D virtual world represented by the target mesh model, such as lighting information, weather information, etc., in the target mesh model. For example, the target mesh model can be imported into a rendering engine. The adjusted environment information indicates the differences between before and after adjusting the target mesh model. For example, adding a weather module causes the color of some mesh faces in the target mesh model to change, resulting in different weather information in the rendered 3D virtual world, and these different weather information constitute the adjusted environment information. It should be noted that the obtained adjusted environment information should be the environment information acquired at the same location and from the same viewpoint as the first image. For example, if the first image is acquired at (1,1,1) in the 3D virtual world rendered based on the first point cloud from a specific viewpoint, then the environment information should be the weather information observed at (1,1,1) from the same viewpoint in the 3D virtual world obtained by rendering the target mesh model and after adjustment by the weather module.

[0036] At box 206, a second image is generated based on the first image and the adjusted environmental information. Because the first image depicts an accurate image, and the adjusted environmental information is edited or adjusted at specific attributes, the second image generated by fusing the two can combine the advantages of both. According to the method of embodiments of this disclosure, by establishing corresponding point cloud and mesh models, images can be acquired from the same perspective in two 3D virtual worlds. Furthermore, by adjusting the environment in the mesh model and then fusing the adjusted environmental information into the image acquired in the 3D virtual environment represented by the point cloud, the generated image possesses both the high accuracy of images acquired in the point cloud and the advantages of adjusting the environment, thereby enabling the generation of accurate and diverse images in the 3D virtual world.

[0037] In some embodiments, this disclosure provides examples of generating a Gaussian point cloud (i.e., a first point cloud) and a target mesh model. In some embodiments, multiple images are acquired in the real world using an image sensor (i.e., a first sensor), and an initial point cloud is acquired using other types of sensors (i.e., a second sensor). For example, a LiDAR sensor can be used to acquire a base point cloud for constructing a 3D virtual world. The embodiments also include generating a Gaussian point cloud based on multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud, and generating a target mesh model based on multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud.

[0038] In this embodiment, the Gaussian point cloud and the target mesh model are generated based on the exact same data. This ensures that the generated target mesh model is specific to the Gaussian point cloud, meaning both can be rendered into the same 3D virtual world. Here, "the same 3D virtual world" does not mean they are completely equivalent, but rather that they can have the same building layout, the same traffic routes, the same weather, etc. However, due to differences in rendering effects such as accuracy, they can have slightly different pixel distributions, slightly different transparency, etc. Furthermore, the Gaussian point cloud is generated based on an initial point cloud detected by the sensor through optimization operations, rather than a randomly initialized point cloud. Therefore, this improves the accuracy of the Gaussian point cloud, thereby improving the accuracy of the fused image.

[0039] Figure 3 A schematic diagram illustrating the determination of a Gaussian point cloud according to an embodiment of the present disclosure is shown. In this embodiment, the initial point cloud 302 may be obtained based on detection such as by lidar. At 304, a Gaussian point cloud is determined based on Gaussian distribution parameters and the initial point cloud. Points in the initial point cloud are converted into Gaussian ellipsoids according to the Gaussian distribution parameters, and these Gaussian ellipsoids constitute the Gaussian point cloud.

[0040] Next, the Gaussian point cloud is iteratively optimized using multiple images to obtain an accurate Gaussian point cloud for the 3D virtual world. Figure 3 In this embodiment, the Gaussian point cloud is projected at position 308 to obtain a two-dimensional mapping. Here, parameters from the image sensor at position 306 are needed to determine the portion of the Gaussian point cloud to be projected (i.e., the Gaussian ellipsoid to be projected). These parameters include the position of the image sensor.

[0041] For the 2D mapping obtained by projection, a differentiable rasterization process at position 310 can be used to render an intermediate image at position 312. This intermediate image corresponds to the viewpoint at a certain location in the current Gaussian point cloud. Then, from multiple images (real images previously acquired in the real world), the corresponding image at this location and viewpoint is used to determine the image loss of the intermediate image. If the image loss is greater than a threshold, the Gaussian distribution parameters, Gaussian density, and 2D projection are optimized.

[0042] The optimization process could, for example, involve calculating the optimization gradient based on the image loss from position 312 to 310. At position 314, the Gaussian density is adjusted according to the optimization gradient, and this adjusted Gaussian density is applied to the current Gaussian point cloud, thus pruning the current Gaussian point cloud. Furthermore, the 2D mapping is adjusted based on the optimization gradient from position 310 to 308. From position 308 to 306, the Gaussian distribution parameters are adjusted based on the optimized 2D mapping, thereby optimizing (i.e., updating) the current Gaussian point cloud based on the adjusted Gaussian distribution parameters. In short, the current Gaussian point cloud is updated based on the Gaussian density and Gaussian distribution parameters.

[0043] After iterative optimization, this embodiment also includes using the optimized Gaussian point cloud as the first point cloud. This first point cloud is also the current Gaussian point cloud obtained from the last optimization during the iterative optimization process. Figure 3 In one embodiment, a two-dimensional projection is determined at 308 based on the optimized Gaussian point cloud and the parameters of the image sensor, and an initial image (i.e., the first image) is obtained by performing differentiable rasterization on the two-dimensional projection at 310.

[0044] As mentioned above, although Gaussian point clouds obtained through Gaussian distributions cannot be edited, information for editing can be obtained through a mesh model, thereby enabling editing. In some embodiments, the target mesh model is imported into a rendering engine and rendered (e.g., through a rendering module) to obtain a rendered mesh model. This embodiment also includes extracting adjusted environmental information from the rendered mesh model.

[0045] As mentioned above, while the mesh model lacks high accuracy, it is editable. Therefore, this embodiment edits the target mesh model and extracts the changed information as adjusted environmental information. For example, after rendering a 3D virtual world using a rendering engine, a light source module can be called or imported. This light source module can be configured with parameters such as brightness and radiation angle. The light radiated by this light source module acts on the 3D virtual world, changing its lighting. This light information can then be extracted as adjusted environmental information. If a light source module is called or imported, a lightmap image can be extracted as adjusted environmental information. A lightmap is a technique used in computer graphics to simulate lighting effects. It is a texture map embedded in the surface of an object, recording the indirect lighting information of the light source module in the 3D virtual world. This can be achieved through baking techniques. Similarly, if a weather module is called or imported, a shadow channel image can be extracted as adjusted environmental information.

[0046] When the adjusted environmental information is in image format, a second image can be generated through a fusion operation. In some embodiments, each pixel in the initial image is assigned a first weight, each pixel in the adjusted environmental information is assigned a second weight, and the initial image and the adjusted environmental information are fused according to the first and second weights to generate a fused image.

[0047] According to some embodiments of this disclosure, the initial image has a high degree of realism and can accurately describe the real world. The adjusted environmental information allows for the editing of some attributes in the initial image. Therefore, by fusing the initial image and the adjusted environmental information in the above manner, some attributes of the initial image can be edited, enabling the fused image to describe a different real world. For example, modifying a high-fidelity image of a sunny day into a high-fidelity image of a rainy day.

[0048] The images generated by this disclosure can be used as training samples to train deep learning techniques for autonomous driving. To achieve this exemplary purpose, the acquired images should be captured from the perspective of the vehicle's image sensors. In some embodiments, a Gaussian point cloud is rendered to obtain a first three-dimensional virtual world, and a simulated vehicle is set up within this first three-dimensional virtual world. This embodiment also includes setting up image acquisition modules, such as simulated cameras, simulated radar, etc., for the simulated vehicle using a rendering engine. The acquisition position is then determined based on the position of the simulated vehicle, for example, the acquisition position could be coordinates (1, 1, 1), and the acquisition perspective is determined based on the viewpoint of the image acquisition module, for example, the orientation of the image acquisition module. This embodiment also includes generating a first image based on the three-dimensional virtual world according to the acquisition position and the acquisition perspective.

[0049] Figure 4A A schematic diagram of a method for generating an image according to an embodiment of the present disclosure is shown. Figure 4A The illustrated embodiment enables real-time interaction between a real-world vehicle and a simulated vehicle. The ROS (Robot Operating System) is deployed within the vehicle, assisting the user in controlling a series of sensors on the vehicle. In this embodiment, at 402, the ROS system uses image sensors to acquire multiple images from the real world. Additionally, sensors such as LiDAR can be used to acquire an initial point cloud for the real world. At 404, these multiple images, the initial point cloud, and the location information from each sensor are used to construct a map, including traffic flow data, to form an initial Gaussian point cloud for the real world. At 406, the initial Gaussian point cloud is optimized based on the acquired images to obtain an optimized Gaussian point cloud. For example, a Gaussian distribution can be used, based on... Figure 3 The optimization method optimizes the initial Gaussian point cloud to obtain an optimized Gaussian point cloud with high accuracy. As an alternative embodiment, a NeRF neural radiation field can be used to optimize the initial point cloud to improve its accuracy.

[0050] To enable the use of the images for vehicle applications during initial image acquisition, a simulated vehicle can be set up in the Gaussian point cloud at point 424, which can be controlled by the user via interface 422. A simulated camera is then set up within the simulated vehicle at point 426. At point 408, the simulated vehicle uses the simulated camera to determine its location within the Gaussian point cloud as the acquisition position, and uses a specific viewpoint of the simulated camera (e.g., a forward viewpoint) as the acquisition viewpoint. Based on the acquisition position and viewpoint, the initial image is generated by capturing images of this 3D virtual world.

[0051] exist Figure 4A In the illustrated embodiment, various data can be integrated into the plugin, allowing the plugin to be directly invoked in a conventional simulation system or rendering engine to implement the graph generation method according to embodiments of this disclosure. For example, at 416, the plugin can be imported to apply the data integrated within it. At 420, the simulation sensors to be used at 426 are invoked and imported into the simulated vehicle. Furthermore, the plugin can also integrate a local high-precision map, enabling the simulated vehicle to be equipped with an autonomous driving system, and at 418, allowing the autonomous driving system to invoke the local high-precision map to plan an autonomous driving route.

[0052] Furthermore, based on the initial point cloud acquired at point 402, multiple images, and the location information of each sensor, a target mesh model is constructed at point 412. At point 414, external modules (such as weather modules, lighting modules, etc.) loaded from the plugin at point 416 can then be invoked to render the target mesh model using the rendering engine, while simultaneously applying the effects of the external modules to obtain adjusted environmental information. Taking the lighting module as an example, the lighting module is used to illuminate the 3D virtual world to obtain the corresponding lightmap image. At point 410, the adjusted environmental information is fused with the initial image to obtain the fused image.

[0053] At point 428, a ROS interface is set up for interaction between the ROS system and the simulated vehicle. For example, it receives the fused image obtained at point 410. This enables interaction between the simulated vehicle and the real vehicle. At point 430, the fused image and the autonomous driving behaviors performed by the simulated vehicle based on the fused image are recorded as a behavior log. At point 432, the behavior logs generated within a certain time period are packetized for developers to inspect and analyze.

[0054] Furthermore, the Gaussian point cloud can be preprocessed to extract relationships between objects. In some embodiments, the Gaussian point cloud includes multiple 3D objects, such as traffic lights, buildings, and road railings. A relationship between the simulated vehicle and a specific 3D object can be extracted from the Gaussian point cloud, such as a positional relationship indicating the distance between the objects. A label for the simulated vehicle relative to the 3D object can then be generated based on this relationship, such as the distance corresponding to the positional relationship. Of course, the 3D object here should be one contained in the initial image. This allows for the generation of labels corresponding to the image simultaneously, which can be used for supervised training of deep learning techniques.

[0055] Figure 4BA schematic diagram of a method for driving a vehicle according to an embodiment of the present disclosure is shown. In block 440, a target image is acquired. The target image is generated based on an initial image and environmental information, wherein the initial image is generated based on a first point cloud, and the environmental information is obtained by adjusting the environment of a target mesh model for the first point cloud. In block 442, driving parameters of the simulated vehicle are determined using a planning network based on a second image. These driving parameters may include speed, direction, acceleration, etc. In this embodiment, the driving parameters may be provided to the vehicle's driving system via a data transmission interface, for example. The vehicle refers to a real vehicle, and the driving system refers to a system used to drive the vehicle. The driving parameters can be provided according to the ROS interface disclosed in the above embodiments. Autonomous driving of the vehicle can be achieved through this driving system. In block 444, the vehicle's driving system drives the vehicle according to the driving parameters. According to the method of the present disclosure, driving parameters in a virtual world can be applied to a real vehicle, allowing the real vehicle to simulate the driving process of the simulated vehicle. Therefore, by modifying the environment of the virtual world, developers can understand the driving state of the real vehicle in different environments (e.g., different weather conditions).

[0056] This disclosure also provides exemplary embodiments for generating and applying images. Figure 5A An architecture diagram of generated and applied images according to embodiments of the present disclosure is shown. In this embodiment, the three-dimensional virtual environment 502 includes multiple simulation data, such as three-dimensional objects 504, other vehicles 506, pedestrians 508, lighting 510, and weather 512. These simulation data can be obtained by rendering Gaussian point clouds. The simulation sensing module 514 includes a simulation camera 516, a simulation radar 518, a simulation laser 520, and a simulation GPS 522. These simulation sensors have the same functions as their corresponding real sensors. These simulation sensors are used to acquire corresponding initial images and transmit these images to the perception module 526, which applies them to generate fused images according to the above embodiments of the present disclosure. In addition, the simulation sensing module 514 also includes map data 524, which contains various positional relationships in the map. This map data 524 is transmitted to the perception module 526 to determine the acquisition location 536 of the simulated vehicle.

[0057] The perception module 526 can generate a target mesh model for the Gaussian point cloud and receive an initial image (i.e., a first image) corresponding to the Gaussian point cloud. The perception module 526 also obtains adjusted environmental information by adjusting the environment within the target mesh model, and generates a fused image (i.e., a second image) based on the initial image and the adjusted environmental information. Furthermore, the perception module 526 analyzes the fused image to perform road detection 528, traffic light detection 530, traffic sign detection 532, pedestrian detection 534, and determine the acquisition location 536. Road detection 528 detects the layout of nearby roads, traffic light detection 530 detects the status of traffic lights around the simulated vehicle, traffic sign detection 532 detects the content of traffic signs around the simulated vehicle, pedestrian detection 534 detects the distribution of pedestrians around the simulated vehicle, and the acquisition location 536 indicates the location of the simulated vehicle.

[0058] The perception module 526 transmits the collected information (including fused images and information obtained from road detection 528, traffic light detection 530, traffic sign detection 532, pedestrian detection 534, and location determination 536) to the planning module 538. Based on this information, the planning module 538 determines driving parameters for the simulated vehicle, including path planning 540, other vehicle prediction 542, behavior planning 544, and trajectory planning 546. Path planning 540 plans the road the simulated vehicle will travel; other vehicle prediction 542 plans possible other vehicles ahead and their positions; behavior planning 544 plans whether the simulated vehicle should merge into or leave the traffic flow currently in which it is located; and trajectory planning 546 plans the specific route for the simulated vehicle to merge into or leave the traffic flow.

[0059] The control module 548 receives driving parameters from the planning module 538 and uses these parameters to control the driving of the simulated vehicle via the simulation controller 550. Furthermore, related control commands can be transmitted via an interface to the remote control system 552 of the real vehicle to apply these driving parameters back to the real vehicle, enabling the real vehicle to simulate the driving of the simulated vehicle. This embodiment helps developers understand the driving status of the real vehicle in different environments (e.g., different weather conditions).

[0060] Figure 5BA schematic diagram of a method for training a planning network according to an embodiment of the present disclosure is shown. The planning network provides a driving trajectory for a simulated or real vehicle. In box 560, a target image is acquired. The target image is generated based on an initial image and environmental information, i.e., the second image generated in the above embodiment. The initial image is generated based on a first point cloud, and the environmental information is obtained by adjusting the environment of the target mesh model for the first point cloud. Additionally, if supervised training of the planning network is required, labels corresponding to the training samples can be determined as part of the training samples. For example, for an image of an intersection, labels can be added indicating which direction the simulated vehicle should continue driving. In box 562, the planning network is trained using training samples containing the target image. This training can be supervised or unsupervised. If unsupervised training of the planning network is required, labels do not need to be determined; training samples can be used directly. For example, data processing techniques can be used to rotate or invert the target image. If supervised training of the planning network is required, a loss can be calculated based on the prediction results and labels of the planning network, and multiple weights of the planning network can be optimized based on the loss to make the predictions of the planning network more accurate. According to the method for training the planning network disclosed herein, a second image with higher realism can be used to train the planning network, thereby improving the performance of the planning network.

[0061] Figure 6 This is a schematic diagram of an apparatus for generating an image according to an embodiment of the present disclosure. Figure 6 The apparatus 600 shown includes a first generation unit 602 configured to generate a target mesh model for a first point cloud and a first image corresponding to the first point cloud. The apparatus 600 also includes an information adjustment unit 604 configured to obtain adjusted environmental information by adjusting the environment in the target mesh model. The apparatus 600 further includes a second generation unit 606 configured to generate a second image based on the first image and the adjusted environmental information.

[0062] In some embodiments, the first generation unit 602 includes a first acquisition unit configured to acquire multiple images using a first sensor. The first generation unit 602 also includes a second acquisition unit configured to determine an initial point cloud using a second sensor. The first generation unit 602 further includes a third generation unit configured to generate a first point cloud based on the multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud. The first generation unit 602 also includes a fourth generation unit configured to generate a target mesh model based on the multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud.

[0063] In some embodiments, the third generation unit includes a first determining unit configured to determine a Gaussian point cloud based on Gaussian distribution parameters and an initial point cloud. The third generation unit further includes an iteration unit configured to iteratively optimize the Gaussian point cloud using multiple images. The third generation unit also includes a second determining unit configured to determine the optimized Gaussian point cloud as the first point cloud.

[0064] In some embodiments, the iteration unit includes a third determining unit configured to determine a two-dimensional projection of the Gaussian point cloud based on the position of the first sensor. The iteration unit further includes a raster processing unit configured to perform differentiable rasterization processing on the two-dimensional projection to obtain an intermediate image. The iteration unit further includes a loss calculation unit configured to determine an image loss of the intermediate image based on multiple images. The iteration unit further includes a pruning unit configured to adjust the Gaussian density and Gaussian distribution parameters based on the image loss in response to an image loss greater than a threshold. The iteration unit further includes an updating unit configured to update the Gaussian point cloud based on the Gaussian density and Gaussian distribution parameters.

[0065] In some embodiments, the information adjustment unit 604 includes an import unit configured to import a target mesh model into a rendering engine. The information adjustment unit 604 also includes a first rendering unit configured to render the target mesh model in the rendering engine to obtain a rendered mesh model. The information adjustment unit 604 includes an extraction unit configured to extract adjusted environmental information from the rendered mesh model.

[0066] In some embodiments, the format of the adjusted environmental information is an image format, and the second generation unit 606 includes a first allocation unit configured to assign a first weight to each pixel in the first image. The second generation unit 606 also includes a second allocation unit configured to assign a second weight to each pixel in the adjusted environmental information. The second generation unit 606 includes a fusion unit configured to fuse the first image and the adjusted environmental information according to the first weight and the second weight to generate a second image.

[0067] In some embodiments, the apparatus 600 further includes a second rendering unit configured to render a first point cloud to obtain a first three-dimensional virtual world. The apparatus 600 also includes a simulated vehicle deployment unit configured to set up a simulated vehicle in the first three-dimensional virtual world. The first generation unit 602 includes a setting unit configured to use a rendering engine to set up an image acquisition module for the simulated vehicle. The first generation unit 602 further includes a fourth determining unit configured to determine an acquisition position based on the position of the simulated vehicle. The first generation unit 602 further includes a fifth determining unit configured to determine an acquisition viewpoint based on the viewpoint of the image acquisition module. The first generation unit 602 further includes a fourth generation unit configured to generate a first image based on the first three-dimensional virtual world according to the acquisition position and the acquisition viewpoint.

[0068] In some embodiments, the first point cloud includes a plurality of three-dimensional objects, and a first three-dimensional object among the plurality of three-dimensional objects is captured in a first image. The device 600 further includes a second extraction unit configured to extract a first association relationship between the simulated vehicle and the first three-dimensional object based on the first point cloud. The device 600 also includes a fifth generation unit configured to generate a label for the simulated vehicle relative to the first three-dimensional object based on the first association relationship.

[0069] Figure 7 A schematic block diagram of a controller 700 suitable for implementing embodiments of the present disclosure is shown. As shown, the controller 700 includes a processor 701, which performs various appropriate actions and processes based on computer program instructions loaded into random access memory (RAM) 703 according to computer program instructions stored in read-only memory (ROM) 702. Various programs and data required for the operation of the controller 700 may also be stored in RAM 703. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0070] The various methods and processes described above can be executed by processor 701. For example, in some embodiments, the various methods and processes described above can be implemented as computer software programs tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded into and / or installed onto controller 700 via ROM 702. When the computer program is loaded into RAM 703 and executed by processor 701, one or more actions of the methods and processes described above can be performed.

[0071] This disclosure can be a method, apparatus, system, electronic device, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0072] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), and any suitable combination thereof. The computer-readable storage medium as used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0073] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0074] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0075] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0076] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0077] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0079] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method (200) for generating an image, comprising: Generate (202) a target mesh model for the first point cloud and a first image corresponding to the first point cloud; (204) Adjusted environmental information is obtained by adjusting the environment in the target mesh model; and A second image is generated (206) based on the first image and the adjusted environmental information.

2. The method according to claim 1, wherein generating (202) a target mesh model for the first point cloud comprises: Multiple images were acquired using the first sensor; The initial point cloud was acquired using a second sensor; The first point cloud is generated based on the multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud. as well as The target mesh model is generated based on the multiple images, the position of the first sensor, the position of the second sensor, and the initial point cloud.

3. The method according to claim 2, wherein generating the first point cloud based on the plurality of images, the position of the first sensor, the position of the second sensor, and the initial point cloud comprises: Determine the Gaussian point cloud based on the Gaussian distribution parameters and the initial point cloud; The Gaussian point cloud is iteratively optimized using the multiple images; as well as The optimized Gaussian point cloud is determined as the first point cloud.

4. The method of claim 3, wherein iteratively optimizing the Gaussian point cloud using the plurality of images comprises: The two-dimensional projection of the Gaussian point cloud is determined based on the position of the first sensor; The two-dimensional projection is processed into a differential rasterization to obtain an intermediate image; Determine the image loss of the intermediate image based on the plurality of images; In response to the image loss exceeding a threshold, the Gaussian density and the Gaussian distribution parameters are adjusted based on the image loss; and The Gaussian point cloud is updated based on the Gaussian density and the Gaussian distribution parameters.

5. The method of claim 1, wherein obtaining (204) adjusted environment information by adjusting the environment in the target mesh model comprises: Import the target mesh model into the rendering engine; The target mesh model is rendered in the rendering engine to obtain a rendered mesh model; as well as The adjusted environmental information is extracted from the rendered mesh model.

6. The method of claim 5, wherein the format of the adjusted environmental information is an image format, and generating (206) a second image based on the first image and the adjusted environmental information comprises: Assign a first weight to each pixel in the first image; Each pixel in the adjusted environmental information is assigned a second weight; as well as The first image and the adjusted environmental information are fused according to the first weight and the second weight to generate the second image.

7. The method according to claim 1, further comprising: Render the first point cloud to obtain the first three-dimensional virtual world; Set up a simulated vehicle in the first three-dimensional virtual world; And generating (202) the first image corresponding to the first point cloud includes: Use a rendering engine to set up an image acquisition module for the simulated vehicle; The data collection location is determined based on the position of the simulated vehicle. The acquisition angle is determined based on the viewing angle of the image acquisition module; and The first image is generated based on the first three-dimensional virtual world according to the acquisition location and the acquisition perspective.

8. The method according to claim 7, wherein the first point cloud comprises a plurality of three-dimensional objects, a first three-dimensional object among the plurality of three-dimensional objects is captured in the first image, and the method further comprises: Based on the first point cloud, extract the first association relationship between the simulated vehicle and the first three-dimensional object; as well as The simulated vehicle is labeled relative to the first three-dimensional object based on the first association relationship.

9. A method for driving a vehicle, comprising: Acquire (440) a target image, wherein the target image is generated based on an initial image and environmental information, the initial image being generated based on a first point cloud, and the environmental information being obtained by adjusting the environment for a target mesh model of the first point cloud; Based on the target image, the driving parameters of the simulated vehicle are determined using a (442) planning network; as well as The vehicle is driven (444) by the vehicle's driving system according to the driving parameters.

10. A method for training a planning network, comprising: Acquire (560) a target image, wherein the target image is generated based on an initial image and environmental information, the initial image being generated based on a first point cloud, and the environmental information being obtained by adjusting the environment for a target mesh model of the first point cloud; as well as The planning network is trained (562) using training samples including the target image.

11. An apparatus (600) for generating an image, comprising: The first generation unit (602) is configured to generate a target mesh model for the first point cloud and a first image corresponding to the first point cloud; The information adjustment unit (604) is configured to obtain adjusted environmental information by adjusting the environment in the target mesh model; as well as The second generation unit (606) is configured to generate a second image based on the first image and the adjusted environmental information.

12. An electronic device, comprising: At least one processor; as well as Coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-10.

13. A computer program product tangibly stored on a non-transient computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform the method according to any one of claims 1 to 10.