Image generation method and device, intelligent agent and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YINWANG INTELLIGENT TECHNOLOGIES CO LTD
- Filing Date
- 2024-09-10
- Publication Date
- 2026-05-19
AI Technical Summary
Intelligent agents suffer from low accuracy in environmental perception, especially when there is camera malfunction, image frame loss, or image quality that does not meet standards. This can lead to blind spots or reduced accuracy in BEV perception and detection, or even cause the vehicle to exit the intelligent driving state.
By identifying cameras that have malfunctioned, lost frames, or whose image quality does not meet standards, images from the target camera's perspective are generated using image data from other cameras and LiDAR data to replace false images, thereby improving the accuracy and robustness of the BEV perception algorithm.
This design avoids blind spots in BEV perception caused by camera failure or image frame loss, even without camera redundancy, thus improving environmental perception accuracy, the robustness of the intelligent driving system, and the user experience.
Smart Images

Figure CN122070701A_ABST
Abstract
Description
Image generation method and device, agent and storage medium TECHNICAL FIELD
[0001] The present application relates to the field of image processing, in particular to an image generation method and device, an agent and a storage medium. BACKGROUND
[0002] An agent is an entity in the field of computer science and artificial intelligence that can perceive the environment and take action to achieve a certain goal. For example, an agent can be a vehicle, a robot, a drone, etc. A visual sensing system is an important sensing system for an agent, providing the agent with comprehensive environmental perception capabilities. For example, a vehicle can be equipped with a single-mode multi-view camera, or a multi-modal multi-view sensor such as a pinhole camera, a surround-view camera, a laser radar, a millimeter wave radar, etc.
[0003] However, in some cases, the accuracy of the agent's environmental perception is not high.
[0004] SUMMARY
[0005] The present application provides an image generation method and device, an agent and a storage medium to improve the accuracy of the agent's environmental perception.
[0006] In a first aspect, the present application provides an image generation method, the steps of which can be executed by an agent, or the method can be executed by a component (such as a chip, a chip system, etc.) configured in the agent, or it can also be implemented by a logic module or software capable of implementing all or part of the agent function, or it can also be executed by an image generation device, which is not limited by the present application.
[0007] For example, the method includes determining a target camera on the agent that meets a target condition, the agent including a plurality of cameras; generating an image under the target camera's perspective based on target data, the target data including images collected by other cameras on the agent other than the target camera; wherein the target condition includes one or more of the following: the camera has failed; or, the camera has experienced image frame loss; or, the image collected by the camera does not meet a predetermined image quality standard.
[0008] Based on the above technical solution, the image under the corresponding perspective is generated based on the target condition, which can achieve bird's eye view (BEV) global non-blind area perception detection without camera redundancy design backup, improving the BEV perception algorithm accuracy and robustness, i.e., improving the accuracy of the agent's environmental perception.
[0009] In a possible implementation manner of the first aspect, the method further includes: obtaining a fault camera code, the fault camera code being used to determine the camera that is in fault.
[0010] According to the fault code, the camera that is in fault (i.e., the fault camera) is determined, and then the image under the view angle of the fault camera is generated, so that the generated image under the view angle of the fault camera is used to replace the false image (i.e., the image that does not contain real-time environmental information), which can avoid the problem that the intelligent driving state is exited due to the camera being in fault, thereby guaranteeing the robustness of the intelligent driving system.
[0011] In a possible implementation manner of the first aspect, the agent is deployed with a BEV perception module, the BEV perception module being used for the agent to perceive the environment in which the agent is located; and the method further includes: detecting, at a preset time, whether there is at least one image of a camera that does not arrive at the BEV perception module on the agent; and in a case where there is at least one image of a camera that does not arrive at the BEV perception module, determining that the image of the at least one camera is in image frame loss.
[0012] In view of the fact that there may be image frame loss in an actual application scenario, the camera that is in image frame loss (i.e., the frame loss camera) is determined according to the image frame loss, and then the image under the view angle of the frame loss camera is generated, so that the generated image under the view angle of the frame loss camera is used to replace the false image, which can avoid the problem that the BEV perception detection has a blind area or low detection accuracy due to the image frame loss.
[0013] In a possible implementation manner of the first aspect, the method further includes: performing quality score calculation on the images collected by the cameras on the agent respectively; and determining, as the image that does not satisfy the preset image quality standard, the image whose quality score is lower than a preset threshold.
[0014] In view of the fact that there may be low-quality images (i.e., images that do not satisfy the preset image quality standard) collected in an actual application scenario, the low-quality image camera (i.e., the camera that collects the low-quality image) is determined according to the low-quality image, and then the image under the view angle of the low-quality image camera is generated, so that the generated image under the view angle of the low-quality image camera is used to replace the low-quality image, which can avoid the problem that the BEV perception detection has a blind area or low detection accuracy due to the low image quality.
[0015] In a possible implementation of the first aspect, the target data further includes one or more of the following: a code of the target camera; or, Lidar data; or, an image before the target condition is met in the target camera view; or, an image before the target condition is met in another camera view; or, a BEV perception result before the target condition is met; or, a moving speed of the agent when the target condition is met; or, a moving direction of the agent when the target condition is met.
[0016] In a possible implementation of the first aspect, the method further includes obtaining a BEV perception result based on the generated image in the target camera view and images captured by other cameras on the agent except the target camera.
[0017] In a possible implementation of the first aspect, the method further includes obtaining a BEV perception result based on the generated image in the target camera view, images captured by other cameras on the agent except the target camera, and Lidar data.
[0018] In a possible implementation of the first aspect, the method further includes displaying the generated image in the target camera view.
[0019] In a possible implementation of the first aspect, displaying the generated image in the target camera view includes displaying the generated image in the target camera view and images captured by other cameras based on a 360° surround view system or a reversing image system.
[0020] Displaying the generated image on the user interface instead of the fake image can improve user experience.
[0021] In a second aspect, an image generation apparatus is provided. The image generation apparatus includes modules for performing the steps of the method in the first aspect and possible implementation manners of the first aspect. The apparatus includes corresponding modules for performing the steps of the method. The modules included in the apparatus can be implemented by software and / or hardware.
[0022] In a third aspect, an image generation apparatus is provided. The image generation apparatus includes a processor. The processor is coupled with a memory and is configured to execute a program in the memory to implement the steps of the method in the first aspect and possible implementation manners of the first aspect.
[0023] Optionally, the image generation apparatus further includes the memory.
[0024] Optionally, the image generation apparatus further comprises a communication interface, and the processor is coupled to the communication interface.
[0025] In a fourth aspect, the present application provides a chip system, which comprises at least one processor configured to support the functions involved in the first aspect and any possible implementation manner of the first aspect, such as receiving or processing the data and / or indication information involved in the above method.
[0026] In a possible design, the chip system further comprises a memory configured to store program instructions and data, and the memory is located in the processor or outside the processor.
[0027] The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0028] In a fifth aspect, the present application provides an intelligent entity, which comprises a module configured to perform the execution steps in the first aspect and any possible implementation manner of the first aspect. The intelligent entity comprises a module configured to perform the above method. The modules included in the intelligent entity can be implemented by software and / or hardware.
[0029] In a sixth aspect, the present application provides an intelligent entity, which comprises the image generation apparatus in the second aspect or the third aspect.
[0030] Optionally, the intelligent entity in the fifth aspect or the sixth aspect can include a vehicle, a robot, a drone or a smart phone, etc.
[0031] In a seventh aspect, the present application provides a computer readable storage medium, which stores a program (also referred to as code or instructions), and when the program is run by a processor, the method in the first aspect and any possible implementation manner of the first aspect is executed.
[0032] In an eighth aspect, the present application provides a computer program product, which comprises a computer program (also referred to as code or instructions), and when the computer program is run, the method in the first aspect and any possible implementation manner of the first aspect is executed.
[0033] It should be understood that the second aspect to the eighth aspect of the present application correspond to the technical solution of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding possible implementation manner are similar, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0034] FIG. 1 is a schematic diagram of a multi-modal multi-view sensor deployment mode;
[0035] FIG. 2 is a schematic flowchart of an image generation method according to an embodiment of the present application;
[0036] FIG. 3 is a schematic flowchart of an image generation method based on fault codes according to an embodiment of the present application;
[0037] FIG. 4 is a schematic diagram of input data and output data of a multi-modal diffusion model according to the present application;
[0038] FIG. 5 is another schematic flowchart of an image generation method based on fault codes according to an embodiment of the present application;
[0039] FIG. 6 is a schematic diagram of input data and output data of a single-modal diffusion model according to the present application;
[0040] FIG. 7 is a schematic diagram of data interaction between modules suitable for the image generation method according to the present application;
[0041] FIG. 8 is a schematic flowchart of an image generation method based on image frame loss according to an embodiment of the present application;
[0042] FIG. 9 is another schematic flowchart of an image generation method based on image frame loss according to an embodiment of the present application;
[0043] FIG. 10 is a schematic flowchart of an image generation method based on image quality according to an embodiment of the present application;
[0044] FIG. 11 is another schematic flowchart of an image generation method based on image quality according to an embodiment of the present application;
[0045] FIG. 12 is a schematic block diagram of an image generation apparatus according to an embodiment of the present application;
[0046] FIG. 13 is another schematic block diagram of an image generation apparatus according to an embodiment of the present application. DETAILED DESCRIPTION
[0047] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0048] First, in the present application, the terms “comprising” and “having” and any variations thereof are intended to cover non-exclusive inclusion, for example, an apparatus, system, product or device comprising a series of modules, modules or units does not have to be limited to those clearly listed, but can include other modules, modules or units that are not clearly listed or inherent to the apparatus, system, product or device.
[0049] Secondly, in the present application, the words "exemplarily", "for example" and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplarily", "for example" and the like are intended to present the relevant concept in a specific manner.
[0050] Thirdly, in the present application, "when", "in the case of", "if" and "whether" all mean that the device will make corresponding processing under certain objective condition, and are not limited to time, and do not require the device to have a judgment action when implemented, nor does it mean that there are other limitations.
[0051] Fourthly, in the present application, "first" and "second" and the like are used to distinguish the same items or similar items with basically the same function and effect. For example, the first target condition and the second target condition are used to distinguish different target conditions, and do not limit the order. Those skilled in the art can understand that "first" and "second" and the like do not limit the quantity and execution order, and "first" and "second" and the like do not necessarily mean different.
[0052] Fifthly, in the present application, "preset" can be understood as predefinition, definition, predefinition, storage, prestorage, prenegotiation or preconfiguration and the like.
[0053] Sixthly, in the present application, "at least one" means one or more. "And / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it, but does not rule out the case that the associated objects before and after it represent an "and" relationship. The specific meaning can be understood in combination with the context.
[0054] Seventhly, in the present application, information C is used for the determination of information D, which includes that information D is determined based on information C only, and information D is determined based on information C and other information. In addition, information C used for the determination of information D can also include the case of indirect determination, such as the case that information D is determined based on information E, and information E is determined based on information C.
[0055] Eighthly, in the present application, the camera can also be understood as a video camera or a camera.
[0056] In order to facilitate the understanding of the embodiments of the present application, some technical terms or words involved in the present application are briefly described below.
[0057] 1、BEV: In autonomous driving and advanced driver assistance systems (ADAS), BEV is a commonly used perception method that fuses data from multiple cameras (or cameras) or sensors to generate a top-down view. This perspective can provide a more comprehensive environmental perception, which helps intelligent agents (such as vehicles) to perform path planning, obstacle detection and obstacle avoidance tasks.
[0058] In the embodiments of the present application, the BEV perception result can be understood as a BEV-based perception result. As an example but not limitation, the BEV perception result can include one or more of the following:
[0059] 1) Obstacle detection: Identify and locate obstacles around the intelligent agent, such as other intelligent agents, pedestrians, bicycles, roadblocks, etc.
[0060] 2) Road signs and markings: Detect and identify signs and markings on the road, such as lane lines, stop signs, speed limit signs, etc.
[0061] 3) Free space detection: Identify areas where the intelligent agent can safely move, usually used for path planning and obstacle avoidance.
[0062] 4) Dynamic object tracking: Track moving objects around, such as pedestrians and other intelligent agents, to predict their motion trajectories.
[0063] 5) Traffic signal recognition: Identify the state of traffic signal lights (red, green, yellow) and other traffic control devices.
[0064] 6) Environment classification: Classify the surrounding environment, such as distinguishing roads, pedestrian areas, buildings, green belts, etc.
[0065] 7) Intelligent agent's own pose and position: Determine the specific position and pose of the intelligent agent itself, including yaw angle, roll angle, etc.
[0066] 2、Image diffusion model: It is a machine learning technique used to generate and process images, especially in the field of computer vision and image processing. It usually involves gradually improving the quality or characteristics of an image through a series of iterative steps. In this application, for ease of description, the image diffusion model is simply referred to as the diffusion model, that is, the diffusion model, single-modal diffusion model and multi-modal diffusion model mentioned in the following refer to the image diffusion model.
[0067] 3. 360° ring car video system: also known as panoramic surround view system or panoramic monitoring system, through the installation of multiple cameras around the vehicle, the images around the vehicle are captured, and the panoramic view around the vehicle can also be provided for the driver, helping the driver better understand the environment around the vehicle, thereby improving driving safety.
[0068] 4. Reverse image system: a kind of auxiliary driving system, mainly used to provide rear view when the vehicle is reversing, helping the driver to operate the vehicle more safely.
[0069] 5. Single modality multi-view: refers to obtaining information through multiple views (such as multiple cameras or cameras) under a single modality (such as vision). This method can provide more rich and comprehensive data, which helps to improve the performance and accuracy of the system.
[0070] 6. Multi-modal multi-view: multi-modal multi-view system combines multiple input modes (such as vision, radar, etc.) and multiple views (such as multiple cameras (or cameras) or sensors) to provide more comprehensive and accurate information.
[0071] 7. Lidar data: Lidar is a technology that uses laser ranging technology to obtain high-precision three-dimensional spatial information. Lidar data is usually generated by laser scanners, global positioning systems (GPS), inertial measurement units (IMU), and other devices, and is widely used in geographic information systems (GIS), autonomous driving, unmanned aerial vehicle surveying and mapping, archaeology, environmental monitoring, and other fields.
[0072] 8. Agent: in the field of computer science and artificial intelligence, an agent refers to an entity that can perceive the environment and take actions to achieve a certain goal. For example, an agent can be a vehicle, a robot, a drone, etc.
[0073] With the development of artificial intelligence related technologies, people have higher and higher requirements for the perception ability of agents. Visual sensing system is an important sensing system for agents, which provides comprehensive environmental perception ability for agents. Taking a vehicle as an example, intelligent driving perception algorithm is an important part of intelligent driving system, and the perception ability of the vehicle to the environment directly affects the safety performance of the intelligent driving vehicle. Single modality multi-view cameras can be deployed on the vehicle; multi-modal multi-view sensors such as pinhole cameras, surround view cameras, laser radars, millimeter wave radars, etc. can also be deployed.
[0074] Figure 1 is a schematic diagram of a multi-modal multi-view sensor deployment mode.
[0075] Exemplarily, as shown in the deployment mode of the multi-modal multi-view sensor in FIG. 1, various sensors can be deployed on a vehicle, and each sensor can be deployed one or more, for example, 1 lidar, 1 long-range camera, 6 short-range cameras, 4 fisheye cameras, 4 millimeter wave radars, 8 short-range ultrasonic radars, and 4 long-range ultrasonic radars are deployed on the vehicle as shown in FIG. 1.
[0076] To improve the detection accuracy and robustness of intelligent driving perception algorithms, in some implementations, BEV perception based on single-modal multi-view fusion (e.g., images under multiple camera views), and BEV perception based on multi-modal multi-view fusion (e.g., images under multiple camera views, and data collected by lidar and millimeter wave radar, etc.).
[0077] Based on the current BEV perception algorithm, on the one hand, when there is a bottleneck in the bandwidth of the intelligent driving computing platform, causing frame loss of data from a certain sensor, the detection result of the single-modal multi-view or multi-modal multi-view fusion perception algorithm will also be lost accordingly. A possible implementation for this scenario is to use an image that does not contain real-time environmental information (e.g., a completely black image, a completely white image, or a completely random noise image, for ease of description, such an image is referred to as a false image in this application) to replace the lost frame image, and then input the false image and other non-lost frame images into the BEV perception algorithm. As a result, there will be a blind area in the "lost frame camera" view of the BEV perception result, that is, objects in the view of the camera that has lost frames will be missed. On the other hand, when a certain sensor fails to normally collect data, the single-modal multi-view or multi-modal multi-view fusion perception algorithm cannot obtain data that meets the requirements, which may cause problems such as exiting the intelligent driving state. On the other hand, when the data quality collected by a certain sensor is affected by external light or external dirt, etc., and does not meet the pre-set image quality standard, the detection result accuracy of the single-modal multi-view or multi-modal multi-view fusion perception algorithm will also decrease. In summary, in some cases, the accuracy of the intelligent agent's perception of the environment will be low.
[0078] To solve the above problems, an image generation method is provided in the embodiments of the present application. By determining a camera that has failed, has lost frames, or has collected images that do not meet the standard (for ease of description, the camera is referred to as a target camera), and combining the images collected by other cameras on the intelligent agent where the target camera is located, an image under the view of the target camera is generated. In this way, not only can the quality of the image be guaranteed and the accuracy of the BEV perception result be improved, but also the problem of exiting the intelligent driving state due to the failure of the camera can be avoided, and the robustness of the intelligent driving system can be improved.
[0079] The image generation method, device, intelligent agent and storage medium provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0080] FIG. 2 is a schematic flowchart of the image generation method provided by the embodiments of the present application.
[0081] The method 200 shown in FIG. 2 can include step 210 and step 220.
[0082] The steps of the method can be executed by an intelligent agent, or the method can be executed by a component (such as a chip, a chip system, etc.) configured in the intelligent agent, or can also be implemented by a logic module or software capable of implementing all or part of the functions of the intelligent agent, or can also be executed by an image generation device, which is not limited in the present application. The following will take the intelligent agent as an example to describe each step in the method 200 in detail.
[0083] In step 210, a target camera satisfying a target condition on the intelligent agent is determined, and the intelligent agent includes a plurality of cameras.
[0084] The target condition includes one or more of the following: the camera fails; or the camera loses image frames; or the image collected by the camera does not meet the preset image quality standard.
[0085] As shown in FIG. 1, for example, a plurality of cameras can be deployed on the intelligent agent. The intelligent agent can determine the target camera on the intelligent agent based on the target condition, and the target camera can include at least one of the following: a camera that fails, a camera that loses image frames, or a camera that collects an image that does not meet the preset image quality standard.
[0086] It can be understood that the number of target cameras can be one or more. For example, there are 11 cameras deployed on the intelligent agent, of which 1 camera fails, 2 cameras lose frames, and 1 camera collects an image that does not meet the preset image quality standard, so there are a total of 4 (obtained by 1+2+1) target cameras in the 11 cameras.
[0087] It can also be understood that the cameras deployed on the intelligent agent can periodically collect image data. As an example but not limitation, for example, the frequency of the cameras on the intelligent agent collecting images is 30 hertz (Hz), that is, an image is collected approximately every 33 milliseconds (ms), so the collection period is 33 ms. Accordingly, in each collection period, the intelligent agent can determine whether there is a target camera, that is, whether there is at least one camera that fails, or whether there is at least one camera that loses image frames, or whether there is at least one camera that collects an image that does not meet the preset image quality standard.
[0088] In step 220, an image under the target camera view is generated based on target data, the target data including images captured by other cameras on the agent other than the target camera.
[0089] The images captured by other cameras on the agent other than the target camera can be understood as images captured by other cameras on the agent other than the target camera within a data capture period (which can be referred to as a capture period for short) at a time when the target condition is met.
[0090] Exemplarily, the agent can input the target data into a diffusion model, and the diffusion model processes and analyzes the target data to obtain the image under the target camera view.
[0091] For brevity, details of the diffusion model are not repeated here, and can be referred to the relevant description in the technical terms section above.
[0092] Optionally, the target data further includes one or more of the following: an encoding of the target camera; or, Lidar data; or, an image under the target camera view before the target condition is met; or, an image under other camera views before the target condition is met; or, a BEV perception result before the target condition is met; or, a moving speed of the agent when the target condition is met; or, a moving direction of the agent when the target condition is met.
[0093] The encoding of the target camera is used to distinguish different target cameras, especially in the case where there are multiple target cameras, the agent can distinguish different target cameras through the encoding of the target camera. Simply put, the Lidar data can be data captured by other sensors (such as the laser radar, millimeter wave radar, short-range ultrasonic radar, and long-range ultrasonic radar shown in FIG. 1, etc.) on the agent other than the camera. The image under the target camera view before the target condition is met, that is, the historical image under the target camera view before the target condition is met. The image under other camera views before the target condition is met, that is, the historical image under other camera views on the agent other than the target camera before the target condition is met.
[0094] It can be understood that the more data items included in the target data, the closer the finally generated image under the target camera view is to the real image under the target camera view.
[0095] In one possible implementation, the method 200 further includes obtaining a faulty camera encoding, the faulty camera encoding being used to determine a camera that has failed.
[0096] Exemplarily, before generating the image in the target camera view based on the target data, the agent can first obtain the code of the faulty camera.
[0097] Optionally, in a possible implementation, a fault code saving and alarm module is deployed on the agent, which can be used to detect whether a sensor on the agent has failed and what type of failure has occurred (such as power failure, etc.), and further generate and save a corresponding fault code (which can include the code of the camera and the code of the type of failure), and generate alarm information, which can include the fault code.
[0098] In a possible implementation, the fault code saving and alarm module can report the alarm information to the agent, and accordingly, the agent can obtain the fault code, that is, the code of the faulty camera, so as to determine the corresponding camera as the target camera based on the code of the faulty camera.
[0099] In another possible implementation, the agent can periodically read the fault status of each camera from the fault code saving and alarm module to determine the faulty camera at the current time to obtain the code of the faulty camera. The current time can be understood as the time when the fault code saving and alarm module detects that the camera has failed.
[0100] In a possible implementation, the agent is deployed with a BEV perception module, which is used for the agent to perceive the environment in which the agent is located; and the method 200 further includes: detecting whether there is at least one camera image that has not arrived at the BEV perception module at a preset time; and in the case where there is at least one camera image that has not arrived at the BEV perception module, determining that the image of the at least one camera has image frame loss.
[0101] The preset time can be a time within each data acquisition period, for example, the preset time can be the Nth ms of each data acquisition period, 0
[0102] Exemplarily, taking an agent deployed with 11 cameras, a data acquisition period of 33 ms, and N = 2 as an example, the agent can detect how many images have arrived at the BEV perception module at the 2nd ms of each data acquisition period, that is, at the 2nd ms within 33 ms (one data acquisition period). For example, if 9 images have arrived at the BEV perception module, the agent can consider that the images that have not arrived have image frame loss. The agent can determine the cameras that have not experienced image frame loss according to the 9 images that have arrived, and further determine the cameras that have experienced image frame loss.
[0103] In a possible implementation, the method 200 further includes: respectively calculating quality scores of images captured by the cameras on the agent; and determining images with quality scores lower than a preset threshold as images that do not meet the preset image quality standard.
[0104] For example, after the cameras capture images, the agent can calculate a quality score for each image captured by each camera, and if the quality score of an image is lower than a preset threshold, the agent can determine the image as an image that does not meet the preset image quality standard.
[0105] In actual application scenarios, some images with transitional exposure, some images covered with water droplets, or some images covered with stains can be determined as images that do not meet the preset image quality standard.
[0106] Optionally, the method 200 can further include: obtaining a BEV perception result based on the generated image in the target camera view and images captured by other cameras on the agent except the target camera.
[0107] After the image in the target camera view is generated, the agent can input the generated image in the target camera view and images captured by other cameras that do not undergo image frame loss and meet the image quality standard into a single-modal BEV perception module or a single-modal BEV perception algorithm for processing and analysis to obtain a BEV perception result. For brevity, details of the BEV perception result are not described herein again, and can be referred to the related description in the technical terms section.
[0108] Optionally, the method 200 can further include: obtaining a BEV perception result based on the generated image in the target camera view, images captured by other cameras on the agent except the target camera, and Lidar data.
[0109] After the image in the target camera view is generated, the agent can input the generated image in the target camera view, images captured by other cameras that do not undergo image frame loss and meet the image quality standard, and Lidar data into a multi-modal BEV perception module or a multi-modal BEV perception algorithm for processing and analysis to obtain a BEV perception result. For brevity, details of the BEV perception result are not described herein again, and can be referred to the related description in the technical terms section.
[0110] In a possible implementation, the method 200 can further include: displaying the generated image in the target camera view.
[0111] That is, after the image in the target camera view is generated, the agent can display the generated image.
[0112] Optionally, the generated image under the target camera perspective is displayed, including: displaying the generated image under the target camera perspective and images captured by other cameras based on the 360° surround view system or the reversing image system.
[0113] For example, when the intelligent entity is a vehicle, the generated image under the target camera perspective and images captured by other cameras on the intelligent entity except the target camera can be displayed on the user interface of the 360° surround view system or the reversing image system. It can be understood that replacing the false image or the image not meeting the preset image quality standard with the generated image under the target perspective can improve the user experience.
[0114] Optionally, when displaying the image generated based on the image generation method provided in the present application on the user interface, the user can be prompted in the form of text or voice which image is generated based on the image of other cameras due to camera failure or frame loss or low image quality of the camera.
[0115] For better understanding, the image generation method provided in the present application is described again below in combination with FIGS. 3-10.
[0116] FIG. 3 is a schematic flowchart of the image generation method based on the fault code provided in an embodiment of the present application.
[0117] Steps 301-304 shown in FIG. 3 can be executed by the intelligent entity, or the method can be executed by a component (such as a chip, a chip system, etc.) configured in the intelligent entity, or can be implemented by a logic module or software capable of implementing all or part of the functions of the intelligent entity, or can be executed by an image generation device, which is not limited in the present application. The steps 301-304 are described in detail below taking the method executed by the intelligent entity as an example.
[0118] In step 301, the faulty camera is determined based on the fault code.
[0119] For example, when the camera fails, the intelligent entity can obtain the fault code, which can be obtained based on the encoding of the camera and the encoding of the fault type. Conversely, based on the fault code, the intelligent entity can also obtain the encoding of the faulty camera and determine the faulty camera based on the encoding of the faulty camera, that is, determine the target camera.
[0120] The detailed description of the intelligent entity obtaining the fault code can refer to the related description in the method 200 above, which is not described again here for brevity.
[0121] In step 302, the images captured by the non-faulty cameras in the current capture period and the Lidar data are obtained.
[0122] The current collection cycle can be understood as the data collection cycle at the moment when the camera failure is detected.
[0123] The intelligent agent can obtain the images collected by the non-failed cameras in the current collection cycle, and can also obtain the Lidar data in the current collection cycle.
[0124] In step 303, the encoding of the failed camera, the images collected by the non-failed cameras, and the Lidar data are input into the trained multi-modal diffusion model to generate the image in the current collection cycle from the perspective of the failed camera.
[0125] The intelligent agent can input the encoding of the failed camera, the images collected by the non-failed cameras in the current collection cycle, and the Lidar data in the current collection cycle into the pre-trained multi-modal diffusion model to generate the image in the current collection cycle from the perspective of the failed camera, which can be regarded as the image that the failed camera should have collected in the current collection cycle.
[0126] FIG. 4 is a schematic diagram of input data and output data of the multi-modal diffusion model provided by the present application.
[0127] As shown in FIG. 4, arrows 1 to 4 can be input data, and arrow 5 can be output data, wherein arrow 1 can represent the images collected by the non-failed cameras, arrow 2 can represent the fake images of the failed cameras, arrow 3 can represent the encoding of the failed cameras, arrow 4 can represent the Lidar data, and arrow 5 can represent the generated image from the perspective of the failed camera.
[0128] In step 304, the generated image in the current collection cycle from the perspective of the failed camera, and the images collected by the non-failed cameras in the current collection cycle and the Lidar data are input into the multi-modal BEV perception algorithm to obtain the multi-modal BEV perception result.
[0129] After generating the image in the current collection cycle from the perspective of the failed camera, the intelligent agent can process and analyze the generated image in the current collection cycle from the perspective of the failed camera, and the images collected by the non-failed cameras in the current collection cycle and the Lidar data based on the multi-modal BEV perception algorithm to obtain the multi-modal BEV perception result, and complete the perception of the surrounding environment of the intelligent agent.
[0130] For detailed description of the BEV perception result, please refer to the related description in the technical term section above. For brevity, no further description is given here.
[0131] FIG. 5 is another schematic flowchart of the image generation method based on the failure code provided by the embodiments of the present application.
[0132] The steps 501 to 504 shown in FIG. 5 can be performed by the intelligent agent, or the method can be performed by a component (such as a chip, a chip system, etc.) configured in the intelligent agent, or can also be implemented by a logic module or software capable of implementing all or part of the functions of the intelligent agent, or can also be performed by the image generation apparatus, which is not limited in the present application. The following takes an example of performing the method by the intelligent agent to explain the steps 501 to 504 in detail.
[0133] In step 501, the faulty camera is determined based on the fault code.
[0134] Exemplarily, when the camera fails, the intelligent agent can obtain a fault code, which can be obtained based on the encoding of the camera and the encoding of the fault type. Conversely, based on the fault code, the intelligent agent can also obtain the encoding of the camera that fails, and based on the encoding of the camera that fails, the camera that fails can be determined, that is, the target camera is determined.
[0135] The detailed description of the intelligent agent obtaining the fault code can refer to the related description in the method 200 above, and for the sake of brevity, it will not be repeated here.
[0136] In step 502, the image collected by the non-faulty camera in the current collection period is obtained.
[0137] The current collection period can be understood as the data collection period at the time when the camera fails is detected.
[0138] The intelligent agent can obtain the image collected by the non-faulty camera in the current collection period.
[0139] In step 503, the encoding of the faulty camera and the image collected by the non-faulty camera are input into the trained single-modal diffusion model to generate an image in the perspective of the faulty camera in the current collection period.
[0140] The intelligent agent can input the encoding of the camera that fails and the image collected by the camera that does not fail in the current collection period into the pre-trained single-modal diffusion model to generate an image in the perspective of the faulty camera in the current collection period, which can be regarded as an image that should have been collected by the camera that fails in the current collection period.
[0141] FIG. 6 is a schematic diagram of input data and output data of the single-modal diffusion model provided by the present application.
[0142] As shown in FIG. 6, arrows 1-3 can be input data, and arrow 5 can be output data, where arrow 1 can represent images captured by non-faulty cameras, arrow 2 can represent false images of the faulty camera, arrow 3 can represent encodings of the faulty camera, and arrow 5 can represent generated images in the perspective of the faulty camera.
[0143] In step 504, the generated images in the perspective of the faulty camera in the current acquisition period and the images captured by non-faulty cameras in the current acquisition period are input into the single-modal BEV perception algorithm to obtain single-modal BEV perception results.
[0144] After generating the images in the perspective of the faulty camera in the current acquisition period, the agent can process and analyze the generated images in the perspective of the faulty camera in the current acquisition period and the images captured by non-faulty cameras in the current acquisition period based on the single-modal BEV perception algorithm to obtain single-modal BEV perception results, thereby completing the perception of the surrounding environment in which the agent is located.
[0145] For detailed descriptions of the BEV perception results, reference can be made to the related descriptions in the technical terms section above, which will not be repeated here for brevity.
[0146] FIG. 7 is a schematic diagram of data interaction between modules suitable for the image generation method provided in the present application.
[0147] As shown in FIG. 7, taking the vehicle field as an example, the agent involved in the above-mentioned FIG. 2, FIG. 3 and FIG. 5 can be a vehicle, and the vehicle can be deployed with a fault code saving and alarming module, a visual sensor, an image generation module and a BEV perception module, etc. Among them, the image generation module and the BEV perception module can be deployed in an intelligent driving computing platform. In addition, the visual sensor shown in FIG. 7 can include one or more sensors as shown in FIG. 1.
[0148] Among them, the image generation module can obtain sensor data from the visual sensor, which can include images captured by normal sensors (i.e., non-faulty sensors) in the current acquisition period, Lidar data, images captured by the faulty camera before the current acquisition period (i.e., historical images), and historical images captured by non-faulty cameras. The image generation module can also obtain sensor fault states from the fault code saving and alarming module to determine the faulty sensor (i.e., the camera), and generate abnormal sensor data (i.e., generate images that should have been captured by the faulty camera in the current acquisition period) based on the sensor fault states and the sensor data.
[0149] The BEV perception module can obtain abnormal sensor data from the image generation module, and normal sensor data from the visual sensor, and process and analyze the data to obtain a BEV perception result.
[0150] FIG. 8 is a schematic flowchart of an image generation method based on image frame loss according to an embodiment of the present application.
[0151] Steps 801 to 804 shown in FIG. 8 can be executed by an intelligent entity, or the method can be executed by a component (such as a chip, a chip system, etc.) configured in the intelligent entity, or can be implemented by a logic module or software capable of implementing all or part of the functions of the intelligent entity, or can be executed by an image generation device, which is not limited in the present application. The following takes the method executed by the intelligent entity as an example to explain steps 801 to 804 in detail.
[0152] In step 801, it is determined whether image frame loss occurs based on an image arrival time threshold, and the camera that has image frame loss is determined.
[0153] For example, the intelligent entity can determine whether image frame loss occurs according to the time when the image arrives at the BEV perception module.
[0154] In one possible implementation, a preset arrival time threshold (equivalent to the preset time in the above method 200) can be used to determine whether image frame loss occurs in each collection period. Then, the camera that has frame loss, i.e., the target camera, is determined.
[0155] For example, the arrival time threshold can be N, 0 < N < collection period. For example, the collection period is 33 ms, and N = 2 or N = 5 or N = 10.
[0156] For detailed description of how the intelligent entity determines the camera that has image frame loss, please refer to the related description in the above method 200, which will not be repeated here for brevity.
[0157] In step 802, the images collected by the non-frame loss camera in the current collection period and the Lidar data are obtained.
[0158] The non-frame loss camera is a camera that does not have image frame loss, and the current collection period can be understood as the data collection period at the time when the camera is detected to have frame loss.
[0159] The intelligent entity can obtain the images collected by the non-frame loss camera in the current collection period, and can also obtain the Lidar data in the current collection period.
[0160] In step 803, the image captured by the non-frame-dropping camera and the Lidar data are input into the trained multi-modal diffusion model to generate an image in the view of the frame-dropping camera in the current acquisition period.
[0161] The frame-dropping camera is a camera that has image frame loss.
[0162] For example, the agent can input the encoding of the camera that has image frame loss, the image captured by the camera that has no image frame loss in the current acquisition period, and the Lidar data in the current acquisition period into the pre-trained multi-modal diffusion model to generate an image in the view of the frame-dropping camera in the current acquisition period, which can be regarded as the image that the camera that has image frame loss should have sent to the BEV perception module in the current acquisition period.
[0163] In step 804, the generated image in the view of the frame-dropping camera in the current acquisition period, and the image captured by the non-frame-dropping camera and the Lidar data in the current acquisition period are input into the multi-modal BEV perception algorithm to obtain a multi-modal BEV perception result.
[0164] After generating the image in the view of the frame-dropping camera in the current acquisition period, the agent can process and analyze the generated image in the view of the frame-dropping camera in the current acquisition period, and the image captured by the non-frame-dropping camera and the Lidar data in the current acquisition period based on the multi-modal BEV perception algorithm to obtain a multi-modal BEV perception result, thereby completing the perception of the surrounding environment of the agent.
[0165] For detailed description of the BEV perception result, please refer to the relevant description in the technical terms section above, which will not be repeated here for brevity.
[0166] FIG. 9 is another schematic flowchart of the image generation method based on image frame loss according to an embodiment of the present application.
[0167] Steps 901 to 904 shown in FIG. 9 can be performed by an agent, or the method can be performed by a component (such as a chip, a chip system, etc.) configured in the agent, or can also be implemented by a logic module or software capable of implementing all or part of the functions of the agent, or can also be performed by an image generation device, which is not limited in the present application. The following will take the agent as an example to explain steps 901 to 904 in detail.
[0168] In step 901, it is determined whether image frame loss occurs based on an image arrival time threshold, and the camera that has image frame loss is determined.
[0169] For detailed description, please refer to the relevant description of step 801 above, which will not be repeated here for brevity.
[0170] In step 902, the image captured by the non-frame-dropped camera in the current acquisition period is obtained.
[0171] The non-frame-dropped camera is a camera that does not have image frame dropping, and the current acquisition period can be understood as the data acquisition period at the time when the camera is detected to have frame dropping.
[0172] The agent can obtain the image captured by the non-frame-dropped camera in the current acquisition period.
[0173] In step 903, the image captured by the non-frame-dropped camera and the image captured by the frame-dropped camera are input into the trained single-modal diffusion model to generate the image in the current acquisition period from the perspective of the frame-dropped camera.
[0174] The frame-dropped camera is a camera that has image frame dropping.
[0175] For example, the agent can input the encoding of the camera that has image frame dropping and the image captured by the camera that does not have image frame dropping in the current acquisition period into the pre-trained single-modal diffusion model to generate the image in the current acquisition period from the perspective of the frame-dropped camera, which can be regarded as the image that should have been sent to the BEV perception module by the frame-dropped camera in the current acquisition period.
[0176] In step 904, the generated image in the current acquisition period from the perspective of the frame-dropped camera and the image captured by the non-frame-dropped camera in the current acquisition period are input into the single-modal BEV perception algorithm to obtain the single-modal BEV perception result.
[0177] After generating the image in the current acquisition period from the perspective of the frame-dropped camera, the agent can process and analyze the generated image in the current acquisition period from the perspective of the frame-dropped camera and the image captured by the non-frame-dropped camera in the current acquisition period based on the single-modal BEV perception algorithm to obtain the single-modal BEV perception result, thereby completing the perception of the surrounding environment of the agent.
[0178] For detailed description of the BEV perception result, please refer to the relevant description in the technical terms section above. For brevity, no further description is given here.
[0179] FIG. 10 is a schematic flowchart of an image generation method based on image quality according to an embodiment of the present application.
[0180] Steps 1001 to 1004 shown in FIG. 10 can be performed by the intelligent agent, or the method can be performed by a component (such as a chip, a chip system, etc.) configured in the intelligent agent, or can also be implemented by a logic module or software capable of implementing all or part of the functions of the intelligent agent, or can also be performed by the image generation apparatus, which is not limited in the present application. The following takes the example of performing the method by the intelligent agent to explain steps 1001 to 1004 in detail.
[0181] In step 1001, based on the images collected by all cameras on the intelligent agent, it is determined whether there is a low-quality image, and the low-quality image camera is determined.
[0182] The low-quality image can be understood as an image that does not meet the preset image quality standard.
[0183] Determining whether there is an image that does not meet the preset image quality standard can include: performing quality score calculation on the images collected by the cameras on the intelligent agent respectively; and determining the images with a quality score lower than a preset threshold as images that do not meet the preset image quality standard. For details, reference can be made to the related description in the above method 200, which will not be described here again for the sake of brevity.
[0184] In step 1002, high-quality images and Lidar data in the current collection period are obtained.
[0185] The high-quality image can be understood as an image that meets the preset image quality standard.
[0186] The intelligent agent can obtain high-quality images in the current collection period, and can also obtain Lidar data in the current collection period.
[0187] In step 1003, the encoding of the low-quality image camera, the high-quality images and the Lidar data are input into the trained multi-modal diffusion model to generate an image in the current collection period in the view of the lost-frame camera.
[0188] The intelligent agent can input the encoding of the low-quality image camera, the high-quality images collected by other cameras in the current collection period, and the Lidar data in the current collection period into the pre-trained multi-modal diffusion model to generate an image in the view of the low-quality image camera in the current collection period.
[0189] In step 1004, the generated image in the view of the low-quality image camera in the current collection period, and the high-quality images and the Lidar data collected in the current collection period are input into the multi-modal BEV perception algorithm to obtain multi-modal BEV perception results.
[0190] After generating the low-quality image from the camera's perspective within the current acquisition period, the agent can process and analyze the generated low-quality image from the camera's perspective within the current acquisition period, as well as the high-quality images and LiDAR data acquired within the current acquisition period, based on the multimodal BEV perception algorithm, to obtain the multimodal BEV perception result and complete the perception of the surrounding environment in which the agent is located.
[0191] For a detailed description of the BEV perception results, please refer to the relevant description in the technical terminology section above. For the sake of brevity, it will not be repeated here.
[0192] Figure 11 is another schematic flowchart of an image generation method based on image quality provided in an embodiment of this application.
[0193] Steps 1101 to 1104 shown in Figure 11 can be executed by an intelligent agent, or the method can be executed by a component (such as a chip, chip system, etc.) configured in the intelligent agent, or it can be implemented by a logic module or software capable of realizing all or part of the functions of the intelligent agent, or it can be executed by an image generation device. This application does not limit this. The following describes steps 1101 to 1104 in detail using the execution of the method by an intelligent agent as an example.
[0194] In step 1101, based on the images captured by all cameras on the agent, it is determined whether there are low-quality images and the low-quality image camera is identified.
[0195] Low-quality images can be understood as images that do not meet the preset image quality standards.
[0196] For a detailed description, please refer to the relevant explanation in step 1001 above. For the sake of brevity, it will not be repeated here.
[0197] In step 1102, high-quality images within the current acquisition period are acquired.
[0198] High-quality images can be understood as images that meet preset image quality standards.
[0199] The agent can acquire high-quality images within the current acquisition period.
[0200] In step 1103, the low-quality image camera code and the high-quality image are input into the trained single-modal diffusion model to generate the image from the perspective of the dropped frame camera within the current acquisition period.
[0201] The agent can input the encoding of the low-quality image camera and the high-quality images captured by other cameras in the current acquisition period into a pre-trained single-modal diffusion model to generate an image from the perspective of the low-quality image camera in the current acquisition period.
[0202] In step 1104, the generated image under the low-quality image camera view in the current acquisition period and the high-quality image acquired in the current acquisition period are input into the single-modal BEV perception algorithm to obtain a single-modal BEV perception result.
[0203] After the image under the low-quality image camera view in the current acquisition period is generated, the agent can process and analyze the generated image under the low-quality image camera view in the current acquisition period and the high-quality image acquired in the current acquisition period based on the single-modal BEV perception algorithm to obtain a single-modal BEV perception result, and complete the perception of the surrounding environment in which the agent is located.
[0204] For detailed description of the BEV perception result, refer to the related description in the technical terms section above, which will not be repeated here for brevity.
[0205] It should be noted that the arrows shown in FIGS. 3, 5, 7 to 11 above can represent data flow direction, and do not represent and limit the execution order between steps.
[0206] It can be understood that the single-modal diffusion model (or algorithm) or multi-modal diffusion model (or algorithm) involved in the embodiments of the present application can be based on the assumption that a certain camera or certain cameras meet the target conditions when training. The real image collected by the assumed target camera is compared with the image generated by the single-modal diffusion model (or algorithm) or multi-modal diffusion model (or algorithm), and the difference is calculated. The network parameters of the diffusion model are iteratively optimized multiple times using gradient backpropagation difference value, so that the generated image is as close as possible to the real image collected by the assumed target camera.
[0207] Exemplarily, the diffusion model suitable for the embodiments of the present application can include but is not limited to: denoising diffusion probabilistic models (DDPM), improved denoising diffusion probabilistic models (IDDPM), score-based generative models (SGMs), variational diffusion models (VDMs), continuous-time diffusion models (CTDMs), or latent diffusion models (LDMs), etc.
[0208] It can be understood that the BEV perception of single-modal multi-view fusion or multi-modal multi-view fusion needs to ensure the consistency of the spatial data collection of multiple sensors, that is, the original data collection time of each view sensor of each mode needs to be almost the same (that is, less than or equal to a preset time error threshold, for example, 1 ms), so as to ensure that the detection accuracy of the BEV perception algorithm is not affected by the "spatial inconsistency" and is reduced.
[0209] Based on the above technical solution, first, the image is generated based on the target condition, which can realize the BEV global blind area-free perception detection without camera redundancy design backup, improve the BEV perception algorithm accuracy and robustness, that is, improve the accuracy of the intelligent agent's environmental perception. More specifically, the generated image under the fault camera view replaces the fake image, which not only improves the BEV perception algorithm accuracy and robustness, but also avoids the problem of exiting the intelligent driving state caused by camera failure, thereby ensuring the robustness of the intelligent driving system; the generated image under the frame loss camera view replaces the fake image, which can avoid the problem of BEV perception detection blind area or low detection accuracy caused by image frame loss; the generated image under the low-quality image camera view replaces the low-quality image, which can avoid the problem of BEV perception detection blind area or low detection accuracy caused by image frame loss or low image quality. Furthermore, displaying the generated image on the user interface to replace the fake image can also improve the user experience.
[0210] Exemplarily, the method provided by the embodiments of the present application can be sunk and hardened to be deployed to the intelligent driving computing platform, and related materials can be provided to related algorithm manufacturers for reference, or can be recorded in the development document of the intelligent driving computing platform.
[0211] The above describes in detail the image generation method provided by the embodiments of the present application in combination with the drawings. The following describes in detail the image generation device provided by the embodiments of the present application in combination with the drawings.
[0212] It should be understood that the image generation device shown in FIG. 12 and FIG. 13 can be used to realize the functions of the intelligent agent in the above method embodiments, and thus the beneficial effects possessed by the above method embodiments can also be realized. In the embodiments of the present application, the image generation device can realize the steps performed by the intelligent agent in any one of the method embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11, and the image generation device can also be a component (such as a chip, a chip system, a processor, etc.) configured in the intelligent agent, and can also be a logic module or software capable of realizing part or all of the functions of the intelligent agent.
[0213] FIG. 12 is a schematic block diagram of an image generation device provided by an embodiment of the present application.
[0214] As shown in FIG. 12, the image generation apparatus 1200 includes a determination module 1210 and a generation module 1220. The image generation apparatus 1200 can be used to implement the functions of the agent in any of the method embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11.
[0215] In an example, when the image generation apparatus 1200 is used to implement the functions of the agent in the method embodiment shown in FIG. 2, the determination module 1210 can be used to determine a target camera on the agent that meets a target condition, the agent including a plurality of cameras; and the generation module 1220 can be used to generate an image under a target camera view based on target data, the target data including images captured by other cameras on the agent other than the target camera; wherein the target condition includes one or more of: a camera failure; or, a camera image frame loss; or, an image captured by a camera not meeting a preset image quality standard.
[0216] Optionally, the determination module 1210 can also be used to obtain a failure camera code, the failure camera code being used to determine a camera that has failed.
[0217] Optionally, the agent is deployed with a BEV perception module, the BEV perception module being used by the agent to perceive an environment in which the agent is located; and the determination module 1210 can also be used to: detect, at a preset time, whether images of at least one camera have not arrived at the BEV perception module; and in a case where images of at least one camera have not arrived at the BEV perception module, determine that the images of the at least one camera have experienced image frame loss.
[0218] Optionally, the determination module 1210 can also be used to: calculate a quality score for each image captured by a camera on the agent; and determine an image with a quality score lower than a preset threshold as an image that does not meet a preset image quality standard.
[0219] Optionally, the target data further includes one or more of: a code of the target camera; or, Lidar data; or, an image under the target camera view before the target condition is met; or, an image under another camera view before the target condition is met; or, a BEV perception result before the target condition is met; or, a moving speed of the agent when the target condition is met; or, a moving direction of the agent when the target condition is met.
[0220] Optionally, the determination module 1210 can also be used to: obtain a BEV perception result based on the generated image under the target camera view and the images captured by other cameras on the agent other than the target camera.
[0221] Optionally, the determining module 1210 can also be configured to obtain the BEV perception result based on the generated image in the target camera view, images captured by other cameras on the agent except the target camera, and Lidar data.
[0222] Optionally, the determining module 1210 or the generating module 1220 can also be configured to display the generated image in the target camera view.
[0223] Optionally, the determining module 1210 or the generating module 1220 can be specifically configured to display the generated image in the target camera view and the images captured by other cameras based on a 360° surround view system or a reversing image system.
[0224] For more details of the above modules, refer to the descriptions of the method embodiments shown in FIG. 2.
[0225] It should be understood that the division of modules in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, another division manner can be used. In addition, each functional module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0226] FIG. 13 is another schematic block diagram of an image generation apparatus provided in an embodiment of the present application.
[0227] The image generation apparatus 1300 can be a chip system, or can be a device configured with a chip system to implement the method described in the above method embodiments. In the embodiments of the present application, the chip system can be composed of a chip, or can include a chip and other discrete devices.
[0228] As shown in FIG. 13, the image generation apparatus 1300 can include a processor 1310, which can be configured to execute computer programs or instructions in a memory to implement the steps performed by the agent in any of the method embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11.
[0229] Optionally, the image generation apparatus 1300 further includes a communication interface 1320. The communication interface 1320 can be configured to communicate with other devices through a transmission medium, thereby enabling the image generation apparatus 1300 to communicate with other devices. The communication interface 1320 can be, for example, a transceiver, an interface, a bus, a circuit, or a device that enables transceiving. The processor 1310 can input and output data through the communication interface 1320, and can be configured to implement the method in any of the embodiments of FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11. Specifically, the image generation apparatus 1300 can be configured to implement the functions of the agent in the above method embodiments.
[0230] Optionally, the image generation apparatus 1300 further includes at least one memory 1330 configured to store program instructions and / or data. The memory 1330 is coupled to the processor 1310. The coupling between the processor 1310 and the memory 1330 in the embodiments of the present application is indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, for information interaction between devices, units or modules. The processor 1310 can operate in conjunction with the memory 1330. The processor 1310 can execute program instructions stored in the memory 1330.
[0231] In the present application, the memory 1330 can be integrated into the processor 1310, and the processor 1310 and the memory 1330 can also be separately provided, which is not limited in the present application.
[0232] It should be understood that the coupling in the embodiments of the present application is indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms, for information interaction between devices, units or modules. The processor 1310 can operate in conjunction with the memory 1330. The specific connection medium between the processor 1310, the communication interface 1320 and the memory 1330 is not limited in the embodiments of the present application. In FIG. 13, the processor 1310, the communication interface 1320 and the memory 1330 are connected through a bus 1340. The bus 1340 is represented by a thick line in FIG. 13, and the connection mode between other components is only schematically illustrated and is not limited. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, only one thick line is used to represent the bus in FIG. 13, but it does not mean that there is only one bus or only one type of bus.
[0233] The embodiment of the present application further provides an intelligent agent, and the vehicle comprises the image generation device.
[0234] The embodiment of the present application further provides an intelligent agent, and the vehicle comprises a module which can implement the method in any of the embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11.
[0235] Optionally, in the embodiment of the present application, the intelligent agent can comprise a vehicle, a robot, a drone or a smart phone, etc.
[0236] The present application further provides a computer program product, which comprises a computer program (also referred to as code or instructions), which, when executed, can implement the steps performed by the intelligent agent in the method in any of the embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11.
[0237] The present application further provides a computer readable storage medium, which stores a computer program (also referred to as code or instructions). When the computer program is executed, it can implement the steps performed by the intelligent agent in the method in any of the embodiments shown in FIG. 2, FIG. 3, FIG. 5, FIG. 7 to FIG. 11.
[0238] It should be understood that the processor in the embodiment of the present application can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the above method embodiments can be completed by an integrated logic circuit of hardware in the processor or by instructions in the form of software. The processor mentioned above can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block diagram disclosed in the embodiment of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can be any conventional processor. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the memory is read by the processor, and the hardware thereof is combined to complete the steps of the above method.
[0239] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0240] The terms "unit", "module" and the like used in the specification can be used to represent a computer-related entity, hardware, firmware, a combination of hardware and software, software, or software in execution. The units and modules in the embodiments of the present application have the same meaning and can be used interchangeably.
[0241] Those of skill would further appreciate that the various illustrative logical blocks, modules, circuits, and steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. The choice of hardware or software, or combinations of both, would be dependent on the specific application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application. In several embodiments provided in the present application, it will be apparent that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the described device embodiments are merely illustrative, and the division into units is merely a logical function division, and actual implementation can have another division, for example, multiple units or components can be combined or integrated into another system, or some features can be omitted or not implemented. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0242] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0243] In addition, the functional units in each of the embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0244] In the above embodiments, the functions of the various functional units can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, the software can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, digital video disc (DVD)), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0245] The functions, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that make a contribution to the technology or parts of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk and various media that can store program codes.
[0246] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An image generation method characterized by, The method comprises: determining a target camera on the agent that meets a target condition, the agent comprising a plurality of cameras; generating an image under a view of the target camera based on target data, the target data comprising images captured by other cameras on the agent other than the target camera; wherein the target condition comprises one or more of: a camera is malfunctioning; or, a camera is experiencing image frame loss; or, an image captured by a camera does not meet a preset image quality standard.
2. The method of claim 1, wherein, The method further comprises: obtaining a malfunctioning camera code, the malfunctioning camera code being used to determine a camera that is malfunctioning.
3. The method of claim 1 or 2, wherein, The agent is deployed with a bird's eye view (BEV) perception module, the BEV perception module being used for the agent to perceive an environment in which the agent is located; and the method further comprises: detecting, at a preset time, whether images of at least one camera are not reaching the BEV perception module on the agent; in a case where images of at least one camera are not reaching the BEV perception module, determining that the images of the at least one camera are experiencing image frame loss.
4. The method of any one of claims 1 to 3, wherein, The method further comprises: performing quality score calculation on images captured by the cameras on the agent respectively; determining, as images that do not meet a preset image quality standard, images whose quality scores are lower than a preset threshold.
5. The method of any one of claims 1 to 4, wherein, The target data further comprises one or more of: a code of the target camera; or, light detection and ranging (Lidar) data; or, images under a view of the target camera before the target condition is met; or, images under a view of the other cameras before the target condition is met; or, BEV perception results before the target condition is met; or, a moving speed of the agent when the target condition is met; or, a moving direction of the agent when the target condition is met.
6. The method of any one of claims 1 to 5, wherein, The method further comprises: obtaining BEV perception results based on the generated images under the view of the target camera and the images captured by the other cameras on the agent other than the target camera.
7. The method of any one of claims 1 to 5, wherein, The method further comprises: obtaining BEV perception results based on the generated images under the view of the target camera, the images captured by the other cameras on the agent other than the target camera, and Lidar data.
8. The method of any one of claims 1 to 7, wherein, The method further comprises: displaying the generated images under the view of the target camera.
9. The method of claim 8, wherein, The displaying of the generated images under the view of the target camera comprises: displaying the generated images under the view of the target camera and the images captured by the other cameras based on a 360° surround view system or a reversing image system.
10. An image generation apparatus characterized by comprising: The apparatus comprises modules for performing the method of any one of claims 1 to 9.
11. An image generation apparatus characterized by comprising: comprises a processor and a memory, wherein the memory is configured to store a program; the processor is configured to invoke the program to cause the apparatus to perform the method of any one of claims 1 to 9.
12. An agent, characterized in that The agent is configured to perform the method of any one of claims 1 to 9.
13. An agent, characterized in that The agent comprises at least two cameras and the image generation apparatus of claim 10 or 11.
14. The agent of claim 12 or 13, wherein, The agent comprises a vehicle, a robot, a drone, or a smartphone.
15. A computer-readable storage medium having stored thereon a program, the program comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 14. The program, when executed, causes the method of any one of claims 1 to 9 to be performed.
16. A computer program product, characterised in that, The computer program, when run, causes the method of any one of claims 1 to 9 to be performed.