Panoramic image generation method, device, medium, product and vehicle

By dividing the background image into multiple sub-background images and mapping them to different bowl layers, combined with the precise mapping of the target object, the problem that panoramic images cannot reflect the depth differences of the scene in the existing technology is solved, and the three-dimensionality and realism of the panoramic image are improved.

WO2026011934A1PCT designated stage Publication Date: 2026-01-15BYD CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/094387
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2025-05-12
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing in-vehicle panoramic imaging technology cannot effectively reflect the distance differences between the vehicle and objects around it, resulting in a poor visual experience.

Method used

The background image is divided into multiple sub-background images and mapped onto different bowl layers of a pre-built bowl-shaped model. At the same time, the target object is mapped onto the corresponding bowl layer, and the image mapping is performed using the different distances of the bowl layers from the center point of the vehicle.

Benefits of technology

It improves the visualization of the three-dimensional background and target objects in panoramic images, enhances the sense of spatial depth, avoids vertical stretching and distortion of ground objects, and improves the realism of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025094387_15012026_PF_FP_ABST
    Figure CN2025094387_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A panoramic image generation method, a device, a medium, a product and a vehicle, relating to the technical field of computers. The method comprises: acquiring an object image of at least one target object and a background image from a collected environment image of a vehicle; dividing the background image into a plurality of sub-background images; determining a first target bowl layer corresponding to each sub-background image from a plurality of bowl layers of a pre-constructed bowl-shaped model; determining a second target bowl layer corresponding to the at least one target object from the plurality of bowl layers; and respectively mapping each sub-background image to the first target bowl layer corresponding to the sub-background image, and mapping the at least one object image to the corresponding second target bowl layer, so as to obtain a target panoramic image of an environment where the vehicle is located.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, equipment, media, products, and vehicles for generating panoramic images

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to Chinese Patent Application No. 202410912750.X, filed on July 9, 2024, entitled "Method, apparatus, medium, product and vehicle for generating panoramic images", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, medium, product, and vehicle for generating panoramic images. Background Technology

[0004] With the rapid development of computer technology, vehicles are often equipped with panoramic surround view functions to further enhance the driver's driving experience. Most existing in-vehicle panoramic images are based on a bowl-shaped model, and their 3D panoramic display effect is a texture or mapping effect based on the stitching of 2D panoramic images taken from multiple perspectives.

[0005] Currently, the common approach is to map pedestrians or obstacles from a 2D panoramic image onto the inner bowl layer of a bowl-shaped model, while mapping the remaining background image onto the outermost bowl layer. However, the panoramic image obtained by this method fails to accurately represent the depth differences between different objects in the scene, and it also fails to accurately represent the distance between objects around the vehicle and the vehicle itself, resulting in a poor visual experience for the user. Summary of the Invention

[0006] This disclosure aims to at least partially address one of the technical problems in the related art.

[0007] Therefore, the first objective of this disclosure is to provide a method, apparatus, medium, product, and vehicle for generating panoramic images.

[0008] According to a first aspect of the present disclosure, a method for generating a panoramic image is provided, the method comprising:

[0009] Obtain at least one object image and background image of a target object from environmental images captured by the vehicle;

[0010] The background image is divided into multiple sub-background images;

[0011] The first target bowl layer corresponding to each sub-background image is determined from multiple bowl layers of a pre-constructed bowl-shaped model; the bowl-shaped model is a model containing multiple bowl layers built with the center of the bowl bottom as the center point of the bowl bottom, and the distance from the bottom of different bowl layers to the center point of the bowl bottom is different; different sub-background images correspond to different first target bowl layers;

[0012] Determine the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers;

[0013] Each sub-background image is mapped onto the first target bowl layer corresponding to the sub-background image, and the at least one object image is mapped onto the corresponding second target bowl layer to obtain a target panoramic image of the environment in which the vehicle is located.

[0014] As one embodiment, the background image includes a first road surface image, and dividing the background image into multiple sub-background images includes:

[0015] Determine a first distance between the road surface in the first road surface image and the vehicle;

[0016] Based on the first distance, the background image is divided into the plurality of sub-background images.

[0017] As one embodiment, dividing the background image into the plurality of sub-background images based on the first distance includes:

[0018] Based on the first distance and multiple preset distance ranges, the road surface in the first road surface image is divided into multiple sub-road surfaces;

[0019] The background image is divided into multiple sub-background images based on the multiple sub-road surfaces, and each sub-road surface corresponds to one of the sub-background images.

[0020] As one embodiment, the background image includes a first road surface image, and dividing the background image into multiple sub-background images includes:

[0021] Identify road surface elements in the first road surface image;

[0022] Based on the position of the road surface elements in the first road surface image, the background image is divided into the plurality of sub-background images.

[0023] As one embodiment, determining the first target bowl layer corresponding to each of the sub-background images from multiple bowl layers of a pre-constructed bowl-shaped model includes:

[0024] For each sub-background image, the first target bowl layer corresponding to the sub-background image is determined based on the second distance between the road surface and the vehicle in the sub-background image.

[0025] As one embodiment, determining the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers includes:

[0026] Determine a third distance between the at least one target object and the vehicle;

[0027] The second target bowl layer corresponding to the at least one target object is determined based on the third distance and the plurality of second distances.

[0028] As one embodiment, the method further includes:

[0029] The background image is input into a pre-trained image generation model to obtain the target background image output by the image generation model;

[0030] The target background image includes the background image and the image corresponding to the background area in the environment image that is occluded by the target object;

[0031] The step of dividing the background image into multiple sub-background images includes:

[0032] The target background image is divided into multiple sub-background images.

[0033] As one embodiment, obtaining an object image and a background image of at least one target object from the environmental images captured by the vehicle includes:

[0034] The environmental image is input into a pre-trained target detection model to obtain the object image and the background image corresponding to the at least one target object output by the target detection model.

[0035] As one embodiment, the background image further includes a first other background image besides the first road surface image, and the sub-background image includes a second road surface image and a second other background image besides the second road surface image. Mapping each sub-background image to the corresponding first target bowl layer includes:

[0036] Map the second road surface image onto the bottom of the bowl corresponding to the first target bowl layer;

[0037] The second other background image is mapped onto the rim of the bowl corresponding to the first target bowl layer.

[0038] As one embodiment, mapping the at least one object image onto the corresponding second target bowl layer includes:

[0039] Obtain the relative position information of the at least one target object in the sub-background image;

[0040] Based on the relative position information, determine the target mapping position of the at least one target object in the second target bowl layer;

[0041] Map the at least one object image to the target mapping position of the corresponding second target bowl layer.

[0042] As one embodiment, the method further includes:

[0043] For each target object, the adjustment scale factor corresponding to the target object is obtained according to the second target bowl layer;

[0044] Adjust the object image corresponding to the target object according to the adjustment scale factor to obtain the target object image;

[0045] The step of mapping the at least one object image to the target mapping position of the corresponding second target bowl layer includes:

[0046] At least one target object image is mapped to the target mapping position of the corresponding second target bowl layer.

[0047] According to a second aspect of the present disclosure, an electronic device is provided, comprising: a memory storing a computer program thereon; and a processor for executing the computer program in the memory to implement the steps of the panoramic image generation method provided in the first aspect of the present disclosure.

[0048] According to a third aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the panoramic image generation method provided in the first aspect of the present disclosure.

[0049] According to a fourth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the panoramic image generation method provided in the first aspect of the present disclosure.

[0050] According to a fifth aspect of the present disclosure, a vehicle is provided, including the electronic equipment provided in the second aspect of the present disclosure.

[0051] The above technical solution involves obtaining an object image and a background image of at least one target object from environmental images captured by the vehicle; dividing the background image into multiple sub-background images; determining a first target bowl layer corresponding to each sub-background image from multiple bowl layers of a pre-constructed bowl-shaped model; the bowl-shaped model is a model containing multiple bowl layers built with the center of the vehicle as the center point of the bowl bottom, with different distances from the bottom of different bowl layers to the center point of the bowl bottom; different first target bowl layers corresponding to different sub-background images; determining a second target bowl layer corresponding to the at least one target object from the multiple bowl layers; mapping each sub-background image to the first target bowl layer corresponding to the sub-background image, and mapping the at least one object image to the corresponding second target bowl layer, to obtain a target panoramic image of the vehicle's environment. By dividing the background image into multiple sub-background images and mapping different sub-background images to different first target bowl layers, the depth differences of different objects in the background can be better reflected. Simultaneously, by mapping the sub-background images and target objects separately, the visualization effect and spatial three-dimensionality of the three-dimensional background and target objects can be improved. Furthermore, the different distances from the bottom of the bowl-shaped model to the center point of the bowl bottom can avoid the problem of the ground scenery at the bottom of different bowl layers being stretched and deformed longitudinally, making the final panoramic image of the target more realistic.

[0052] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0053] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:

[0054] Figure 1 is a flowchart illustrating a method for generating a panoramic image according to an exemplary embodiment;

[0055] Figure 2 is a side view schematic diagram of a bowl-shaped model according to an exemplary embodiment;

[0056] Figure 3 is a flowchart illustrating another method for generating panoramic images according to an exemplary embodiment;

[0057] Figure 4 is a flowchart illustrating another method for generating panoramic images according to an exemplary embodiment;

[0058] Figure 5 is a flowchart illustrating another method for generating panoramic images according to an exemplary embodiment;

[0059] Figure 6 is a flowchart illustrating another method for generating panoramic images according to an exemplary embodiment;

[0060] Figure 7 is a top view of a bowl-shaped model according to an exemplary embodiment;

[0061] Figure 8 is a block diagram illustrating a panoramic image generation apparatus according to an exemplary embodiment;

[0062] Figure 9 is a block diagram of another panoramic image generation apparatus according to an exemplary embodiment;

[0063] Figure 10 is a block diagram of another panoramic image generation apparatus according to an exemplary embodiment;

[0064] Figure 11 is a block diagram illustrating an electronic device according to an exemplary embodiment;

[0065] Figure 12 is a block diagram illustrating a vehicle according to an exemplary embodiment. Detailed Implementation

[0066] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.

[0067] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.

[0068] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily construed as referring to a specific order or sequence. Furthermore, in the description with reference to the accompanying drawings, the same reference numerals in different drawings denote the same elements.

[0069] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0070] Although operations or steps are described in a specific order in the accompanying drawings in the embodiments of this disclosure, it should not be construed as requiring these operations or steps to be performed in the specific order or serial order shown, or requiring all of the shown operations or steps to be performed to obtain the desired result. In the embodiments of this disclosure, these operations or steps may be performed serially; they may be performed in parallel; or a portion of these operations or steps may be performed.

[0071] In the description of this disclosure, unless otherwise stated, "multiple" means two or more, and other quantifiers are similar; "at least one," "one or more," or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one 'a' can represent any number of 'a's; as another example, one or more of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple; "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " indicates that the preceding and following related objects are in an "or" relationship.

[0072] Before introducing the panoramic image generation method, device, medium, product, and vehicle provided in this disclosure, the application scenarios involved in the various embodiments of this disclosure are first described. This disclosure can be applied to scenarios where vehicles provide panoramic images. In such scenarios, objects such as pedestrians or obstacles in the 2D panoramic image are typically mapped onto the inner bowl layer of a bowl-shaped model, while the remaining background image is mapped onto the outermost bowl layer. However, the panoramic image obtained by this method cannot well reflect the depth differences of different objects in the scene, nor can it well reflect the distance between objects around the vehicle and the vehicle, resulting in a poor visual experience for the user.

[0073] In some related technologies, target obstacles in an environmental image are detected, along with their width, length, and distance from the vehicle. The obstacles are then mapped to different layers of the environmental image. Specifically, the image of the target obstacle is mapped to at least one inner layer, while other parts of the environmental image are mapped to the outermost layer. However, the inventors have found that in this method, both the outermost and inner layers have bottom circular surfaces of the same radius. After mapping, objects at the bottom of layers farther from the vehicle experience significant distortion, thus affecting the visualization effect of the final panoramic image.

[0074] To address the aforementioned technical problems, this invention provides a method, device, medium, product, and vehicle for generating panoramic images. By dividing a background image into multiple sub-background images and mapping different sub-background images onto different first target bowl layers, the depth differences of different objects in the background can be better represented. Simultaneously, by mapping the sub-background images and target objects separately, the visualization effect and spatial depth of the three-dimensional background and target objects can be improved. Furthermore, the different distances from the bottom of different bowl layers to the center point of the bowl bottom in the bowl-shaped model can avoid the problem of longitudinal stretching and deformation of ground objects at the bottom of different bowl layers, making the final target panoramic image more realistic.

[0075] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0076] Figure 1 is a flowchart illustrating a panoramic image generation method according to an exemplary embodiment. This method can be applied to a vehicle equipped with multiple image acquisition devices, each used to acquire environmental images from different perspectives. As shown in Figure 1, the method may include the following steps:

[0077] In step S101, an object image and a background image of at least one target object are obtained from the environmental images collected by the vehicle.

[0078] The environmental image can be an image from different perspectives captured by multiple image acquisition devices, such as front view, rear view, left view, and right view. For example, after the vehicle is powered on, the image acquisition devices (such as panoramic surround view cameras, binocular cameras, etc.) for the front view, rear view, left view, and right view can be activated separately to acquire four-way views of the environment surrounding the vehicle.

[0079] In some embodiments, depth information corresponding to each pixel in an environmental image can be obtained through an image acquisition device, which is convenient for subsequent determination of the distance between each location in the environmental image and the vehicle.

[0080] Furthermore, based on the acquired environmental image, by identifying the target object and background in the environmental image, at least one object image and a background image corresponding to the target object can be obtained. The target object can be, for example, a pedestrian, vehicle, or other obstacle. The background image is the remaining portion of the environmental image excluding the object image; that is, the background image does not include at least one target object from the environmental image. The background image may include a first road surface image corresponding to the road surface and a first other background image excluding the first road surface image. For ease of subsequent analysis, the identified target object and road surface can also be labeled.

[0081] For example, the environmental image can be input into a pre-trained object detection model to obtain the object image and the background image corresponding to the at least one target object output by the object detection model. The object detection model can be any detection model with target object recognition capabilities in related technologies, and this disclosure does not specifically limit it. Furthermore, the object detection model can also identify a first road surface image corresponding to the area where the road surface is located in the background image, thereby obtaining the first road surface image and a first other background image other than the first road surface image in the background image.

[0082] In step S102, the background image is divided into multiple sub-background images.

[0083] In this step, to display the background image in a more three-dimensional way, it can be divided into multiple sub-background images. Specifically, the background image can be divided into multiple sub-background images based on road surface information in the background image. This road surface information includes the first distance between the road surface and vehicles in the first road surface image, or road surface elements (e.g., curbs, green belts, etc.). In other words, dividing the background image into multiple sub-background images based on road surface information can be understood as dividing the background image according to the road surface information in the background image to obtain multiple sub-background images.

[0084] For example, in one approach, the road surface can be divided into multiple regions based on a first distance between the road surface and the vehicle in the first road surface image, thereby obtaining a sub-background image corresponding to each region. The first distance can be obtained based on the depth information corresponding to each pixel in the environmental image acquired in step 101. The distances between the road surface and the vehicle contained in different regions are different, and the width of each region can be the same or different; this disclosure does not specifically limit this. In another approach, the road surface can also be divided into multiple regions based on road surface elements, thereby obtaining a sub-background image corresponding to each region. For example, the background image can be divided into multiple sub-background images based on the position of the road surface element in the first road surface image.

[0085] In step S103, the first target bowl layer corresponding to each sub-background image is determined from multiple bowl layers of the pre-constructed bowl-shaped model.

[0086] The bowl-shaped model is a multi-layered model built with the center of the vehicle as the center point of the bowl's bottom. The distance from the bottom of each layer to the center point of the bowl's bottom varies; different sub-background images correspond to different first target bowl layers. As shown in Figure 2, point O is the center point of the bowl's bottom, and L1, L2, L3, and L4 represent different bowl layers. The distance from the bottom of each bowl layer to the center point of the bowl's bottom is different, meaning that different bowl layers have bottom circular surfaces with different radii. This solves the problem in related technologies where the outer and inner bowl layers have bottom circular surfaces of the same radius, resulting in distortion of the scenery at the bottom of bowl layers farther from the vehicle after mapping, thus improving the visualization effect of the subsequently generated panoramic target image.

[0087] Furthermore, in this embodiment, the curvature of the rim of each bowl layer can be the same or different. For example, the central bowl layer can be cylindrical, while the rims of the other bowl layers can be arc-shaped. This disclosure does not impose any specific limitations on this. By setting the curvature of the rim of each bowl layer to be different, the depth differences of different scenes can be better reflected, thereby enabling a more three-dimensional presentation of different scenes.

[0088] In this step, when dividing the sub-background images according to step S102, a second distance between the road surface and the vehicle in each sub-background image can be obtained. This second distance can be a range, that is, the distance range between the road surface closest to the vehicle and the road surface farthest from the vehicle in the sub-background image, for example, 3 to 5 meters. This second distance can be obtained based on the depth information corresponding to each pixel in the environmental image obtained in step 101. In this step, the sub-background images can be allocated sequentially according to the second distance between the road surface and the vehicle in each sub-background image, from smallest to largest, to obtain the first target bowl layer corresponding to each sub-background image. For example, the sub-background image with the largest second distance can be allocated to the outermost bowl layer of the bowl-shaped model, that is, the outermost bowl layer is taken as the first target bowl layer corresponding to the sub-background image, and the allocation is carried out sequentially from the outermost bowl layer to the inner bowl layers according to the size of the second distance, to obtain the first target bowl layer corresponding to each sub-background image.

[0089] In step S104, the second target bowl layer corresponding to the at least one target object is determined from the plurality of bowl layers.

[0090] In this step, different target objects can be assigned to different bowl layers based on a third distance between at least one target object and the vehicle, thus obtaining a second target bowl layer corresponding to at least one target object. This third distance can also be obtained based on the depth information corresponding to each pixel in the environmental image acquired in step 101. The closer a target object is to the vehicle, the closer its corresponding second target bowl layer is to the model's center line (i.e., the straight line passing through the center point of the bowl bottom and perpendicular to the bowl bottom). This effectively maps target objects with different depths of field onto different second target bowl layers, improving the three-dimensionality and realism of the target objects in the panoramic image.

[0091] In some embodiments, to improve the accuracy of target object mapping, distance information corresponding to target objects within a preset range (which can be the range of distances detectable by the ultrasonic radar detection device) around the vehicle can be obtained through an ultrasonic radar detection device installed on the vehicle. Specifically, the ultrasonic radar detection device can be activated simultaneously after the vehicle is powered on. Then, based on the third distance and the distance information corresponding to the target object, different target objects can be assigned to different target layers to obtain at least one second target target target layer. In other words, when assigning target layers, the third distance detected by the ultrasonic radar detection device can also be considered to improve the accuracy of the target layer assignment.

[0092] In step S105, each sub-background image is mapped onto the first target bowl layer corresponding to the sub-background image, and the at least one object image is mapped onto the corresponding second target bowl layer to obtain a target panoramic image of the environment in which the vehicle is located.

[0093] In this step, when mapping the sub-background image and the object image, you can first map the sub-background image and then overlay the object image onto the sub-background image.

[0094] In some embodiments, the sub-background image may include a second road surface image and a second other background image other than the second road surface image. For each sub-background image, the second road surface image can be mapped to the bottom of the first target bowl layer, and the second other background image can be mapped to the rim of the first target bowl layer. This allows for more targeted mapping of different types of backgrounds to different positions within the model bowl layer, more closely resembling the visual effect in a real-world scene.

[0095] In other embodiments, the relative position information of the at least one target object in the sub-background image can be obtained. That is, the position of the at least one target object relative to the sub-background image is determined. Then, based on the relative position information, the target mapping position of the at least one target object in the second target bowl layer can be determined, and the image of the at least one object can be mapped to the corresponding target mapping position in the second target bowl layer.

[0096] Furthermore, considering the visual perception of objects appearing larger when closer and smaller when farther away in real-world scenarios, to improve the realism of the target object mapping, in one implementation, an adjustment scale factor corresponding to each target object can be obtained based on the second target bowl layer. This adjustment scale factor is used to characterize the parameters of the adjusted object image size and shape. For example, preset adjustment scale factors corresponding to different bowl layers can be pre-set, and the adjustment scale factor can be determined from multiple preset adjustment scale factors based on the second target bowl layer. Then, the object image corresponding to the target object is adjusted according to the adjustment scale factor to obtain the target object image. For example, the object image can be scaled, deformed, etc., based on the adjustment scale factor. Afterwards, at least one adjusted target object image is mapped to the corresponding target mapping position in the second target bowl layer.

[0097] In another implementation, the distance information between the target object and the vehicle (detected by an ultrasonic radar detection device) can be further considered. Then, based on the second target bowl layer and the distance information between the target object and the vehicle, the adjustment scale factor corresponding to the target object is obtained. For example, preset adjustment scale factors corresponding to different bowl layers and different distances can be pre-set, and the adjustment scale factor can be determined from multiple preset adjustment scale factors based on the second target bowl layer and distance information. Then, the object image corresponding to the target object is adjusted according to the adjustment scale factor to obtain the target object image. For example, the object image can be scaled, deformed, etc., according to the adjustment scale factor. Afterwards, at least one adjusted target object image is mapped to the corresponding target mapping position of the second target bowl layer.

[0098] Furthermore, considering that users pay more attention to target objects while driving, in some embodiments, to make the target object's location more clearly visible to the user, the target object image can be mapped onto the target mapping position of the second target bowl layer in a form that is perpendicular or nearly perpendicular to the target mapping position. That is, the target object does not need to be directly attached to the rim of the second target bowl layer. Taking a person as an example, the person's feet can be mapped onto the target mapping position of the second target bowl layer, while the rest of the person's body can be mapped onto that target mapping position of the second target bowl layer in a form that is perpendicular or nearly perpendicular to the target mapping position, rather than being mapped onto the rim of the second target bowl layer. This improves the intuitiveness of the user's observation of the target object, making it easier for the user to observe changes in the target object and enhancing the user experience.

[0099] Using the above method, multiple sub-background images are obtained by dividing the background image, and different sub-background images are mapped onto different first target bowl layers, which can better reflect the depth differences of different objects in the background. Simultaneously, by mapping the sub-background images and target objects separately, the visualization effect and spatial depth of the 3D background and target objects can be improved. Furthermore, the different distances from the bottom of the different bowl layers to the center point of the bowl bottom in the bowl-shaped model can avoid the problem of vertical stretching and deformation of the ground objects at the bottom of different bowl layers, making the final panoramic image of the target more realistic.

[0100] The following is a detailed explanation of step S102.

[0101] In one implementation, as shown in Figure 3, dividing the background image into multiple sub-background images in step S102 may include the following steps:

[0102] In step S1021, a first distance is determined between the road surface in the first road surface image and the vehicle.

[0103] In this step, the first distance between the road surface and the vehicle in the first road surface image can be determined based on the depth information corresponding to each pixel in the environmental image determined above.

[0104] In step S1022, the background image is divided into the multiple sub-background images according to the first distance.

[0105] In this step, the road surface in the first road surface image can be divided into multiple sub-road surfaces based on the first distance and multiple preset distance ranges. Then, based on the multiple sub-road surfaces, the background image is divided into multiple sub-background images, with each sub-road surface corresponding to a specific sub-background image. The preset distance ranges can be understood as pre-defined distance ranges corresponding to the bottom of each bowl layer. Based on the preset distance ranges and the first distance, the road surface can be divided into multiple sub-road surfaces, each corresponding to a different road surface area.

[0106] For example, for each preset distance range, the road surface within that preset distance range can be divided into a sub-road surface based on a first distance. For instance, a road surface within a distance of three to five meters from a vehicle can be considered a sub-road surface. Then, based on the distance corresponding to the sub-road surface, an image within that preset distance range in the background image is used as a sub-background image. The image area containing the sub-road surface in this sub-background image is the first road surface image, and the remaining portion is the first other background image. In other words, the image corresponding to the road surface within the preset distance range is the first road surface image, and the image corresponding to the remaining portion within the preset distance range (excluding the road surface) is the first other background image.

[0107] In another implementation, as shown in Figure 4, dividing the background image into multiple sub-background images in step S102 may include the following steps:

[0108] In step S1023, road surface elements in the first road surface image are identified.

[0109] In this implementation, different levels of segmentation can be performed according to different scenarios. When a vehicle travels on different roads, the road surface elements contained on the road surface are different. Therefore, in this step, the road surface elements in the first road surface image can be identified first. These road surface elements may include, but are not limited to, elements such as curbs and green belts.

[0110] In step S1024, the background image is divided into multiple sub-background images according to the position of the road surface element in the first road surface image.

[0111] In this step, based on the position of the road surface element in the first road surface image, the area within that position (assuming the road surface element is inside the direction of the vehicle) can be considered as one sub-region, and the area outside that position (assuming the direction of the vehicle is outside the direction of the road surface element) as another sub-region, resulting in multiple sub-background images corresponding to multiple sub-regions. Alternatively, if multiple road surface elements are included, the area between each road surface element can be considered as one sub-region, thus obtaining multiple sub-background images. The image containing the road surface in each sub-background image becomes the second road surface image, and the remaining portion becomes the second other background image.

[0112] For example, if the road surface element includes the curb, the location of the curb can be used as the dividing criterion to divide the road surface between the curb and the vehicle into one area, and the area outside the curb into another area, thus obtaining two sub-background images.

[0113] In some embodiments, determining the first target bowl layer corresponding to each sub-background image from multiple bowl layers of a pre-built bowl-shaped model in step S103 may include: for each sub-background image, determining the first target bowl layer corresponding to the sub-background image based on a second distance between the road surface and the vehicle in the sub-background image.

[0114] In this step, the layers can be assigned sequentially according to the second distance between the road surface and the vehicle in each sub-background image, from smallest to largest, to obtain the first target bowl layer corresponding to each sub-background image.

[0115] For example, taking a scenario comprising four sub-background images (i.e., sub-background image 1, sub-background image 2, sub-background image 3, and sub-background image 4) and four bowl layers (i.e., L1, L2, L3, and L4 in Figure 2), with the second distances corresponding to sub-background images 1, 2, 3, and 4 gradually increasing, the first target bowl layer corresponding to each sub-background image can be determined sequentially according to the second distance of each sub-background image, in order of increasing distance from the vehicle body. That is, the first target bowl layer corresponding to sub-background image 1 is L1, the first target bowl layer corresponding to sub-background image 2 is L2, the first target bowl layer corresponding to sub-background image 3 is L3, and the first target bowl layer corresponding to sub-background image 4 is L4.

[0116] In some embodiments, as shown in FIG5, the step S104 described above, which involves determining the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers, may include the following steps:

[0117] In step S1041, a third distance is determined between the at least one target object and the vehicle.

[0118] Similarly, the third distance between the at least one target object and the vehicle can be determined based on the depth information corresponding to each pixel in the environmental image determined above.

[0119] In step S1042, the second target bowl layer corresponding to the at least one target object is determined based on the third distance and the plurality of the second distances.

[0120] In this step, a fourth distance that matches the third distance can be determined from multiple second distances, and the first target bowl layer where the sub-background image corresponding to the fourth distance is located is taken as the second target bowl layer corresponding to the target object. That is, by matching the third distance and the second distance, the first target bowl layer where the sub-background image corresponding to the target object is located can be determined, and then the target object can be mapped onto the first target bowl layer where the sub-background image is located, which is also the second target bowl layer.

[0121] The third distance can be defined as the closest distance between the target object and the vehicle body. If the third distance is equal to the second distance, or falls within the range corresponding to the second distance, then the second distance can be defined as the fourth distance.

[0122] In addition, if different target objects are assigned to different bowl layers based on the third distance and the distance information corresponding to the target object, a new third distance can be obtained first based on the average of the third distance and the distance information, and then the different target objects can be assigned to different bowl layers based on the new third distance using the above method.

[0123] In practical applications, the inventors discovered that during the process of mapping the object image and the background image separately, there might be a certain degree of missing or misaligned edge pixels in the object image, thus affecting the display effect of the target panoramic image. Based on the above, as shown in Figure 6, the method may further include the following steps:

[0124] In step S106, the background image is input into the pre-trained image generation model to obtain the target background image output by the image generation model.

[0125] The target background image includes the background image itself and the image corresponding to the background area occluded by the target object in the environment image. It's understood that the target object occludes the background behind it in the environment image. When the object image is extracted, the background image will not contain the background of the area where the target object is located; that is, the background image contains a missing portion. In this step, an image generation model can be used to recover the background image of the area occluded by the target object, thus ensuring the integrity of the background image.

[0126] In some embodiments, the network structure of the image generation model can be, for example, a Generative Adversarial Network (GAN). This GAN consists of a generator and a discriminator. During use, the generator can be trained first using images from a database, and then the discriminator can be trained while keeping the generator constant. For example, a real image can be used as input, the generator can produce a similar image, and the discriminator can then determine whether the similar image is real, thus outputting a generated image similar to the input image.

[0127] Accordingly, dividing the background image into multiple sub-background images in step S102 above may include: dividing the target background image into multiple sub-background images.

[0128] In this way, by further dividing the complete target background image into multiple sub-background images to achieve subsequent mapping, and then superimposing the object image corresponding to the mapped target object, the pixel effect of the target object's edge contour can be effectively improved, thereby enhancing the display effect of the target panoramic image.

[0129] The following example illustrates the target object mapping method in the above embodiments. Figure 7 is a top view of a bowl-shaped model according to an exemplary embodiment. As shown in Figure 7, the layered bowl bottoms are concentric circular surfaces centered on the vehicle center. Each bowl layer is an arc surface whose radius increases progressively from the vertical line from the vehicle center, extending to the upper edge of the bowl model. When a target object is detected in the environmental image, the distance information of the target object around the vehicle and the background objects is obtained based on the depth information corresponding to each pixel in the environmental image, thereby matching different bowl layers for different target objects. Furthermore, the adjustment scale factor corresponding to the target object can be obtained based on the third distance between the target object and the vehicle, and the object image corresponding to the target object can be adjusted accordingly using the adjustment scale factor, thereby mapping it onto different bowl layers. As shown in Figure 7, based on the third distance (depth information) between the target object and the vehicle, it can be determined that a single pedestrian, a group of pedestrians, a car, and a truck are identified in the L1, L2, L3, and L4 bowl layers, respectively. After the unobstructed target background image is mapped onto the multi-bowl model, the target object image is adjusted according to the scale factor, and then superimposed and mapped to different positions in the L1, L2, L3, and L4 bowl layers.

[0130] Using the above method, multiple sub-background images are obtained by dividing the background image, and different sub-background images are mapped onto different first target bowl layers, which can better reflect the depth differences of different objects in the background. Simultaneously, by mapping the sub-background images and target objects separately, the visualization effect and spatial depth of the 3D background and target objects can be improved. Furthermore, the different distances from the bottom of the different bowl layers to the center point of the bowl bottom in the bowl-shaped model can avoid the problem of vertical stretching and deformation of the ground objects at the bottom of different bowl layers, making the final panoramic image of the target more realistic.

[0131] Figure 8 is a block diagram of a panoramic image generation apparatus according to an exemplary embodiment. As shown in Figure 8, the apparatus 200 includes:

[0132] The acquisition module 201 is used to acquire an object image and a background image of at least one target object from the environmental images collected by the vehicle;

[0133] The segmentation module 202 is used to divide the background image into multiple sub-background images;

[0134] The first determining module 203 is used to determine the first target bowl layer corresponding to each sub-background image from multiple bowl layers of the pre-constructed bowl-shaped model; the bowl-shaped model is a model containing multiple bowl layers built with the center of the vehicle as the center point of the bowl bottom, and the distance from the bottom of different bowl layers to the center point of the bowl bottom is different; different sub-background images correspond to different first target bowl layers;

[0135] The second determining module 204 is used to determine the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers;

[0136] The mapping module 205 is used to map each sub-background image to the first target bowl layer corresponding to the sub-background image, and to map the at least one object image to the corresponding second target bowl layer, so as to obtain a target panoramic image of the environment in which the vehicle is located.

[0137] As one embodiment, the background image includes a first road surface image, and the segmentation module 202 is used to determine a first distance between the road surface in the first road surface image and the vehicle; and to segment the background image into the plurality of sub-background images according to the first distance.

[0138] As one embodiment, the segmentation module 202 is used to divide the road surface in the first road surface image into multiple sub-road surfaces according to the first distance and multiple preset distance ranges; and to divide the background image into multiple sub-background images according to the multiple sub-road surfaces, wherein each sub-road surface corresponds to a sub-background image.

[0139] As one embodiment, the background image includes a first road surface image, and the segmentation module 202 is used to identify road surface elements in the first road surface image; and to segment the background image into multiple sub-background images according to the position of the road surface elements in the first road surface image.

[0140] As one embodiment, the first determining module 203 is used to determine the first target bowl layer corresponding to each sub-background image based on the second distance between the road surface and the vehicle in the sub-background image.

[0141] As one embodiment, the second determining module 204 is used to determine a third distance between the at least one target object and the vehicle; and to determine the second target bowl layer corresponding to the at least one target object based on the third distance and a plurality of the second distances.

[0142] As one embodiment, as shown in FIG9, the device 200 further includes:

[0143] The prediction module 206 is used to input the background image into a pre-trained image generation model to obtain the target background image output by the image generation model;

[0144] The target background image includes the background image and the image corresponding to the background area in the environment image that is occluded by the target object;

[0145] The partitioning module 202 is used to partition the target background image into multiple sub-background images.

[0146] As one embodiment, the acquisition module 201 is used to input the environmental image into a pre-trained target detection model to obtain the object image and the background image corresponding to the at least one target object output by the target detection model.

[0147] As one embodiment, the background image also includes a first other background image in the background image besides the first road surface image. The sub-background image includes a second road surface image and a second other background image in the sub-background image besides the second road surface image. The mapping module 205 is used to map the second road surface image onto the bottom of the bowl corresponding to the first target bowl layer; and to map the second other background image onto the rim of the bowl corresponding to the first target bowl layer.

[0148] As one embodiment, the mapping module 205 is used to obtain the relative position information of the at least one target object in the sub-background image; determine the target mapping position of the at least one target object in the second target bowl layer according to the relative position information; and map the at least one object image to the target mapping position of the corresponding second target bowl layer.

[0149] As one embodiment, the acquisition module 201 is further configured to acquire, for each target object, the adjustment scale factor corresponding to the target object based on the second target bowl layer;

[0150] As shown in Figure 10, the device 200 further includes: an adjustment module 207, used to adjust the object image corresponding to the target object according to the adjustment scale factor to obtain the target object image;

[0151] The mapping module 205 is used to map at least one target object image to the target mapping position of the corresponding second target bowl layer.

[0152] Using the aforementioned device, multiple sub-background images are obtained by dividing the background image, and different sub-background images are mapped onto different first target bowl layers, which can better reflect the depth differences of different objects in the background. Simultaneously, by mapping the sub-background images and target objects separately, the visualization effect and spatial depth of the 3D background and target objects can be improved. Furthermore, the different distances from the bottom of different bowl layers to the center point of the bowl bottom in the bowl-shaped model can avoid the problem of vertical stretching and deformation of ground objects at the bottom of different bowl layers, making the final panoramic image of the target more realistic.

[0153] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0154] Figure 11 is a block diagram illustrating an electronic device 300 according to an exemplary embodiment. As shown in Figure 11, the electronic device 300 may include a processor 301 and a memory 302. The electronic device 300 may also include one or more of a multimedia component 303, an input / output (I / O) interface 304, and a communication component 305.

[0155] The processor 301 controls the overall operation of the electronic device 300 to complete all or part of the steps in the panoramic image generation method described above. The memory 302 stores various types of data to support the operation of the electronic device 300. This data may include, for example, instructions for any application or method operating on the electronic device 300, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Multimedia component 303 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 302 or transmitted via communication component 305. The audio component also includes at least one speaker for outputting audio signals. I / O interface 304 provides an interface between processor 301 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 305 is used for wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IoT, eMTC, or other 5G technologies, or combinations thereof, is not limited here. Therefore, the corresponding communication component 305 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.

[0156] In an exemplary embodiment, the electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the panoramic image generation method described above.

[0157] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the panoramic image generation method described above. For example, the computer-readable storage medium may be the memory 302 including the program instructions described above, which may be executed by the processor 301 of the electronic device 300 to complete the panoramic image generation method described above.

[0158] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described panoramic image generation method when executed by the programmable device.

[0159] As shown in Figure 12, this disclosure also provides a vehicle 400, including the electronic device 300 shown in Figure 11.

[0160] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0161] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0162] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

Claims

1. A method for generating panoramic images, characterized in that, The method includes: Obtain at least one object image and background image of a target object from environmental images captured by the vehicle; The background image is divided into multiple sub-background images; The first target bowl layer corresponding to each sub-background image is determined from multiple bowl layers of a pre-constructed bowl-shaped model; the bowl-shaped model is a model containing multiple bowl layers built with the center of the bowl bottom as the center point of the bowl bottom, and the distance from the bottom of different bowl layers to the center point of the bowl bottom is different; different sub-background images correspond to different first target bowl layers; Determine the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers; Each sub-background image is mapped onto the first target bowl layer corresponding to the sub-background image, and the at least one object image is mapped onto the corresponding second target bowl layer to obtain a target panoramic image of the environment in which the vehicle is located.

2. The method according to claim 1, characterized in that, The background image includes a first road surface image, and dividing the background image into multiple sub-background images includes: Determine a first distance between the road surface in the first road surface image and the vehicle; Based on the first distance, the background image is divided into the plurality of sub-background images.

3. The method according to claim 2, characterized in that, The step of dividing the background image into the plurality of sub-background images based on the first distance includes: Based on the first distance and multiple preset distance ranges, the road surface in the first road surface image is divided into multiple sub-road surfaces; The background image is divided into multiple sub-background images based on the multiple sub-road surfaces, and each sub-road surface corresponds to one of the sub-background images.

4. The method according to any one of claims 1 to 3, characterized in that, The background image includes a first road surface image, and dividing the background image into multiple sub-background images includes: Identify road surface elements in the first road surface image; Based on the position of the road surface elements in the first road surface image, the background image is divided into the plurality of sub-background images.

5. The method according to any one of claims 1 to 4, characterized in that, Determining the first target bowl layer corresponding to each of the sub-background images from multiple bowl layers of a pre-constructed bowl-shaped model includes: For each sub-background image, the first target bowl layer corresponding to the sub-background image is determined based on the second distance between the road surface and the vehicle in the sub-background image.

6. The method according to claim 5, characterized in that, Determining the second target bowl layer corresponding to the at least one target object from the plurality of bowl layers includes: Determine a third distance between the at least one target object and the vehicle; The second target bowl layer corresponding to the at least one target object is determined based on the third distance and the plurality of second distances.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The background image is input into a pre-trained image generation model to obtain the target background image output by the image generation model; The target background image includes the background image and the image corresponding to the background area in the environment image that is occluded by the target object; The step of dividing the background image into multiple sub-background images includes: The target background image is divided into multiple sub-background images.

8. The method according to any one of claims 1 to 7, characterized in that, The process of obtaining at least one target object image and background image from the environmental images captured by the vehicle includes: The environmental image is input into a pre-trained target detection model to obtain the object image and the background image corresponding to the at least one target object output by the target detection model.

9. The method according to any one of claims 1 to 8, characterized in that, The background image further includes a first other background image besides the first road surface image, and the sub-background image includes a second road surface image and a second other background image besides the second road surface image. Mapping each sub-background image to the corresponding first target bowl layer includes: Map the second road surface image onto the bottom of the bowl corresponding to the first target bowl layer; The second other background image is mapped onto the rim of the bowl corresponding to the first target bowl layer.

10. The method according to any one of claims 1 to 9, characterized in that, The step of mapping the at least one object image onto the corresponding second target bowl layer includes: Obtain the relative position information of the at least one target object in the sub-background image; Based on the relative position information, determine the target mapping position of the at least one target object in the second target bowl layer; Map the at least one object image to the target mapping position of the corresponding second target bowl layer.

11. The method according to claim 10, characterized in that, The method further includes: For each target object, the adjustment scale factor corresponding to the target object is obtained according to the second target bowl layer; Adjust the object image corresponding to the target object according to the adjustment scale factor to obtain the target object image; The step of mapping the at least one object image to the target mapping position of the corresponding second target bowl layer includes: At least one target object image is mapped to the target mapping position of the corresponding second target bowl layer.

12. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1 to 11.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program performs the steps of the method described in any one of claims 1 to 11.

14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.

15. A vehicle, characterized in that, Includes the electronic device as described in claim 12.

Citation Information

Patent Citations

  • Vehicle panorama generating method and apparatus

    CN106355546A

  • Pavement detecting method and device, terminal and storage medium

    CN108197590A

  • Vehicle panoramic look-around image generation method and system

    CN113362232A

  • Image processing method, image processing device and electronic equipment

    CN114519680A

  • Panoramic image generation method and device, medium, product and vehicle

    CN118570060A