A panoramic street view image generation method and system

By constructing the data set and using distortion noise modulation and conditional attention mechanisms, panoramic street scene images with real distortion are directly generated, which solves the problem of lack of consistency in the generation of three-dimensional street scene images in the prior art, and significantly improves the authenticity of the image and the performance of the autonomous driving system.

CN119850864BActive Publication Date: 2025-06-10HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323125.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-10
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

When generating three-dimensional street scene images in the prior art, two-dimensional annotations cannot effectively capture three-dimensional geometric information, resulting in the lack of consistency in the height, distance, depth and other dimensions of images at different perspectives, affecting the authenticity of the image and the performance of the autonomous driving system.

Method used

By constructing the data set, the features of the original panoramic image, BEV map and 3D target frame are extracted, and multi-scale coded features are generated through distortion noise modulation and conditional attention mechanisms, and panoramic street scene images with real distortion are directly generated to avoid the complexity of multi-view stitching.

Benefits of technology

It significantly improves the authenticity, consistency and visual quality of generated images, reduces the complexity of multi-view generation, and improves the applicability and real-timeness of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_20
    Figure QLYQS_20
  • Figure QLYQS_24
    Figure QLYQS_24
Patent Text Reader

Abstract

The present invention provides a method and system for generating panoramic street view images, belonging to the technical field of image data processing. The present invention directly generates a complete panoramic image with real distortion without the need for multi-view image stitching. It can generate a high-fidelity panoramic street view image with global brightness consistency, seamless connection, and real distortion while retaining precise geometric control. Moreover, it improves the authenticity and consistency of the image, reduces the complexity of multi-view generation, and significantly improves the applicability and real-time performance in the autonomous driving scenario. By introducing multi-scale geometric control and conditional encoding, combined with a pre-trained diffusion model, a panoramic street view image with real distortion is generated from multi-condition inputs such as road BEV maps, 3D object bounding boxes, camera poses, and text descriptions. During the generation process, geometric details such as road elevation and object height can be precisely controlled, significantly improving the training effect of 3D perception tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image data processing, and particularly relates to a panoramic street view image generation method and system. Background Art

[0002] In the field of autonomous driving, especially in the tasks of generating and enhancing autonomous driving datasets, generating high-fidelity three-dimensional street view images with complex geometric control has become a key requirement. The generation quality of three-dimensional street view images directly affects the autonomous driving system's understanding of the environment and decision-making ability.

[0003] In the related art, the generation of three-dimensional street view images is mainly achieved by stitching. That is, images from multiple perspectives (such as front view, left view, right view, etc.) are first generated, and then two-dimensional annotations are performed on the images from each perspective, such as two-dimensional bounding boxes or two-dimensional semantic segmentation. Then, the images from multiple perspectives are stitched together to form a complete three-dimensional street view image. Since two-dimensional annotations cannot effectively capture three-dimensional geometric information, the images generated from different perspectives lack consistency in key dimensions such as height, distance, and depth, resulting in problems such as uneven brightness and incoherent edge connection in the generated three-dimensional street view images, affecting the authenticity of the three-dimensional street view images, and further affecting the performance of the autonomous driving system in perception tasks. At the same time, it also limits the precise positioning of objects in the scene, thereby reducing the autonomous driving system's understanding of the environment and decision-making ability.

[0004] Therefore, it is necessary to provide a panoramic street view image generation method and system to solve the above problems. Summary of the Invention

[0005] The present invention provides a panoramic street view image generation method and system, which directly generates a complete panoramic image with real distortion, rather than generating a panoramic image with unrealistic distortion by stitching multiple perspective images. Through this method, on the basis of retaining precise geometric control, a high-fidelity panoramic street view image with global brightness consistency, seamless connection, and real distortion can be generated, and the authenticity and consistency of the image are improved, the complexity of multi-perspective generation is reduced, and the applicability and real-time performance in autonomous driving scenarios are significantly improved, which can effectively solve at least one of the technical problems involved in the background art.

[0006] To solve the above technical problems, the present invention is implemented as follows:

[0007] A panoramic street view image generation method includes the following steps:

[0008] Step S1, construct a dataset, which includes the original panoramic images of streets and road BEV maps under multiple autonomous driving scenarios, and label the scenario conditions for each panoramic image and road BEV map, where the scenario conditions include 3D target boxes, camera poses, and text descriptions;

[0009] Step S2, extract the image features of the original panoramic images, BEV maps, and 3D target boxes, the geometric features of the camera poses, and the text features of the text descriptions respectively; modulate the extracted image features through distortion noise and fuse them with the geometric features and text features to generate multi-scale encoded features;

[0010] Step S3, decode the multi-scale encoded features, and embed the scenario condition information into each layer in the decoding process through a conditional attention mechanism to generate a corrected panoramic image that can accurately reflect the scenario conditions;

[0011] Step S4, optimize the generated corrected panoramic image based on a classifier-free guidance training strategy, randomly discard some scenario conditions during the training process to improve the consistency between the corrected panoramic image and the scenario condition information, and avoid over-reliance on the scenario conditions.

[0012] As a preferred improvement, the autonomous driving scenarios include urban roads, highways, and rural roads.

[0013] As a preferred improvement, the following steps are also included before step S2: divide the dataset into different categories according to the different contents of the text descriptions, where:

[0014] The text descriptions of the first type of dataset are related to environmental elements and are used to describe the overall environment in the image; the text descriptions of the second type of dataset are related to dynamic targets in the scenario; the text descriptions of the third type of dataset are related to static targets in the scenario.

[0015] As a preferred improvement, the extracted image features are represented as , where represents the number of layers of the extraction network; R represents the real number space; represents the height of the feature map; represents the width of the feature map; represents the number of channels; the extracted geometric features are represented as: ; the extracted text features are represented as: , where represents the number of words in the text description; represents the embedding dimension of the text features.

[0016] As a preferred improvement, in step S2, the distortion noise modulation process of the image features specifically includes the following steps: The image features are first added with noise and then the noise is removed through a diffusion model. During the process of adding noise, distortion noise is introduced to simulate the possible geometric distortion in the panoramic image. The distortion noise adjusts the noise characteristics through a control function to make it conform to the characteristics of the road geometry and environmental optical distortion. The distortion noise at time step is generated by the formula:

[0017]

[0018] In the formula, represents the distortion noise generated at time step ; represents the distortion noise intensity; represents the distortion mapping function; represents the image feature input at time step ; represents the intensity of the control random noise; represents the standard Gaussian noise.

[0019] As a preferred improvement, the distortion mapping function is specifically selected as the radial distortion mapping function, which is expressed as:

[0020]

[0021] In the formula, represents the weight coefficient for controlling the radial distortion; represents the normalized radial distance of the pixel point to the image optical center.

[0022] As a preferred improvement, the multi-scale encoded feature is expressed as:

[0023]

[0024] In the formula, represents the activation function with batch normalization; represents the element-wise multiplication operation; and respectively represent the trainable weights of the text feature and the geometric feature.

[0025] As a preferred improvement, step S3 specifically includes the following steps:

[0026] Step S31: Perform an upsampling operation on the multi-scale encoded feature , and the upsampling process is expressed as:

[0027]

[0028] In the formula, Represents the feature map after upsampling; Represents the upsampling operation;

[0029] Step S32: Construct a noise filtering gate to remove noise and supplement missing details, avoiding information loss during the upsampling process. The noise filtering gate is implemented through a convolutional layer with batch normalization, expressed as:

[0030]

[0031] In the formula, Represents the convolutional weight of the noise filtering gate; Represents the filtered feature map;

[0032] Step S33, supplement the missing detailed information during the noise filtering process through a convolutional projection operation. The supplementation process is expressed as:

[0033]

[0034] In the formula, Is the convolutional projection operation, used to supplement the useful information that has been filtered out; Represents the decoded feature after supplementation;

[0035] Step S34, embed the scene condition information into each layer in the decoding process through a conditional attention mechanism to generate a panoramic image that can accurately reflect the scene conditions. The process is expressed as:

[0036]

[0037] In the formula, Represents the conditional attention mechanism; Represents the text feature; Represents the camera pose information; Represents the 3D object bounding box information; Represents the feature map after the conditional attention mechanism;

[0038] Step S35, restore the multi-scale encoded features to the original resolution through step-by-step upsampling, and enhance the details and quality of the image through the bilinear interpolation algorithm to output the final corrected panoramic image , expressed as:

[0039]

[0040] In the formula, Represents the bilinear interpolation upsampling operation.

[0041] As a preferred improvement, the objective function formed in the training stage of the classifier-free guidance strategy is expressed as:

[0042]

[0043] Wherein, represents the prediction error of the model; represents the true distortion; represents the weight coefficient of the distortion loss; represents the weight coefficient of the time step.

[0044] A system for implementing the above panoramic street view image generation method, comprising:

[0045] A dataset construction module for constructing a dataset, which includes original panoramic images of streets and road BEV maps under multiple autonomous driving scenarios, and annotating scene conditions for each panoramic image and road BEV map, where the scene conditions include 3D target boxes, camera poses, and text descriptions;

[0046] A multi-scale feature encoding network for respectively extracting the image features of the original panoramic image, BEV map, and 3D target box, the geometric features of the camera pose, and the text features of the text description; modulating the extracted image features by distortion noise and fusing them with the geometric features and text features to generate multi-scale encoded features;

[0047] A multi-scale feature decoding network for decoding the multi-scale encoded features, embedding the scene condition information into each layer in the decoding process through a conditional attention mechanism, and generating a corrected panoramic image that can accurately reflect the scene conditions;

[0048] An optimization module for optimizing the generated corrected panoramic image based on a classifier-free guidance training strategy, randomly discarding some scene conditions during the training process to improve the consistency between the corrected panoramic image and the scene condition information and avoid over-reliance on the scene conditions.

[0049] The beneficial effects of the present invention are as follows:

[0050] (1) Abandoning the traditional multi-view stitching method and directly generating a complete panoramic image effectively avoids the problems of uneven brightness and discontinuous edges between multi-view images in the prior art, and at the same time solves the problem of unrealistic distortion in the existing methods of directly generating panoramic images, such as excessive edge stretching and depth perception distortion, thereby significantly improving the overall consistency, realism, and visual quality of the generated images;

[0051] (2)Introduce the technology based on dynamic distortion noise control. During the noise addition process of the diffusion model, by adding radial distortion, a type of geometric deformation noise, to simulate the geometric distortion characteristics of real panoramic images, not only the detailed information in the image is retained, but also high-fidelity panoramic images that conform to real optical distortion can be generated, significantly improving the consistency of street view images in edge connection and depth perception, and providing more accurate perception data support for the autonomous driving system. At the same time, the combination of distortion noise and conditional inputs (such as camera pose, environmental factors) enables the generated images to more realistically reflect the geometric characteristics and visual effects in complex environments;

[0052] (3)Introduce multi-scale geometric control technology. Combining with the pre-trained diffusion model, it can accurately generate geometric details such as road elevation and target object height, which not only improves the authenticity of the image, but also significantly enhances the training effect of 3D perception tasks such as BEV map segmentation and 3D object detection, and enhances the performance of the autonomous driving model;

[0053] (4)By combining multi-conditional inputs such as camera pose, text description, 3D object bounding box, and road BEV map, it provides flexible image generation control capabilities; the text description can include additional conditions such as weather and time, and street view images with different environmental scenes can be generated to meet diverse application requirements;

[0054] (5)Different from the existing solutions that need to rely on multi-view geometric transformation, it directly generates high-quality panoramic images through conditional encoding, without relying on cumbersome view transformation or geometric stitching, simplifies the generation process and improves the computational efficiency. Specific implementation manner

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] This embodiment provides a method for generating panoramic street view images, including the following steps:

[0057] Step S1, construct a data set, which includes the original panoramic images of streets in multiple autonomous driving scenarios and road BEV maps, and label scene conditions for each panoramic image and road BEV map, where the scene conditions include 3D object bounding boxes, camera poses, and text descriptions.

[0058] The described autonomous driving scenarios include urban roads, highways, and rural roads. According to the different content described in the text, the dataset is divided into different categories, and through round-by-round training, the dynamic effect of the model is improved and the generation effect is refined.

[0059] The dataset is divided into three categories in total:

[0060] The text description of the first type of dataset is related to environmental elements, such as weather, time, street type, etc. This type of text is relatively concise and is used to describe the overall environment in the image;

[0061] The text description of the second type of dataset is related to dynamic targets in the scenario, such as the movement state, position of pedestrians and animals, and their interaction with the surrounding environment;

[0062] The text description of the third type of dataset is related to static targets in the scenario, such as detailed information about building styles, floor heights, building materials, etc.

[0063] Step S2: Extract the image features of the original panoramic image, BEV map, and 3D target box, the geometric features of the camera pose, and the text features of the text description respectively; modulate the extracted image features through distortion noise and fuse them with the geometric features and text features to generate multi-scale encoded features.

[0064] The extracted image features are represented as , where represents the number of layers of the extraction network; R represents the real number space; represents the height of the feature map; represents the width of the feature map; represents the number of channels. The image features are extracted through a multi-scale convolutional neural network.

[0065] The extracted geometric features are represented as: . The extraction process of the geometric features can capture key information such as road elevation, the height and position of the target object, and then embed the extracted geometric features into the same dimension as the image features through a geometric mapping function. The geometric features are extracted through Camera Pose Auto-Encoders.

[0066] The extracted text features are represented as: , where represents the number of words in the text description; represents the embedding dimension of the text features. The text features are extracted through a text encoder (such as BiGRU or Transformer).

[0067] The distortion noise modulation process of image features specifically includes the following steps: removing noise from the image features through a diffusion model, introducing distortion noise during the process of adding noise, and simulating the possible radial distortion in the original panoramic image. The distortion noise adjusts the noise characteristics through a control function to conform to the characteristics of the road geometry and environmental optical distortion. The distortion noise at time step is generated according to the formula:

[0068]

[0069] In the formula, represents the distortion noise generated at time step ; represents the distortion noise intensity; represents the distortion mapping function; represents the input of image features at time step ; represents the intensity of controlling random noise; represents the standard Gaussian noise.

[0070] In this embodiment, the distortion mapping function is specifically selected as the radial distortion mapping function, expressed as:

[0071]

[0072] In the formula, represents the weight coefficient for controlling radial distortion; represents the normalized radial distance from the pixel point to the image optical center.

[0073] Then the final multi-scale encoded feature is expressed as:

[0074]

[0075] In the formula, represents the activation function with batch normalization; represents the element-wise multiplication operation; and respectively represent the trainable weights of the text feature and the geometric feature.

[0076] During the noise addition stage of the diffusion model, the distortion noise affects the image generation process through dynamic adjustment, so as to ensure that the generated image can not only conform to the geometric consistency of the panorama, but also retain the real optical distortion characteristics. By introducing dynamic distortion noise, the present invention further improves the authenticity and visual consistency of the generated panoramic image on the basis of fusing geometric features and text features, especially the consistency in edge connection and depth perception. At the same time, the distortion perception mechanism significantly enhances the diversity and environmental adaptability of the generated image, making it more valuable in the application of the autonomous driving scenario.

[0077] In step S3, the multi-scale encoded features are decoded, and the scene condition information is embedded into each layer in the decoding process through the conditional attention mechanism to generate a corrected panoramic image that can accurately reflect the scene conditions.

[0078] Step S3 specifically includes the following steps:

[0079] Step S31: Perform upsampling on the multi-scale encoded features The upsampling process is expressed as:

[0080]

[0081] In the formula, represents the upsampled feature map; represents the upsampling operation;

[0082] Step S32: Construct a noise filtering gate to remove noise and supplement missing details, avoiding information loss during the upsampling process. The noise filtering gate is implemented through a convolutional layer with batch normalization and is expressed as:

[0083]

[0084] In the formula, represents the convolutional weight of the noise filtering gate; represents the feature map after noise filtering;

[0085] Step S33: Supplement the missing detailed information during the noise filtering process through convolutional projection operation. The supplementation process is expressed as:

[0086]

[0087] In the formula, is the convolutional projection operation used to supplement the useful information that has been filtered out; represents the decoded features after supplementation.

[0088] Step S34: Embed the scene condition information into each layer in the decoding process through the conditional attention mechanism to generate a panoramic image that can accurately reflect the scene conditions. The process is expressed as:

[0089]

[0090] In the formula, represents the conditional attention mechanism; represents the camera pose information; represents the 3D target box information; represents the feature map after the conditional attention mechanism;

[0091] Use text descriptions, camera pose information, and 3D bounding box information to conditionally constrain the image, ensuring the consistency of the generated panoramic image in both geometric and semantic dimensions, so that the generated image can accurately reflect road elevation, object height, and other scene condition information in a specific scene.

[0092] Step S35, restore the multi-scale encoded features to the original resolution through gradual upsampling, and enhance the details and quality of the image through the bilinear interpolation algorithm, and output the final corrected panoramic image , expressed as:

[0093]

[0094] In the formula, represents the bilinear interpolation upsampling operation.

[0095] After decoding, a high-fidelity and detail-rich panoramic image can be obtained, avoiding the disadvantages of traditional multi-view stitching methods, and effectively improving the quality and consistency of the generated image.

[0096] Step S4, optimize the generated corrected panoramic image based on a classifier-free guidance training strategy. During the training process, randomly discard some scene conditions to enhance the consistency between the corrected panoramic image and the scene condition information, and avoid over-reliance on scene conditions.

[0097] The classifier-free guidance strategy is based on the existing diffusion model. By introducing randomness during the training stage, the generalization ability of the model is increased. The formed objective function is expressed as:

[0098]

[0099] In the formula, represents the prediction error of the model; represents the true distortion; represents the weight coefficient of the distortion loss; represents the weight coefficient of the time step.

[0100] The key to classifier-free guidance is to introduce a "null" placeholder for the missing scene conditions during the inference stage, so that the model can still maintain the generation effect when facing the lack of real data, thus achieving the consistency of panoramic street view images. Through this random discard strategy, the situation where the model over-relies on certain specific scene conditions is effectively solved. Especially in the self-driving scenario, this classifier-free guidance training strategy can improve the robustness and applicability of the model under different environmental conditions, and further enhance the global consistency of the generated panoramic street view images from multiple perspectives.

[0101] Through multiple conditions input such as road BEV maps, 3D object bounding boxes, camera poses, and text descriptions, the present invention directly generates a complete panoramic street view image with real distortion by combining a diffusion model with added distortion noise, without the need for traditional multi-view image stitching processes, thus avoiding problems such as brightness differences between views and inconsistent edge connections. During the encoding process, road geometric structures and text description information are extracted simultaneously. Through precise geometric control, accurate generation of geometric details such as road elevation and object height is achieved, thereby significantly improving the global consistency and visual quality of the generated images. In addition, the text description can include additional conditions such as weather and time, making the generated street view images more adaptable and realistic in different environments and meeting the requirements of complex application scenarios.

[0102] To further improve the training effect of the system in panoramic street view image generation, the present invention also proposes a training strategy based on classifier-free guidance. In this strategy, during the training process, some scene-level conditions (such as camera poses, text descriptions, etc.) are randomly discarded. By introducing a classifier-free guidance mechanism, the model can still generate high-quality street view images without specific conditions. This method effectively avoids the over-reliance of the model on certain conditions and increases the generalization ability of the model. During the inference stage, the system introduces a "null" placeholder for the missing conditions and adjusts the generation process through a conditional denoising network to ensure that the generated images still have consistency and high-quality performance. Through this strategy, the present invention not only improves the global consistency of street view images but also reduces the impact of noise interference during the training process on the image generation quality, enabling the system to adapt to the requirements of more scenarios.

[0103] This embodiment also provides a panoramic street view image generation system, including:

[0104] A dataset construction module for constructing a dataset, which includes original panoramic images of streets in multiple autonomous driving scenarios and road BEV maps, and annotating scene conditions for each panoramic image and road BEV map. The scene conditions include 3D object bounding boxes, camera poses, and text descriptions;

[0105] A multi-scale feature encoding network for respectively extracting image features of the original panoramic image, BEV map, and 3D object bounding box, geometric features of the camera pose, and text features of the text description; modulating the extracted image features with distortion noise and fusing them with the geometric features and text features to generate multi-scale encoded features;

[0106] A multi-scale feature decoding network for decoding the multi-scale encoded features, embedding scene condition information into each layer of the decoding process through a conditional attention mechanism, and generating a corrected panoramic image that can accurately reflect the scene conditions.

[0107] Optimization module, which optimizes the generated corrected panoramic image based on a training strategy without classifier guidance. During the training process, some scene conditions are randomly discarded to improve the consistency between the corrected panoramic image and the scene condition information, and to avoid over-reliance on the scene conditions.

[0108] The embodiments of the present invention have been described above, but the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them belong to the protection scope of the present invention.

Claims

1. A method for generating a panoramic street view image, characterized in that: The steps include: Step S1, constructing a data set, wherein the data set includes original panoramic images of streets in multiple autonomous driving scenarios and road BEV maps, and annotating scene conditions for each panoramic image and road BEV map, wherein the scene conditions include a 3D target box, a camera pose, and a text description; Step S2, respectively extracting image features of the original panoramic image, BEV map and 3D target frame, geometric features of the camera posture and text features of the text description; fusing the extracted image features with the geometric features and text features after being modulated by distortion noise to generate multi-scale coding features; Step S3, decoding the multi-scale encoded features, embedding the scene condition information into each layer of the decoding process through the conditional attention mechanism, and generating a corrected panoramic image that can accurately reflect the scene conditions; Step S4, optimizing the generated corrected panoramic image based on a training strategy without classifier guidance, randomly discarding some scene conditions during the training process, improving the consistency between the corrected panoramic image and the scene condition information, and avoiding excessive dependence on the scene conditions; In step S2, the distortion noise modulation process of the image feature specifically includes the following steps: adding noise to the image feature through the diffusion model and then removing the noise, introducing distortion noise in the noise adding process to simulate the geometric distortion that may exist in the panoramic image, and adjusting the noise characteristics of the distortion noise through the control function to make it conform to the characteristics of the road geometry and the environmental optical distortion. The distortion noise is adjusted at the time step The generation formula is: In the formula, Represents the time step Distortion noise generated when Indicates the distortion noise intensity; represents the distortion mapping function; Represents the time step Image feature input at the time; Indicates the intensity of controlling random noise; represents standard Gaussian noise.

2. The method for generating a panoramic street view image according to claim 1, characterized in that: The autonomous driving scenarios include urban roads, highways and rural roads.

3. The method for generating a panoramic street view image according to claim 1, characterized in that: Before step S2, the following steps are also included: dividing the data set into different categories according to different contents of the text description, wherein: The text description of the first type of data set is related to environmental elements and is used to describe the overall environment in the image; the text description of the second type of data set is related to dynamic targets in the scene; the text description of the third type of data set is related to static targets in the scene.

4. The method for generating a panoramic street view image according to claim 1, characterized in that: The extracted image features are expressed as , where Indicates the number of layers of the extraction network; R represents the real number space; Indicates the height of the feature map; Indicates the width of the feature map; Represents the number of channels; the extracted geometric features are expressed as: ; The extracted text features are expressed as: , where Indicates the number of words in the text description; Represents the embedding dimension of text features.

5. The method for generating a panoramic street view image according to claim 1, characterized in that: The distortion mapping function is specifically selected as a radial distortion mapping function, which is expressed as: In the formula, Represents the weight coefficient for controlling radial distortion; Represents the normalized radial distance from the pixel to the optical center of the image.

6. The method for generating a panoramic street view image according to claim 5, characterized in that: Multi-scale encoding features It is expressed as: In the formula, represents the activation function with batch normalization; Represents element-wise multiplication operation; and Represent the trainable weights of text features and geometric features respectively.

7. The method for generating a panoramic street view image according to claim 6, characterized in that: Step S3 specifically includes the following steps: Step S31: Multi-scale encoding features Perform upsampling operation, the upsampling process is expressed as: In the formula, Represents the upsampled feature map; Represents an upsampling operation; Step S32: Construct a noise filter gate to remove noise and supplement missing details to avoid information loss during upsampling. The noise filter gate is implemented through a convolutional layer with batch normalization, which is expressed as: In the formula, Represents the convolution weight of the noise filtering gate; Represents the filtered feature map; Step S33, the missing detail information in the noise filtering process is supplemented by convolution projection operation, and the supplementation process is expressed as: In the formula, It is a convolution projection operation, which is used to supplement the useful information that has been filtered out; represents the decoded features after supplementation; Step S34, embedding the scene condition information into each layer of the decoding process through the conditional attention mechanism to generate a panoramic image that can accurately reflect the scene conditions. The process is expressed as: In the formula, Represents the conditional attention mechanism; Represents text features; Represents the camera posture information; Represents 3D target frame information; Represents the feature map after the conditional attention mechanism; Step S35: restore the multi-scale coding features to the original resolution by stepwise upsampling, and enhance the image details and quality by bilinear interpolation algorithm, and output the final corrected panoramic image. , expressed as: In the formula, Represents a bilinear interpolation upsampling operation.

8. The method for generating a panoramic street view image according to claim 7, characterized in that: The objective function formed in the training phase of the strategy without classifier guidance is expressed as: In the formula, represents the prediction error of the model; Indicates the real distortion; Represents the weight coefficient of distortion loss; Represents the weight coefficient of the time step.

9. A system for executing the method for generating a panoramic street view image according to any one of claims 1 to 8, characterized in that: include: A data set construction module is used to construct a data set, wherein the data set includes original panoramic images of streets in multiple autonomous driving scenarios and road BEV maps, and annotates scene conditions for each panoramic image and road BEV map, wherein the scene conditions include a 3D target box, a camera pose, and a text description; The multi-scale feature coding network is used to extract the image features of the original panoramic image, BEV map and 3D target box, the geometric features of the camera posture and the text features of the text description respectively; the extracted image features are modulated by distortion noise and fused with the geometric features and text features to generate multi-scale coding features; The multi-scale feature decoding network is used to decode the multi-scale encoded features and embed the scene condition information into each layer of the decoding process through the conditional attention mechanism to generate a corrected panoramic image that can accurately reflect the scene conditions; The optimization module optimizes the generated corrected panoramic image based on a training strategy without classifier guidance. During the training process, some scene conditions are randomly discarded to improve the consistency between the corrected panoramic image and the scene condition information and avoid excessive dependence on the scene conditions.