Road network image generation method and device, storage medium and electronic equipment
Through feature coding and iterative denoising processing technology, road network images that meet the road network generation conditions are automatically generated, solving the problem of low image generation efficiency of road network in the existing technology, and achieving more efficient and rich road network image generation.
Patent Information
- Application Number
- CN202311818723.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the generation efficiency of road network images in the road system is low, requiring a large number of experienced designers to participate, resulting in inefficiency.
By obtaining the road network generation conditions, feature encoding is used to obtain the road network feature vector, and iteratively denoising the reference noise based on the vector, obtain the road network hidden vector, and decode it to generate a matching road network image.
It realizes that the road network image that meets actual needs is automatically generated without the manual participation of designers, improving the generation efficiency and content richness of road network images in the road system.
Smart Images

Figure CN120219212A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular, to a method and apparatus for generating a road network image, a storage medium, and an electronic device. Background Art
[0002] In order to express the road system within a certain geographical area in a planar image, a road network image is usually used. The road network image here can show the distribution of main roads, highways, and arterial roads such as expressways in the geographical area, as well as branch roads such as associated streets and alleys.
[0003] For the above-mentioned road network image, the image construction methods provided by the current related technologies mainly include: experienced urban planners or game planners browse a large amount of road network data and then perform personalized manual design according to actual needs. However, when applying the above method to construct a road network image for a road system in a larger geographical area, a large number of experienced designers are required to participate in the process of generating the road network image, resulting in a problem of low generation efficiency of the road system.
[0004] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a method and apparatus for generating a road network image, a storage medium, and an electronic device, so as to at least solve the technical problem of low generation efficiency of the road network image of the road system.
[0006] According to one aspect of the embodiments of the present application, a method for generating a road network image is provided, including: obtaining road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image; performing feature encoding on the road network generation conditions to obtain a road network feature vector; performing iterative denoising processing on reference noise based on the road network feature vector to obtain a road network latent vector; and decoding the road network latent vector to generate a road network image matching the road network generation conditions.
[0007] According to another aspect of the embodiments of the present application, an apparatus for generating a road network image is further provided, including: an obtaining unit, configured to obtain road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image; an encoding unit, configured to perform feature encoding on the road network generation conditions to obtain a road network feature vector; a denoising processing unit, configured to perform iterative denoising processing on reference noise based on the road network feature vector to obtain a road network latent vector; and a decoding unit, configured to decode the road network latent vector to generate a road network image matching the road network generation conditions.
[0008] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the above method for generating a road network image when running.
[0009] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for generating a road network image as described above.
[0010] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above method for generating a road network image through the computer program.
[0011] In the embodiments of the present application, road network generation conditions are obtained, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image. Then, feature encoding is performed on the road network generation conditions to obtain a road network feature vector. Next, denoising iterative processing is performed on the reference noise based on the road network feature vector to obtain a road network latent vector. Furthermore, the road network latent vector is decoded to generate a road network image matching the road network generation conditions. In other words, by using the embodiments of the present application, denoising iterative processing is performed on the reference noise by using the road network feature vector to obtain a road network latent vector; and then the road network latent vector is decoded to generate a road network image matching the road network generation conditions. It achieves the effect of automatically generating a road network image incorporating the road network generation conditions specified in advance according to actual needs without the need for a designer to manually participate in the process of generating the road network image. Thus, it solves the problem of low efficiency in generating road network images of the road system in the prior art, and realizes the technical effects of improving the efficiency of generating road network images of the road system and enhancing the content richness of the road network images. Description of the Drawings
[0012] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0013] Figure 1 is a schematic diagram of an application environment of an optional method for generating a road network image according to an embodiment of the present application;
[0014] Figure 2 is a flowchart of an optional method for generating a road network image according to an embodiment of the present application;
[0015] Figure 3 It is a flowchart of an optional method for generating a road network image according to an embodiment of the present application;
[0016] Figure 4 It is a schematic diagram of an optional method for generating a road network image according to an embodiment of the present application;
[0017] Figure 5 It is a schematic diagram of an optional method for generating a road network image according to an embodiment of the present application;
[0018] Figure 6 It is a schematic diagram of an optional method for generating a road network image according to an embodiment of the present application;
[0019] Figure 7 It is a flowchart of an optional method for generating a road network image according to an embodiment of the present application;
[0020] Figure 8 It is a flowchart of an optional method for generating a road network image according to an embodiment of the present application;
[0021] Figure 9 It is a schematic structural diagram of an optional device for generating a road network image according to an embodiment of the present application;
[0022] Figure 10 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the description, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0025] According to one aspect of the embodiments of the present application, a method for generating a road network image is provided. Optionally, as an alternative implementation, the generation of the above-mentioned road network image can be, but is not limited to, applied to an environment such as Figure 1 shown. As Figure 1 shown, the terminal device 102 includes a memory 104 for storing various data generated during the operation of the terminal device 102, a processor 106 for processing and computing the above-mentioned various data, and a display 108 for displaying the road network image. The terminal device 102 can perform data interaction with the server 112 through the network 110. The server 112 is connected to the database 114, and the database 114 is used to store various data. The terminal device 102 can run an application program for generating a road network image.
[0026] Furthermore, the specific application process of the above method in the Figure 1 shown environment is as follows:
[0027] Execute step S102, and the terminal device 102 sends request information for requesting the generation of a road network image to the server 112 through the network 110.
[0028] Then execute step S104, and the server 112 obtains road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image.
[0029] Next, execute steps S106 - S110. The server 112 performs feature encoding on the road network generation conditions to obtain a road network feature vector. The server 112 performs denoising iteration processing on the reference noise based on the road network feature vector to obtain a road network latent vector. The server 112 decodes the road network latent vector to generate a road network image that matches the road network generation conditions.
[0030] Then execute step S112, and the server 112 sends the road network image to the terminal device 102 through the network 110.
[0031] Then, step S114 is executed, and the terminal device 102 displays a road network image as shown in Figure 1 (a).
[0032] In the embodiment of the present application, road network generation conditions are obtained, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image. Then, feature encoding is performed on the road network generation conditions to obtain a road network feature vector. Next, denoising iterative processing is performed on the reference noise based on the road network feature vector to obtain a road network latent vector. Furthermore, the road network latent vector is decoded to generate a road network image that matches the road network generation conditions. In other words, in the embodiment of the present application, by using the road network feature vector to perform denoising iterative processing on the reference noise to obtain a road network latent vector; and then decoding the road network latent vector to generate a road network image that matches the road network generation conditions. It achieves the automatic generation of a road network image incorporating the road network generation conditions specified in advance according to actual needs without the need for a designer to manually participate in the road network image generation process. Thus, it solves the problem of low generation efficiency of road network images in the existing technology and realizes the technical effects of improving the generation efficiency of road network images in the road system and enhancing the content richness of road network images.
[0033] Optionally, in this embodiment, the above terminal device may be a terminal device configured with a target client, and may include but are not limited to at least one of the following: mobile phones (such as Android mobile phones, iOS mobile phones, etc.), laptop computers, tablet computers, handheld computers, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, etc. The target client may be a video client, an instant messaging client, a browser client, an education client, etc. The above network may include but are not limited to: wired networks, wireless networks, where the wired network includes: local area networks, metropolitan area networks, and wide area networks, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The above server may be a single server, or a server cluster composed of multiple servers, or a cloud server. The above is only an example, and no limitation is made thereto in this embodiment.
[0034] Optionally, as an alternative solution, as shown in Figure 2 , the above method for generating a road network image includes:
[0035] S202, obtaining road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image.
[0036] Optionally, the above method for generating a road network image can be but is not limited to being applied to the following scenarios: during game production, generating a road network image for simulating the real road network structure of a city to generate a virtual game scene using the road network image; during the production of a city map, generating a road network image for displaying the road network structure of a city to produce a city map scene using the road network image; during the process of urban planning, generating a city road network image for assisting in urban road network planning to reasonably plan the road network structure in the city using the road network image.
[0037] Furthermore, the above road network generation conditions can be but are not limited to including: a road network condition atlas and a road network style label. The road network condition atlas further includes a road network attribute condition image for indicating the distribution information of surface ecological characteristics, a density condition image for indicating the density of surface feature distribution, and a road network style label. Among them, the road network attribute condition image for indicating the distribution information of surface ecological characteristics can be but is not limited to an image for indicating environmental characteristics of a specified type, such as an image including main roads, water bodies, and green spaces of a specified type. The density condition image for indicating the density of surface feature distribution can be but is not limited to an image for indicating that the surface feature distribution density is a specified density. The road network style label can be but is not limited to being used to indicate a label for identifying the style characteristics of a specified city. Specifically, different road network style labels are respectively used to identify different city styles.
[0038] S204, perform feature encoding on the road network generation conditions to obtain a road network feature vector.
[0039] It should be noted that in this embodiment, performing feature encoding on the road network generation conditions to obtain a road network feature vector can be but is not limited to including: performing feature encoding on the road network condition atlas to obtain a condition feature vector, and performing feature encoding on the road network style label to obtain a style feature vector.
[0040] Furthermore, performing feature encoding on the road network condition atlas to obtain a condition feature vector can be but is not limited to including: in the case where the road network generation conditions include a road network attribute condition image for indicating the distribution information of surface ecological characteristics, performing binarization processing on the road network attribute condition image to obtain a grayscale image; performing size scaling adjustment on the grayscale image to obtain an attribute condition feature vector, where the condition feature vector includes the attribute condition feature vector; in the case where the road network generation conditions include a density condition image for indicating the density of surface feature distribution, mapping each pixel value in the density condition image to a density vector; concatenating the density vectors to obtain a density condition feature vector, where the condition feature vector includes the density condition feature vector.
[0041] Furthermore, the above-mentioned feature encoding of the road network style label to obtain the style feature vector may include, but is not limited to: finding the style feature vector matching the road network style label in the style feature vector library; or encoding the road network style label according to the agreed vector dimension to obtain the style feature vector. It should be noted that the above-mentioned style feature vector library pre-stores a plurality of style feature vectors generated based on the urban style features corresponding to each of the plurality of cities. Among them, each style feature vector in the style feature vector library corresponds to the style feature of a city.
[0042] S206, perform denoising iterative processing on the reference noise based on the road network feature vector to obtain the road network hidden vector.
[0043] Optionally, in this embodiment, a multi-layer iterative denoising processing network may be used, but is not limited to, to perform denoising iterative processing on the reference noise based on the road network feature vector to obtain the road network hidden vector. Among them, each layer of the denoising processing network in the multi-layer iterative denoising processing network is used to perform denoising processing on the denoising processing result output by the previous layer of the denoising processing network by using the road network feature vector to obtain the denoising processing result of this layer.
[0044] S208, decode the road network hidden vector to generate a road network image matching the road network generation condition.
[0045] It should be noted that the above-mentioned decoding of the road network hidden vector to generate a road network image matching the road network generation condition may include, but is not limited to: using the decoder Decoder in the Variational Autoencoder Network (VAE) to decode the road network hidden vector to generate a road network image matching the road network generation condition.
[0046] Optionally, in this embodiment, before the above-mentioned obtaining of the road network generation condition, the above method may further include, but is not limited to: respectively training the decoding network for generating the road network image and the above-mentioned multi-layer iterative denoising processing network; or training the decoding network for generating the road network image, and in the case where the above decoding network training is completed, jointly training the multi-layer iterative denoising processing network by using the above decoding network.
[0047] As an optional implementation manner, the following steps are used to give an overall example explanation of the above-mentioned method for generating a road network image:
[0048] Obtain the road network generation condition, where the road network generation condition includes a road network condition atlas and a road network style label. Then, input the road network generation condition into the condition encoding module, and further use the condition encoding module to perform feature encoding on the road network generation condition to obtain a road network feature vector, where the road network feature vector includes a condition feature vector and a style feature vector.
[0049] Next, based on the road network feature vector, denoising iterative processing is performed on the reference noise to obtain the road network latent vector. Furthermore, the road network latent vector is input into the decoder VAE-Decoder in the VAE model, and the VAE-Decoder is used to decode the road network latent vector to generate a road network image corresponding to the road network condition map set and the road network style label.
[0050] In the embodiment of the present application, road network generation conditions are obtained, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image. Then, feature encoding is performed on the road network generation conditions to obtain a road network feature vector. Next, denoising iterative processing is performed on the reference noise based on the road network feature vector to obtain a road network latent vector. Furthermore, the road network latent vector is decoded to generate a road network image matching the road network generation conditions. In other words, by adopting the embodiment of the present application, denoising iterative processing is performed on the reference noise by using the road network feature vector to obtain a road network latent vector; and then the road network latent vector is decoded to generate a road network image matching the road network generation conditions. The technical effect of automatically generating a road network image incorporating the road network generation conditions specified in advance according to actual needs is achieved without the need for a designer to manually participate in the road network image generation process. Thus, the problem of low generation efficiency of the road network image in the road system in the prior art is solved, and the technical effects of improving the generation efficiency of the road network image in the road system and enhancing the content richness of the road network image are realized.
[0051] Optionally, as an alternative solution, it is characterized in that performing denoising iterative processing on the reference noise based on the road network feature vector to obtain a road network latent vector includes:
[0052] In the multi-layer iterative denoising processing network, denoising processing is performed on the reference noise by using the road network feature vector to obtain a road network latent vector, where each layer of the denoising processing network in the multi-layer iterative denoising processing network is used to perform denoising processing on the denoising result output by the previous layer of the denoising processing network by using the road network feature vector.
[0053] It should be noted that the above multi-layer iterative denoising processing network can, but is not limited to, using an image semantic segmentation network (Convolutional Networks for Biomedical Image Segmentation, abbreviated as U-Net network). Further, in the above multi-layer iterative denoising processing network, using the road network feature vector to denoise the reference noise to obtain the road network hidden vector can, but is not limited to, including: performing the following operations in the multi-layer iterative denoising processing network: when the i-th layer denoising processing network is the first layer denoising processing network, in the first layer denoising processing network, using the road network feature vector to predict the noise residual of the reference noise to obtain the first layer predicted noise vector, where i is a natural number greater than or equal to 1 and less than or equal to N, and N is the number of layers of the multi-layer iterative denoising processing network; when the i-th layer denoising processing network is not the first layer denoising processing network, obtaining the (i - 1)-th layer predicted noise vector output by the (i - 1)-th layer denoising processing network before the i-th layer denoising processing network; in the i-th layer denoising processing network, using the road network feature vector and the (i - 1)-th layer predicted noise vector to predict the noise residual to obtain the i-th layer predicted noise vector; performing denoising processing based on the i-th layer predicted noise vector to obtain the i-th layer denoising processing result; when the i-th layer denoising processing network is the last layer denoising processing network, using the i-th layer denoising processing result as the road network hidden vector.
[0054] As an alternative implementation, assume that the above multi-layer iterative denoising processing network is a U-Net network, and assume that the decoder in the VAE model is used to decode the road network hidden vector to generate a road network image, as shown by Figure 3 the following steps and use the architecture shown by Figure 4 to give an overall example explanation of the above road network image generation method:
[0055] Step S302, obtain the road network generation condition.
[0056] Step S304, input the road network generation condition into the condition encoding module 402 shown in Figure 4 , and then use the condition encoding module to perform feature encoding on the road network generation condition to obtain the road network feature vector.
[0057] Step S306, input the road network feature vector into the multi-attention denoising module 404 shown in Figure 4 , and in the U-Net network 404-1 included in the multi-attention denoising module 404, use the road network feature vector to denoise the reference noise to obtain the road network hidden vector.
[0058] Step S308, input the road network hidden vector into Figure 4The decoder VAE-Decoder 406 in the VAE model shown in [Figure 0], and use VAE-Decoder to decode the road network latent vector to generate a road network image.
[0059] In the embodiment of the present application, in the multi-layer iterative denoising processing network, the reference noise is denoised using the road network feature vector to obtain a road network latent vector, where each layer of the denoising processing network in the multi-layer iterative denoising processing network is used to denoise the denoising processing result output by the previous layer of the denoising processing network using the road network feature vector. In other words, by adopting the embodiment of the present application, in the multi-layer iterative denoising processing network, the reference noise is iteratively denoised using the road network feature vector to obtain a road network latent vector; and then the road network latent vector is decoded to generate a road network image that matches the road network generation condition. The technical effect of automatically generating a road network image incorporating the road network generation conditions specified in advance according to actual needs is achieved without the need for a designer to manually participate in the road network image generation process. Thereby solving the problem of low generation efficiency of the road network image of the road system in the prior art, and realizing the technical effects of improving the generation efficiency of the road network image of the road system and enhancing the content richness of the road network image.
[0060] Optionally, as an alternative solution, in the multi-layer iterative denoising processing network, denoising the reference noise using the road network feature vector to obtain a road network latent vector includes:
[0061] Perform the following operations in the multi-layer iterative denoising processing network:
[0062] When the i-th layer of the denoising processing network is the first layer of the denoising processing network, in the first layer of the denoising processing network, use the road network feature vector to predict the noise residual of the reference noise to obtain the first layer of predicted noise vector, where i is a natural number greater than or equal to 1 and less than or equal to N, and N is the number of layers of the multi-layer iterative denoising processing network;
[0063] When the i-th layer of the denoising processing network is not the first layer of the denoising processing network, obtain the (i - 1)-th layer of predicted noise vector output by the (i - 1)-th layer of the denoising processing network before the i-th layer of the denoising processing network; in the i-th layer of the denoising processing network, use the road network feature vector and the (i - 1)-th layer of predicted noise vector to predict the noise residual to obtain the i-th layer of predicted noise vector;
[0064] Perform denoising processing based on the i-th layer of predicted noise vector to obtain the denoising processing result of the i-th layer;
[0065] When the i-th layer of the denoising processing network is the last layer of the denoising processing network, use the denoising processing result of the i-th layer as the road network latent vector.
[0066] As an alternative implementation, assume that the above multi-layer iterative denoising processing network is a U-Net network. Assume that the multi-layer iterative denoising processing network includes 3 U-Net networks, namely U-Net network 1, U-Net network 2, and U-Net network 3. Using the architecture as shown in Figure 5 , the following steps are used to give an overall example and explanation of the above method for generating road network images:
[0067] As shown in Figure 5 , the conditional feature vector, the reference noise, and the style feature vector are input into U-Net network 1 in the multi-layer iterative denoising processing network 502. In U-Net network 1, using the conditional feature vector, the reference noise, and the style feature vector, the predicted noise vector 1 corresponding to the reference noise is predicted. Then, the predicted noise vector 1 is removed from the reference noise to obtain the first-layer denoising processing result.
[0068] Then, the processing time vector of U-Net network 2 is determined using the current time. Next, as still shown in Figure 5 , the conditional feature vector, the first-layer denoising processing result, and the style feature vector are input into U-Net network 2 in the multi-layer iterative denoising processing network 502. Using U-Net network 2 with the conditional feature vector, the first-layer denoising processing result, and the style feature vector, the predicted noise vector 2 corresponding to the first-layer denoising processing result is predicted. Then, the predicted noise vector 2 is removed from the first-layer denoising processing result to obtain the second-layer denoising processing result.
[0069] The processing time vector of U-Net network 3 is determined using the current time. Next, as still shown in Figure 5 , the conditional feature vector, the second-layer denoising processing result, and the style feature vector are input into U-Net network 3 in the multi-layer iterative denoising processing network 502. Using U-Net network 3 with the conditional feature vector, the second-layer denoising processing result, and the style feature vector, the predicted noise vector 3 corresponding to the second-layer denoising processing result is predicted. Then, the predicted noise vector 3 is removed from the second-layer denoising processing result to obtain the third-layer denoising processing result.
[0070] Furthermore, it is determined that U-Net network 3 is the last denoising processing network in the multi-layer iterative denoising processing network 502, and the third-layer denoising processing result is determined as the road network latent vector.
[0071] It should be noted that the above embodiment is a reference embodiment of the above method for generating road network images. The number of layers of the above multi-layer iterative denoising processing network is not limited to 3 layers, and the number of layers of the multi-layer iterative denoising processing network can also be flexibly set according to requirements. No limitation is made in this embodiment.
[0072] In the embodiment of the present application, in a multi-layer iterative denoising processing network, a reference noise is subjected to denoising iteration processing by using a road network feature vector to obtain a road network hidden vector; and then the road network hidden vector is decoded to generate a road network image that matches the road network generation condition. In this way, without the need for manual participation of a designer in the process of generating the road network image, a road network image incorporating road network generation conditions pre-specified according to actual requirements is automatically generated. Thus, the problem of low generation efficiency of the road network image of the road system in the prior art is solved, and the technical effects of improving the generation efficiency of the road network image of the road system and enhancing the content richness of the road network image are achieved.
[0073] Optionally, as an alternative solution, when obtaining the (i-1)-th layer predicted noise vector output by the (i-1)-th layer denoising processing network before the i-th layer denoising processing network, it further includes: determining the processing time vector of the i-th layer denoising processing network.
[0074] Optionally, in this embodiment, it is also possible but not limited to determining the processing time vector of the i-th layer denoising processing network before obtaining the (i-1)-th layer predicted noise vector output by the (i-1)-th layer denoising processing network before the i-th layer denoising processing network, or after obtaining the (i-1)-th layer predicted noise vector output by the (i-1)-th layer denoising processing network before the i-th layer denoising processing network, where i is a natural number greater than or equal to 1 and less than or equal to N, and N is the number of layers of the multi-layer iterative denoising processing network. No limitation is made in this regard in this embodiment.
[0075] In the i-th layer denoising processing network, performing noise residual prediction by using the road network feature vector and the (i-1)-th layer predicted noise vector to obtain the i-th layer predicted noise vector includes: concatenating the (i-1)-th layer denoising processing result and the conditional feature vector in the road network feature vector to obtain the i-th layer conditional vector, where the conditional feature vector is a vector in the road network feature vector for indicating the surface environment feature distribution information; performing weighted summation on the style feature vector and the processing time vector in the road network feature vector to obtain the i-th layer style vector, where the style feature vector is a vector in the road network feature vector for indicating the road network style feature distribution information; and performing noise residual prediction based on the i-th layer conditional vector and the i-th layer style vector to obtain the i-th layer predicted noise vector.
[0076] Optionally, as an alternative implementation manner, the concatenating the (i-1)-th layer denoising processing result and the conditional feature vector to obtain the i-th layer conditional vector may but is not limited to include:
[0077] C = Concat([z t , c]) (1)
[0078] where the above z tUsed to represent the denoising processing result of the (i - 1)-th layer, where c above is used to represent the conditional feature vector, and C above is used to represent the conditional vector of the i-th layer.
[0079] Further, the weighted sum of the style feature vector and the processing time vector in the road network feature vector to obtain the style vector of the i-th layer may include, but is not limited to:
[0080] S = t emb + c style (2)
[0081] Wherein, t above emb is used to represent the processing time vector, c above style is used to represent the style feature vector, and S above is used to represent the style vector of the i-th layer.
[0082] Still further, the noise residual prediction based on the conditional vector of the i-th layer and the style vector of the i-th layer may include, but is not limited to:
[0083] R1 = F φ (C, S) = F φ ((Concat([z t , c]))), (t emb + c style )) (3)
[0084] Wherein, R1 above is used to represent the predicted noise vector of the i-th layer, F above is used to represent the function corresponding to the denoising processing network of the i-th layer, and φ above is used to represent the network parameters in F.
[0085] In the embodiments of the present application, by adopting the method of integrating the time vector in the process of denoising the reference noise with the conditional feature vector and the style feature vector in the multi-layer iterative denoising processing network, the purpose of distinguishing the denoising processing of each denoising processing network in the multi-layer iterative denoising processing network is achieved. Furthermore, the denoising processing of each denoising processing network in the multi-layer iterative denoising processing network is not confused, thereby achieving the technical effect of improving the accuracy of the denoising processing.
[0086] Optionally, as an alternative solution, concatenating the denoising processing result of the (i - 1)-th layer and the conditional feature vector in the road network feature vector to obtain the conditional vector of the i-th layer includes:
[0087] S1. When the conditional feature vector includes the attribute conditional feature vector and the density conditional feature vector, concatenating the denoising processing result of the (i - 1)-th layer and the attribute conditional feature vector to obtain the attribute vector of the i-th layer;
[0088] S2. After dimension adjustment of the attribute vector of the i-th layer, perform weighted summation with the density condition feature vector to obtain the condition vector of the i-th layer.
[0089] It should be noted that the above-mentioned attribute condition feature vector can be but is not limited to being used to indicate the feature vector corresponding to the environmental feature of a specified type, such as the feature vector corresponding to the main road, water body, and green space of a specified type. The above-mentioned density condition feature vector can be but is not limited to being used to indicate the feature vector corresponding to the specified image density.
[0090] Furthermore, the above-mentioned dimension adjustment of the attribute vector of the i-th layer can be but is not limited to including: adjusting the feature dimension of the attribute vector of the i-th layer to the first dimension, where the above-mentioned first dimension is the feature dimension corresponding to the density condition feature vector.
[0091] Optionally, in this embodiment, the concatenation of the denoising processing result of the (i - 1)-th layer and the attribute condition feature vector to obtain the attribute vector of the i-th layer can be but is not limited to being exemplified and explained based on the following example:
[0092] C1 = Concat([z t , c cond ) (4)
[0093] Among them, the above-mentioned z t is used to represent the denoising processing result of the (i - 1)-th layer, the above-mentioned c cond is used to represent the attribute condition feature vector, and the above-mentioned C1 is used to represent the attribute vector of the i-th layer.
[0094] Furthermore, the weighted summation of the dimension-adjusted attribute vector of the i-th layer and the density condition feature vector to obtain the condition vector of the i-th layer can be but is not limited to being exemplified and explained based on the following example:
[0095] C = C2 + c density (5)
[0096] Among them, the above-mentioned C2 is used to represent the vector obtained after feature dimension adjustment of the attribute vector C1 of the i-th layer, the above-mentioned c density is used to represent the density condition feature vector, and the above-mentioned C is used to represent the condition vector of the i-th layer.
[0097] Still further, the noise residual prediction based on the condition vector of the i-th layer and the style vector of the i-th layer can be but is not limited to being exemplified and explained based on the following example:
[0098] R1 = F φ (C, S) = F φ ((C2 + c density ), (t emb+c style )) (6)
[0099] Wherein, the above R1 is used to represent the denoising processing result of the i-th layer, the above F is used to represent the function corresponding to the denoising processing network of the i-th layer, and the above φ is used to represent the network parameters in F.
[0100] It should be noted that the above embodiment is an optional embodiment listed for easy explanation. Other conventional calculation methods can also be used to splice the denoising processing result of the (i - 1)-th layer and the attribute condition feature vector to obtain the attribute vector of the i-th layer. After dimension adjustment processing of the attribute vector of the i-th layer, weighted summation is performed with the density condition feature vector to obtain the condition vector of the i-th layer. In this embodiment, no limitation is imposed on this.
[0101] In the embodiment of the present application, when the condition feature vector includes an attribute condition feature vector and a density condition feature vector, the denoising processing result of the (i - 1)-th layer and the attribute condition feature vector are spliced to obtain the attribute vector of the i-th layer; the attribute vector of the i-th layer after dimension adjustment is weighted and summed with the density condition feature vector to obtain the condition vector of the i-th layer. In other words, in the embodiment of the present application, by integrating the density condition feature vector in the process of denoising the reference noise with the condition feature vector and the style feature vector in the multi-layer iterative denoising processing network, the road network hidden vector obtained after denoising processing carries the specified density condition feature. Furthermore, the road network image generated through the road network hidden vector can be presented with the specified image density. That is to say, by adopting the embodiment of the present application, without manual analysis, a road network image integrated with the specified condition features and density features can be automatically generated, thereby achieving the technical effects of improving the generation efficiency of the road network image of the road system and enriching the content of the road network image.
[0102] Optionally, as an alternative solution, after decoding the road network hidden vector to generate a road network image matching the road network generation conditions, it further includes:
[0103] S1, obtaining a mask image;
[0104] S2, in the multi-layer iterative denoising processing network, using the road network feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image to update the road network image area blocked by the mask image in the road network image.
[0105] Optionally, the obtaining of the mask image described above may include, but is not limited to: after obtaining the initialized mask image, performing encoding processing on the initialized mask image to obtain the above-mentioned mask image. Specifically, performing binarization processing on the obtained initialized mask image to obtain a grayscale image. Then, performing compression processing on the grayscale image to obtain the above-mentioned mask image. For example, for an initialized mask image with a size of W×H, performing binarization processing to obtain a grayscale image with a size of W×H. Furthermore, reducing the size of the above grayscale image from W×H to obtain the above-mentioned mask image, where both W and H are positive integers.
[0106] It should be noted that in this embodiment, the above-mentioned mask image may be used to indicate, but is not limited to, the image corresponding to the target area in the road network image. Specifically, the above-mentioned target area is the image area to be updated in the above-mentioned road network image.
[0107] Furthermore, in the multi-layer iterative denoising processing network, updating the road network image area occluded by the mask image in the road network image by using the road network feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image may include, but is not limited to: in the multi-layer iterative denoising processing network, updating the road network image area occluded by the mask image in the road network image by using the conditional feature vector, the style feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image.
[0108] In the embodiment of the present application, a mask image is obtained. Then, in the multi-layer iterative denoising processing network, the road network image area occluded by the mask image in the road network image is updated by using the road network feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image. Furthermore, without manual analysis, the road network image can be automatically updated according to requirements, thereby achieving the technical effect of improving the update efficiency of the road network image.
[0109] Optionally, as an alternative solution, in the multi-layer iterative denoising processing network, updating the road network image area occluded by the mask image in the road network image by using the road network feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image includes:
[0110] Perform the following operations in the multi-layer iterative denoising processing network:
[0111] Obtain the denoising processing result of the (j - 1)-th layer output by the (j - 1)-th layer denoising processing network before the j-th layer denoising processing network, where j is a natural number greater than or equal to 1 and less than or equal to N;
[0112] In the denoising processing network of the j-th layer, the masked denoising processing result of the j-th layer is obtained in the following manner:
[0113] Concatenate the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer, and perform weighted summation on the style feature vector and the processing time vector of the denoising processing network of the j-th layer to obtain the style vector of the j-th layer;
[0114] Based on the conditional vector of the j-th layer and the style vector of the j-th layer, perform noise residual prediction to obtain the intermediate predicted noise vector of the j-th layer;
[0115] Fuse the intermediate predicted noise vector of the j-th layer with the mask vector to obtain the occlusion region vector corresponding to the road network image region occluded by the masked image;
[0116] Add noise to the image region not occluded by the masked image to obtain a noise-added result, and fuse the noise-added result with the mask reference vector corresponding to the mask vector to obtain the non-occluded region vector;
[0117] Perform weighted summation on the occlusion region vector and the non-occluded region vector to obtain the masked predicted noise vector of the j-th layer;
[0118] Based on the masked predicted noise vector of the j-th layer, perform denoising processing to obtain the masked denoising processing result of the j-th layer.
[0119] Optionally, as an alternative implementation, the concatenation of the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer can be exemplified and explained by the following examples, but are not limited thereto:
[0120] C3 = Concat([z t1 , c1]) (7)
[0121] Wherein, the above z t1 is used to represent the denoising processing result of the (j - 1)-th layer, the above c1 is used to represent the conditional feature vector, and the above C3 is used to represent the conditional vector of the j-th layer.
[0122] Furthermore, the weighted summation of the style feature vector and the processing time vector of the denoising processing network of the j-th layer to obtain the style vector of the j-th layer can be exemplified and explained by the following examples, but are not limited thereto:
[0123] S1 = t emb1 + c style1 (8)
[0124] Wherein, the above t emb1 is used to represent the processing time vector, and the above c style1For representing the style feature vector, the above S1 is used to represent the style vector of the j-th layer, and the weight of the weighted sum is 1.
[0125] It should be noted that the above embodiment is an optional embodiment listed for easy explanation. Other conventional calculation methods can also be used to splice the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer, and the style feature vector and the processing time vector of the denoising processing network of the j-th layer are weighted and summed to obtain the style vector of the j-th layer. In this embodiment, no limitation is imposed on this.
[0126] Optionally, it can be but is not limited to the following examples to give an illustrative explanation of predicting the noise residual based on the conditional vector of the j-th layer and the style vector of the j-th layer to obtain the intermediate predicted noise vector of the j-th layer:
[0127] R2 = F φ (C3, S1) = F φ ((Concat([z t1 , c1]))), (t emb1 + c style1 )) (9)
[0128] Among them, the above R2 is used to represent the intermediate predicted noise vector of the j-th layer, the above F is used to represent the function corresponding to the denoising processing network of the j-th layer, and the above φ is used to represent the network parameters in F.
[0129] Furthermore, it can be but is not limited to the following examples to give an illustrative explanation of fusing the intermediate predicted noise vector of the j-th layer with the mask vector to obtain the occlusion region vector corresponding to the road network image region occluded by the masked image:
[0130] R3 = R2 × c mask = F φ ((Concat([z t1 , c1]))), (t emb1 + c style1 )) × c mask (10)
[0131] Among them, the above R3 is used to represent the occlusion region vector, and the above c mask is used to represent the mask vector.
[0132] It should be noted that the above-mentioned adding noise to the image region not occluded by the masked image to obtain the noise addition result can be but is not limited to including: performing multi-layer Gaussian noise addition on the road network image to obtain the noise addition result.
[0133] Optionally, in this embodiment, the fusion of the above-mentioned noise addition result and the mask reference vector corresponding to the mask vector to obtain the unoccluded region vector, and the weighted summation of the occluded region vector and the unoccluded region vector to obtain the j-th layer mask prediction noise vector can be, but is not limited to, illustrated by the following examples:
[0134] R5 = R3 + R4 = R3 + q(x′) * (1 - c mask ) (11)
[0135] Among them, the above-mentioned R5 is used to represent the j-th layer mask prediction noise vector, the above-mentioned x′ is used to represent the road network image, the above-mentioned q(x′) is used to represent the noise addition result, the above-mentioned (1 - c mask ) is used to represent the mask reference vector, the above-mentioned R4 is used to represent the unoccluded region vector, and the weight of the weighted summation is 1.
[0136] As an alternative implementation, the above-mentioned concatenation of the j - 1-th layer denoising processing result and the conditional feature vector to obtain the j-th layer conditional vector: when the conditional feature vector includes an attribute conditional feature vector and a density conditional feature vector, the j - 1-th layer denoising processing result and the attribute conditional feature vector are concatenated to obtain the j-th layer attribute vector; after dimension adjustment processing of the j-th layer attribute vector, it is weighted and summed with the density conditional feature vector to obtain the j-th layer conditional vector.
[0137] Optionally, the above-mentioned concatenation of the j - 1-th layer denoising processing result and the conditional feature vector to obtain the j-th layer conditional vector, and the weighted summation of the style feature vector and the processing time vector of the j-th layer denoising processing network to obtain the j-th layer style vector, and the noise residual prediction based on the j-th layer conditional vector and the j-th layer style vector to obtain the j-th layer intermediate prediction noise vector, and the fusion of the j-th layer intermediate prediction noise vector and the mask vector to obtain the occluded region vector corresponding to the road network image region occluded by the mask image, and the noise addition processing of the image region not occluded by the mask image to obtain the noise addition result, and the fusion of the noise addition result and the mask reference vector corresponding to the mask vector to obtain the unoccluded region vector can be, but is not limited to, illustrated by the following examples:
[0138] C4 = f conv (Concat([z t1 , c cond1 )) (12)
[0139] R6 = F φ ((C5 + c density1 ), (t emb1 + c style1 )) * c mask + q(x′) * (1 - cmask ) (13)
[0140] Among them, the above R6 is used to represent the predicted noise vector of the j-th layer mask, the above C4 is used to represent the attribute vector of the j-th layer, the above C5 is used to represent the vector obtained by adjusting the feature dimension of the attribute vector C4 of the j-th layer, and the above c cond1 is used to represent the attribute conditional feature vector, and the above c density1 is used to represent the density conditional feature vector.
[0141] It should be noted that the above embodiments are optional embodiments listed for easy explanation. Other conventional calculation methods can also be used to splice the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer, and perform weighted summation on the style feature vector and the processing time vector of the denoising processing network of the j-th layer to obtain the style vector of the j-th layer. In this embodiment, no limitation is made in this regard.
[0142] Optionally, in this embodiment, the above denoising processing based on the predicted noise vector of the j-th layer mask to obtain the denoising processing result of the j-th layer mask may include, but is not limited to: removing the predicted noise vector of the j-th layer mask from the denoising processing result of the (j - 1)-th layer to obtain the denoising processing result of the j-th layer mask.
[0143] Optionally, assume that the above multi-layer iterative denoising processing network is a U-Net network, and assume that there are 2 U-Net networks in the multi-layer iterative denoising processing network, namely Figure 6 the U-Net network 602 and the U-Net network 604 shown. Using the architecture shown in Figure 6 the following steps are used to give an overall example and explanation of the above road network image generation method:
[0144] As Figure 6 shown, in the U-Net network 602, the first intermediate predicted noise vector in the reference noise is obtained by using the conditional feature vector, the style feature vector and the reference noise, and then the first intermediate predicted noise vector is multiplied by the mask vector to obtain the occlusion area vector. Then, multi-layer Gaussian noise addition processing is performed on the road network image to obtain the noise addition result. Further, the noise addition result is multiplied by the mask reference vector corresponding to the mask vector to obtain the unoccluded area vector. Then, the predicted noise vector of the first layer mask is obtained by using the occlusion area vector and the unoccluded area vector. Further, the predicted noise vector of the first layer mask is removed from the reference noise to obtain the denoising processing result of the first layer mask.
[0145] Then, still as Figure 6As shown in the figure, in the U-Net network 604, the conditional feature vector, the style feature vector, and the denoising processing result of the first-layer mask are used to obtain the second intermediate predicted noise vector in the denoising processing result of the first-layer mask. Then, the second intermediate predicted noise vector is multiplied by the mask vector to obtain the occlusion region vector. Next, the second-layer mask predicted noise vector is obtained by using the occlusion region vector and the unoccluded region vector. Thus, the second-layer mask predicted noise vector is removed from the denoising processing result of the first-layer mask to obtain the denoising processing result of the second-layer mask. It is determined that the U-Net network 604 is the last denoising network, and the denoising processing result of the second-layer mask is determined as the road network latent vector.
[0146] Further, still as Figure 6 shown, the road network latent vector is input into the decoder 606, and the decoder 606 is used to decode the road network latent vector to generate an updated road network image.
[0147] In the embodiment of the present application, the following operations are performed in the multi-layer iterative denoising processing network: obtaining the denoising processing result of the (j - 1)-th layer output by the (j - 1)-th layer denoising processing network before the j-th layer denoising processing network, where j is a natural number greater than or equal to 1 and less than or equal to N; in the j-th layer denoising processing network, the j-th layer predicted noise vector output is obtained by the following method: concatenating the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the j-th layer conditional vector, and performing a weighted sum of the style feature vector and the processing time vector of the j-th layer denoising processing network to obtain the j-th layer style vector; performing noise residual prediction based on the j-th layer conditional vector and the j-th layer style vector to obtain the j-th layer intermediate predicted noise vector; fusing the j-th layer intermediate predicted noise vector with the mask vector to obtain the occlusion region vector corresponding to the road network image region occluded by the masked image; adding noise to the image region not occluded by the masked image to obtain a noise-added result, and fusing the noise-added result and the mask reference vector corresponding to the mask vector to obtain the unoccluded region vector; performing a weighted sum of the occlusion region vector and the unoccluded region vector to obtain the j-th layer mask predicted noise vector; performing denoising processing based on the j-th layer mask predicted noise vector to obtain the denoising processing result of the j-th layer mask. Furthermore, without manual analysis, the road network image can be automatically updated according to requirements, thus achieving the technical effect of improving the update efficiency of the road network image.
[0148] Optionally, as an alternative solution, feature encoding of the road network generation conditions to obtain the road network feature vector includes:
[0149] When the road network generation conditions include a road network attribute condition image for indicating the distribution information of surface environmental feature, the road network attribute condition image is binarized to obtain a grayscale image; the size of the grayscale image is scaled and adjusted to obtain an attribute condition feature vector, where the condition feature vector includes the attribute condition feature vector.
[0150] It should be noted that the above road network attribute condition image for indicating the distribution information of surface environmental feature can be, but is not limited to, an image for indicating an image including specified types of environmental features, such as an image including specified types of main roads, water bodies, and green spaces.
[0151] Assume that the above road network attribute condition image is an image including specified types of environmental features. The above method is illustrated and explained by the following steps: A road network attribute condition image with a size of W×H including main roads, water bodies, and green spaces of the target type is binarized to obtain a grayscale image of the road network attribute condition image with a size of W×H. Then, the size of the above grayscale image is reduced from W×H to W / 4×H / 4 to obtain an attribute condition feature vector, where both W and H are positive integers.
[0152] When the road network generation conditions include a density condition image for indicating the density of surface element distribution, each pixel value in the density condition image is mapped to a density vector respectively; the respective density vectors are concatenated to obtain a density condition feature vector, where the condition feature vector includes the density condition feature vector.
[0153] It should be noted that the above density condition image for indicating the density of surface feature distribution can be, but is not limited to, an image for indicating that the surface feature distribution density is a specified density. Further, the above mapping each pixel value in the density condition image to a density vector respectively can be, but is not limited to, including: mapping each pixel value in the density condition image to a density vector according to the density value corresponding to the target density level. Among them, different density levels correspond to different density values. Specifically, the higher the density level, the larger the corresponding density value.
[0154] For example, assume that the size of the density condition image is 8×8, and assume that the size of the density value corresponding to the target density level is 128. The above method is illustrated and explained by the following steps: Each pixel value in the density condition image is mapped to a density vector according to the density value 128. Then, the respective density vectors are concatenated to obtain a density condition tensor with a size of 128×8×8 (i.e., the density condition feature vector).
[0155] It should be noted that the above-mentioned feature encoding of the road network generation conditions to obtain the road network feature vector further includes: feature encoding the road network style label to obtain the style feature vector. Specifically, the style feature vector matching the road network style label is found in the style feature vector library; or, the road network style label is encoded according to the agreed vector dimension to obtain the style feature vector.
[0156] Optionally, in this embodiment, the above-mentioned road network style label can be used, but is not limited to, uniquely identifying the style features corresponding to a specified city. Specifically, different road network style labels are respectively used to identify different urban styles. It should be noted that the above-mentioned style feature vector library pre-stores a plurality of style feature vectors generated based on the urban style features corresponding to each of a plurality of cities. Among them, each style feature vector in the style feature vector library is associated with its corresponding road network style label.
[0157] For example, assume that the style feature vector library includes style feature vectors corresponding to 49 cities respectively, and among them, the style feature vector corresponding to city A among the above 49 cities is associated with the above-mentioned road network style label. Then, after obtaining the above-mentioned road network style label, the style feature vector corresponding to city A that matches the road network style label can be found in the style feature vector library.
[0158] In the embodiment of the present application, in other words, in the embodiment of the present application, by means of feature encoding the road network generation conditions, the conditional feature vector is obtained. Thereby, the road network image is generated by using the conditional vector, and further, without manual analysis, the road network image incorporating the specified conditional features can be automatically generated, thereby achieving the technical effects of improving the generation efficiency of the road network image of the road system and enhancing the content richness of the road network image.
[0159] Optionally, as an alternative solution, before obtaining the road network generation conditions, it further includes:
[0160] S1. Obtain a sample road network generation condition carrying different types of surface feature distribution information;
[0161] S2. Use the sample road network generation condition to train the decoding network for generating the road network image until the first convergence condition is reached, where the first convergence condition is used to indicate that the difference loss value between the input image and the output image of the decoding network is less than the first predetermined threshold;
[0162] S3. Using the sample road network generation conditions and the decoding network that meets the first convergence condition, jointly train the initialized multi-layer iterative denoising processing network until the second convergence condition is met, so as to obtain the multi-layer iterative denoising processing network, where the second convergence condition is used to indicate that the difference value between the predicted noise and the label noise output by the multi-layer iterative denoising processing network during training is less than the second predetermined threshold.
[0163] Optionally, in this embodiment, the above-mentioned obtaining of the sample road network generation conditions carrying different types of surface feature distribution information may include, but is not limited to: obtaining a sample image set and a sample road network style label set. Specifically, the sample image set and the sample road network style label set can be obtained by, but are not limited to, the following steps: Obtain the road network data of multiple cities from the map database (Open StreetMap, abbreviated as OSM). Then use code to convert the road network data into image data to generate the above-mentioned sample image set, where the sample image set includes a sample urban road network image set, a sample water body image set, a sample main road image set, a sample green space image set, a sample density image set, and so on. At the same time, generate a sample road network style feature vector set and a sample road network style label set respectively according to the road network data of each city in the above-mentioned road network data of multiple cities.
[0164] It should be noted that OSM will include annotations of water body ranges, green space ranges, and urban roads, and each point in OSM is marked with these data using longitude and latitude. In this embodiment, the entire city is divided into small blocks of 4×4 km. In each small block, the longitude and latitude coordinates are converted into pixel coordinates and then drawn into pictures. In this way, the urban road network images, water body images, main road images, and green space images in the above-mentioned sample image set can be obtained, and these images are all binary grayscale images.
[0165] For example, assuming that the above-mentioned decoding network is a VAE network, the above-mentioned training of the decoding network for generating road network images using the sample image set and the sample road network style label until the first convergence condition is met may include, but is not limited to:
[0166] Take each sample urban road network image in the sample urban road network image set in the sample image set as the current sample urban road network image in turn, and execute as Figure 7 shown in the following steps:
[0167] Execute step S702, input the current sample urban road network image into the encoder of the VAE network, so as to use the encoder of the VAE network to encode the current sample urban road network image to obtain the current reference noise prediction noise vector.
[0168] Next, perform step S704, input the current reference noise prediction noise vector into the decoder of the VAE network, and use the decoder of the VAE network to decode the current reference noise prediction noise to obtain the current training urban road network image.
[0169] Then, perform steps S706 - S708 to obtain the difference loss value between the current sample urban road network image and the current training urban road network image. Determine whether the difference loss value between the current sample urban road network image and the current training urban road network image is less than the first predetermined threshold.
[0170] When the difference loss value is not less than the first predetermined threshold, obtain the next sample urban road network image as the current sample urban road network image, and repeat steps S702 - S708. When the difference loss value is less than the first predetermined threshold, perform step S710, determine that the VAE network has reached the third convergence condition, and increment the condition value corresponding to the third convergence condition by 1.
[0171] Next, perform step S712 to determine whether the condition value corresponding to the third convergence condition is greater than the third predetermined threshold. When the condition value corresponding to the third convergence condition is greater than the third predetermined threshold, perform step S714 to determine that the VAE network has reached the first convergence condition and stop training. When the condition value corresponding to the third convergence condition is less than the third predetermined threshold, obtain the next sample urban road network image as the current sample urban road network image, and repeat steps S702 - S708.
[0172] It should be noted that in this embodiment, the above - mentioned difference loss value can be jointly obtained by, but not limited to, using a reconstruction loss function, a generative adversarial network loss function, and a divergence loss function, where the reconstruction loss function includes a deep convolutional neural network loss function and an absolute value loss function.
[0173] Specifically, as shown in formula (14), use the deep convolutional neural network loss function to obtain the first loss value between the current sample urban road network image and the current training urban road network image:
[0174]
[0175] Among them, the above - mentioned Loss1 is used to represent the first loss value, and the above - mentioned L vgg is used to represent the deep convolutional neural network loss function, the above - mentioned x is used to represent the current sample urban road network image, and the above is used to represent the current training urban road network image.
[0176] As shown in formula (15), use the absolute value loss function to obtain the second loss value between the current sample urban road network image and the current training urban road network image:
[0177]
[0178] Among them, the above-mentioned Loss2 is used to represent the second loss value, and the above-mentioned L1 is used to represent the absolute value loss function.
[0179] As shown in formulas (16)-(17), using the first loss value and the second loss value, the third loss value between the current sample urban road network image and the current training urban road network image calculated by the reconstruction loss function is obtained:
[0180]
[0181]
[0182] Among them, the above-mentioned Loss3 is used to represent the third loss value, and the above-mentioned L rec is used to represent the reconstruction loss function, and the above-mentioned α is used to represent the parameter in the reconstruction loss function. The above-mentioned epoch is used to represent the number of rounds of this training. When epoch is greater than 50, the value of α is 0, and when epoch is less than or equal to 50, the value of α is 1.
[0183] As shown in formula (18), using the adversarial network loss function, the fourth loss value between the current sample urban road network image and the current training urban road network image is obtained:
[0184]
[0185] Among them, the above-mentioned Loss4 is used to represent the fourth loss value, and the above-mentioned L kl is used to represent the adversarial network loss function.
[0186] As shown in formula (19), using the divergence loss function, the fifth loss value between the current sample urban road network image and the current training urban road network image is obtained:
[0187]
[0188] Among them, the above-mentioned Loss5 is used to represent the fifth loss value, and the above-mentioned L adv is used to represent the divergence loss function.
[0189] As shown in formula (20), using the above-mentioned third loss value, fourth loss value, and fifth loss value, the difference loss value between the current sample urban road network image and the current training urban road network image is obtained:
[0190]
[0191] Among them, the above Loss is used to represent the difference loss value between the current sample urban road network image and the current training urban road network image.
[0192] As an optional implementation manner, assume that the above multi-layer iterative denoising processing network is a U-Net network. The following steps are used to jointly train the initialized multi-layer iterative denoising processing network with the decoding network that reaches the first convergence condition until the second convergence condition is reached for example and explanation:
[0193] Each label noise included in the label noise set is respectively determined as the current label noise, where the above current label noise is generated by performing multi-layer Gaussian noise addition processing on the label road network image.
[0194] Then, in the sample image set, the current sample road network attribute condition image (i.e., the sample water body image, sample main road image, and sample green space image that match the road network data included in the label road network image), the current sample density condition image (i.e., the sample density image that matches the label road network image) are obtained. And the current sample road network style label that matches the road network data included in the label road network image is obtained in the sample road network style label set.
[0195] Next, the current sample road network attribute condition image, the current sample density condition image, and the current sample road network style label are respectively encoded to obtain the current sample road network attribute condition vector, the current sample density condition vector, and the current sample style feature vector.
[0196] Furthermore, in the U-Net network, as shown in formula (21), using the random reference noise, the current sample road network attribute condition vector, the current sample density condition vector, the current time vector, and the current sample style feature vector, the current predicted noise is obtained:
[0197] ε φ =F φ ((Concat([z t2 ,c cond2 ))+c density2 ,t emb2 +c style2 ) (21)
[0198] Among them, ε φ is used to represent the current predicted noise, z t2 is used to represent the random reference noise, c cond2 is used to represent the current road network attribute condition vector of this network, c density2 is used to represent the current sample density condition vector, t emb2 is used to represent the current time vector, c style2Used to represent the current sample style feature vector.
[0199] Then, the current reference road network hidden vector obtained by removing the above-mentioned current prediction noise from the random reference noise is input into the above-mentioned decoding network that reaches the first convergence condition. And the decoding network is used to decode the above-mentioned current prediction road network hidden vector to obtain the above-mentioned current prediction road network image.
[0200] Next, use the following formula (22) to obtain the difference value between the current prediction road network image and the label road network image corresponding to the above-mentioned current label noise, and use the following formula (23) to calculate the difference value between the current prediction noise and the current label noise:
[0201] Loss1 = ||X1 - X2|| 2 (22)
[0202] Loss ε = ||ε - ε φ (z t2 , t emb2 , c cond2 , c style2 , c density2 )|| 2 (23)
[0203] Wherein, the above-mentioned Loss1 is used to represent the difference value between the current prediction road network image and the label road network image corresponding to the above-mentioned current label noise, the above-mentioned X1 is used to represent the label road network image, and the above-mentioned X2 is used to represent the current prediction road network image. The above-mentioned Loss ε is used to represent the difference value between the current prediction noise and the current label noise, and the above-mentioned ε is used to represent the current label noise.
[0204] Furthermore, when the difference value between the current prediction noise and the current label noise is less than the second threshold, and the difference value between the current prediction road network image and the label road network image corresponding to the above-mentioned current label noise is less than the fourth threshold, it is determined that the U-Net network has reached the fourth convergence condition, and the condition value corresponding to the fourth convergence condition is incremented by 1. Determine whether the condition value corresponding to the fourth convergence condition is greater than the fifth predetermined threshold. When the condition value corresponding to the fourth convergence condition is greater than the fifth predetermined threshold, it is determined that the U-Net network has reached the second convergence condition, and the training is stopped. When the condition value corresponding to the fourth convergence condition is less than or equal to the fifth predetermined threshold, continue to obtain the next label noise in each label noise included in the label noise set as the current label noise, and repeat the above training process.
[0205] When the difference value between the current predicted noise and the current labeled noise is greater than the second threshold, and / or the difference value between the current predicted road network image and the labeled road network image corresponding to the above current labeled noise is greater than the fourth threshold, continue to obtain the next labeled noise from each of the labeled noises included in the labeled noise set as the current labeled noise, and repeat the above training process.
[0206] In the embodiments of the present application, sample road network generation conditions carrying different types of surface feature distribution information are obtained. Then, the decoding network for generating the road network image is trained using the sample road network generation conditions until the first convergence condition is reached, where the first convergence condition is used to indicate that the difference loss value between the input image and the output image of the decoding network is less than the first predetermined threshold. Next, the initialized multi-layer iterative denoising processing network is jointly trained using the sample road network generation conditions and the decoding network that has reached the first convergence condition until the second convergence condition is reached, where the second convergence condition is used to indicate that the difference value between the predicted noise and the labeled noise output by the multi-layer iterative denoising processing network during training is less than the second predetermined threshold. In other words, in the embodiments of the present application, by training the decoding network for generating the road network image and the multi-layer iterative denoising processing network using the sample road network generation conditions, the robustness of the decoding network and the multi-layer iterative denoising processing network is enhanced.
[0207] As an alternative implementation, the application process of the above method is explained by way of example as a whole by the following steps shown in Figure 8 :
[0208] Step S802, prepare training sample data. Specifically, road network data of multiple cities are obtained from the map database (Open Street Map, abbreviated as OSM). Then, the road network data is converted into image data using code to generate the above sample image set, where the sample image set includes a sample urban road network image set, a sample water body image set, a sample main road image set, a sample green space image set, a sample density image set, and so on. At the same time, a sample road network style feature vector and a sample road network style label are generated respectively according to the road network data of each city in the above road network data of multiple cities. It should be noted that the OSM will include the annotation of the water body range, the green space range, and the urban roads, and each point in the OSM is labeled with these data using longitude and latitude. In this embodiment, the entire city is divided into small blocks of 4×4 km, and in each small block, the longitude and latitude coordinates are converted into pixel coordinates and then drawn into pictures. In this way, the urban road network images, water body images, main road images, and green space images in the above sample image set can be obtained, and these images are all binary grayscale images.
[0209] Step S804: Train the VAE network. Specifically, use the sample urban road network images in the sample image set to train the VAE network until the first convergence condition is met.
[0210] Step S806: Train the U-Net network. Specifically, use the sample image set, the sample style label set, and the label noise set to train the initialized U-Net network until the second convergence condition is met.
[0211] Step S808: Obtain a road network attribute condition image for indicating the distribution information of surface ecological features, a density condition image for indicating the density of surface feature distribution, and a road network style label.
[0212] Step S810: Perform feature encoding on the road network attribute condition image for indicating the distribution information of surface ecological features and the density condition image for indicating the density of surface feature distribution to obtain conditional feature vectors, and perform feature encoding on the road network style label to obtain style feature vectors.
[0213] Step S812: In the multi-layer iterative denoising processing network, use the conditional feature vectors and the style feature vectors to perform denoising processing on the reference noise to obtain the target road network latent vector.
[0214] Step S814: Decode the road network latent vector to generate a road network image corresponding to the road network generation condition.
[0215] In the embodiment of the present application, obtain the road network generation condition, where the road network generation condition is used to indicate various different surface feature distribution information referred to when generating the road network image. Then, perform feature encoding on the road network generation condition to obtain the road network feature vector. Next, perform denoising iterative processing on the reference noise based on the road network feature vector to obtain the road network latent vector. Furthermore, decode the road network latent vector to generate a road network image matching the road network generation condition. In other words, adopting the embodiment of the present application, by using the road network feature vector to perform denoising iterative processing on the reference noise to obtain the road network latent vector; and then decoding the road network latent vector to generate a road network image matching the road network generation condition. It achieves the effect of automatically generating a road network image incorporating the road network generation condition specified in advance according to actual needs without the manual participation of a designer in the process of generating the road network image. Thus, it solves the problem of low generation efficiency of the road network image of the road system in the prior art, and realizes the technical effects of improving the generation efficiency of the road network image of the road system and enhancing the content richness of the road network image.
[0216] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should understand that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0217] According to another aspect of the embodiments of the present application, there is also provided a road network image generation device for implementing the above-mentioned road network image generation method. As Figure 9 shown, the device includes:
[0218] An acquisition unit 902, configured to acquire road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image;
[0219] An encoding unit 904, configured to perform feature encoding on the road network generation conditions to obtain a road network feature vector;
[0220] A denoising processing unit 906, configured to perform denoising iterative processing on a reference noise based on the road network feature vector to obtain a road network latent vector;
[0221] A decoding unit 908, configured to decode the road network latent vector to generate a road network image that matches the road network generation conditions.
[0222] Optionally, in this embodiment, the above-mentioned denoising processing unit includes: a denoising processing module, configured to perform denoising processing on the reference noise by using the road network feature vector in a multi-layer iterative denoising processing network to obtain a road network latent vector, where each layer of the multi-layer iterative denoising processing network is used to perform denoising processing on the denoising processing result output by the previous layer of the denoising processing network by using the road network feature vector.
[0223] Optionally, in this embodiment, the above denoising processing module is further configured to perform the following operations in the multi-layer iterative denoising processing network: when the i-th layer denoising processing network is the first layer denoising processing network, in the first layer denoising processing network, use the road network feature vector to predict the noise residual of the reference noise to obtain the first layer predicted noise vector, where i is a natural number greater than or equal to 1 and less than or equal to N, and N is the number of layers of the multi-layer iterative denoising processing network; when the i-th layer denoising processing network is not the first layer denoising processing network, obtain the (i-1)-th layer predicted noise vector output by the (i-1)-th layer denoising processing network before the i-th layer denoising processing network; in the i-th layer denoising processing network, use the road network feature vector and the (i-1)-th layer predicted noise vector to predict the noise residual to obtain the i-th layer predicted noise vector; perform denoising processing based on the i-th layer predicted noise vector to obtain the i-th layer denoising processing result; when the i-th layer denoising processing network is the last layer denoising processing network, use the i-th layer denoising processing result as the road network hidden vector.
[0224] Optionally, in this embodiment, the above denoising processing module is further configured to: determine the processing time vector of the i-th layer denoising processing network; splice the (i-1)-th layer denoising processing result and the conditional feature vector in the road network feature vector to obtain the i-th layer conditional vector, where the conditional feature vector is the vector in the road network feature vector used to indicate the surface environment feature distribution information; perform weighted summation on the style feature vector and the processing time vector in the road network feature vector to obtain the i-th layer style vector, where the style feature vector is the vector in the road network feature vector used to indicate the road network style feature distribution information; perform noise residual prediction based on the i-th layer conditional vector and the i-th layer style vector to obtain the i-th layer predicted noise vector.
[0225] Optionally, in this embodiment, the above denoising processing module is further configured to: when the conditional feature vector includes an attribute conditional feature vector and a density conditional feature vector, splice the (i-1)-th layer denoising processing result and the attribute conditional feature vector to obtain the i-th layer attribute vector; perform weighted summation on the i-th layer attribute vector after adjusting the vector dimension and the density conditional feature vector to obtain the i-th layer conditional vector.
[0226] Optionally, in this embodiment, the above device further includes: a first acquisition unit, configured to acquire a mask image; an update unit, configured to update the road network image area occluded by the mask image in the road network image by using the road network feature vector, the (j-1)-th layer denoising processing result, and the mask vector corresponding to the mask image in the multi-layer iterative denoising processing network.
[0227] Optionally, in this embodiment, the above update unit is configured to perform the following operations in the multi-layer iterative denoising processing network: obtain the denoising result of the (j-1)-th layer output by the denoising processing network of the (j-1)-th layer before the denoising processing network of the j-th layer, where j is a natural number greater than or equal to 1 and less than or equal to N; splice the denoising result of the (j-1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer, and perform weighted summation on the style feature vector and the processing time vector of the denoising processing network of the j-th layer to obtain the style vector of the j-th layer; perform noise residual prediction based on the conditional vector of the j-th layer and the style vector of the j-th layer to obtain the intermediate predicted noise vector of the j-th layer; fuse the intermediate predicted noise vector of the j-th layer with the mask vector to obtain the occlusion region vector corresponding to the road network image region occluded by the masked image; perform noise addition processing on the image region not occluded by the masked image to obtain a noise addition result, and fuse the noise addition result and the mask reference vector corresponding to the mask vector to obtain the non-occluded region vector; perform weighted summation on the occlusion region vector and the non-occluded region vector to obtain the masked predicted noise vector of the j-th layer; perform denoising processing based on the masked predicted noise vector of the j-th layer to obtain the masked denoising processing result of the j-th layer.
[0228] Optionally, in this embodiment, the above encoding unit includes: a binarization processing module, configured to binarize the road network attribute condition image when the road network generation condition includes the road network attribute condition image for indicating the surface environment feature distribution information, to obtain a grayscale image; perform size scaling adjustment on the grayscale image to obtain an attribute condition feature vector, where the conditional feature vector includes the attribute condition feature vector; an adjustment module, configured to map each pixel value in the density condition image to a density vector when the road network generation condition includes the density condition image for indicating the density degree of the surface element distribution; splice the density vectors to obtain a density condition feature vector, where the conditional feature vector includes the density condition feature vector.
[0229] Optionally, in this embodiment, the above device further includes: a second acquisition unit, configured to acquire sample road network generation conditions carrying different types of surface feature distribution information; a first training unit, configured to train the decoding network for generating a road network image by using the sample road network generation conditions until a first convergence condition is reached, where the first convergence condition is used to indicate that the difference loss value between the input image and the output image of the decoding network is less than a first predetermined threshold; a second training unit, configured to jointly train the initialized multi-layer iterative denoising processing network by using the sample road network generation conditions and the decoding network that reaches the first convergence condition until a second convergence condition is reached, to obtain a multi-layer iterative denoising processing network, where the second convergence condition is used to indicate that the difference value between the predicted noise and the label noise output by the multi-layer iterative denoising processing network during training is less than a second predetermined threshold.
[0230] Specific embodiments may refer to the examples shown in the above method for generating road network images, and will not be elaborated herein.
[0231] According to another aspect of the embodiments of the present application, an electronic device for implementing the above method for generating road network images is further provided. In this embodiment, the electronic device is taken as an example of a terminal for illustration. As Figure 10 shown, the electronic device includes a memory 1002 and a processor 1004. A computer program is stored in the memory 1002, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0232] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.
[0233] Optionally, in this embodiment, the above processor may be configured to execute the following steps through the computer program:
[0234] S1, obtaining road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image;
[0235] S2, performing feature encoding on the road network generation conditions to obtain a road network feature vector;
[0236] S3, performing denoising iteration processing on the reference noise based on the road network feature vector to obtain a road network latent vector;
[0237] S4, decoding the road network latent vector to generate a road network image matching the road network generation conditions.
[0238] Optionally, those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic. The electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown Figure 10 in, or have a different configuration from that shown Figure 10 in.
[0239] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for generating road network images in the embodiments of the present application. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, implements the above-mentioned method for generating road network images. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1002 may further include a memory remotely disposed relative to the processor 1004, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. As an example, as Figure 10 shown, the above-mentioned memory 1002 may include, but is not limited to, the acquisition unit 902, encoding unit 904, denoising processing unit 906, and decoding unit 908 in the above-mentioned device for generating road network images. In addition, it may further include, but is not limited to, other module units in the above-mentioned device for generating road network images, which will not be elaborated in this example.
[0240] Optionally, the above-mentioned transmission device 1006 is used to receive or send data via a network. Specific examples of the above network may include a wired network and a wireless network. In one instance, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, thereby enabling communication with the Internet or a local area network. In one instance, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0241] In addition, the above-mentioned electronic device further includes: a connection bus 1008, which is used to connect each module component in the above-mentioned electronic device.
[0242] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a point-to-point network, and any form of computing device, such as a server, a terminal, and other electronic devices, can become a node in the blockchain system by joining the point-to-point network.
[0243] According to one aspect of the present application, there is provided a computer program product, which includes computer programs / instructions containing program codes for performing the above-mentioned method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions provided by the embodiments of the present application are performed.
[0244] According to one aspect of the present application, there is provided a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned method.
[0245] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for performing the following steps:
[0246] S1. Obtain road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image;
[0247] S2. Perform feature encoding on the road network generation conditions to obtain a road network feature vector;
[0248] S3. Perform denoising iteration processing on a reference noise based on the road network feature vector to obtain a road network latent vector;
[0249] S4. Decode the road network latent vector to generate a road network image that matches the road network generation conditions.
[0250] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.
[0251] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods can be completed by instructing the relevant hardware of a terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0252] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0253] In the above embodiments of this application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0254] In the several embodiments provided by this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0255] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0256] In addition, the functional units in the various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0257] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for generating a road network image, characterized in that, Comprising: Obtaining road network generation conditions, where the road network generation conditions are used to indicate various different surface feature distribution information referred to when generating a road network image; Performing feature encoding on the road network generation conditions to obtain a road network feature vector; Performing iterative denoising processing on reference noise based on the road network feature vector to obtain a road network latent vector; Decoding the road network latent vector to generate a road network image that matches the road network generation conditions.
2. The method according to claim 1, wherein The performing iterative denoising processing on reference noise based on the road network feature vector to obtain a road network latent vector includes: In a multi-layer iterative denoising processing network, using the road network feature vector to perform denoising processing on the reference noise to obtain the road network latent vector, where each layer of the denoising processing network in the multi-layer iterative denoising processing network is used to perform denoising processing on the denoising processing result output by the previous layer of the denoising processing network using the road network feature vector.
3. The method according to claim 2, wherein The in the multi-layer iterative denoising processing network, using the road network feature vector to perform denoising processing on the reference noise to obtain the road network latent vector includes: Performing the following operations in the multi-layer iterative denoising processing network: When the i-th layer of the denoising processing network is the first layer of the denoising processing network, in the first layer of the denoising processing network, using the road network feature vector to perform noise residual prediction on the reference noise to obtain a first-layer predicted noise vector, where i is a natural number greater than or equal to 1 and less than or equal to N, and N is the number of layers of the multi-layer iterative denoising processing network; When the i-th layer of the denoising processing network is not the first layer of the denoising processing network, obtaining the (i - 1)-th layer predicted noise vector output by the (i - 1)-th layer of the denoising processing network before the i-th layer of the denoising processing network; in the i-th layer of the denoising processing network, using the road network feature vector and the (i - 1)-th layer predicted noise vector to perform noise residual prediction to obtain the i-th layer predicted noise vector; Performing denoising processing based on the i-th layer predicted noise vector to obtain the i-th layer denoising processing result; When the i-th layer of the denoising processing network is the last layer of the denoising processing network, using the i-th layer denoising processing result as the road network latent vector.
4. The method according to claim 3, wherein Comprising: When obtaining the (i - 1)-th layer predicted noise vector output by the (i - 1)-th layer of the denoising processing network before the i-th layer of the denoising processing network, it further includes: Determining the processing time vector of the i-th layer of the denoising processing network; The in the i-th layer of the denoising processing network, using the road network feature vector and the (i - 1)-th layer predicted noise vector to perform noise residual prediction to obtain the i-th layer predicted noise vector includes: Concatenate the denoising processing result of the (i - 1)-th layer and the conditional feature vector in the road network feature vector to obtain the conditional vector of the i-th layer, where the conditional feature vector is the vector in the road network feature vector for indicating the distribution information of surface environment features; perform weighted summation on the style feature vector in the road network feature vector and the processing time vector to obtain the style vector of the i-th layer, where the style feature vector is the vector in the road network feature vector for indicating the distribution information of road network style features; perform noise residual prediction based on the conditional vector of the i-th layer and the style vector of the i-th layer to obtain the predicted noise vector of the i-th layer.
5. The method according to claim 4, characterized in that The concatenating the denoising processing result of the (i - 1)-th layer and the conditional feature vector in the road network feature vector to obtain the conditional vector of the i-th layer includes: When the conditional feature vector includes an attribute conditional feature vector and a density conditional feature vector, concatenate the denoising processing result of the (i - 1)-th layer and the attribute conditional feature vector to obtain the attribute vector of the i-th layer; Perform weighted summation on the attribute vector of the i-th layer after adjusting the vector dimension and the density conditional feature vector to obtain the conditional vector of the i-th layer.
6. The method according to claim 4, wherein After decoding the road network latent vector to generate a road network image matching the road network generation condition, it further includes: Obtain a mask image; In the multi-layer iterative denoising processing network, use the road network feature vector, the denoising processing result of the (j - 1)-th layer, and the mask vector corresponding to the mask image to update the road network image area occluded by the mask image in the road network image.
7. The method according to claim 6, characterized in that In the multi-layer iterative denoising processing network, using the road network feature vector, the predicted noise vector of the (j - 1)-th layer, and the mask vector corresponding to the mask image to update the road network image area occluded by the mask image in the road network image includes: Perform the following operations in the multi-layer iterative denoising processing network: Obtain the denoising processing result of the (j - 1)-th layer output by the (j - 1)-th layer denoising processing network before the j-th layer denoising processing network, where j is a natural number greater than or equal to 1 and less than or equal to N; In the j-th layer denoising processing network, obtain the denoising processing result output by the following method: Concatenate the denoising processing result of the (j - 1)-th layer and the conditional feature vector to obtain the conditional vector of the j-th layer, and perform weighted summation on the style feature vector and the processing time vector of the j-th layer denoising processing network to obtain the style vector of the j-th layer; Perform noise residual prediction based on the conditional vector of the j-th layer and the style vector of the j-th layer to obtain the intermediate predicted noise vector of the j-th layer; Fuse the intermediate predicted noise vector of the j-th layer with the mask vector to obtain an occlusion area vector corresponding to the road network image area occluded by the mask image; Perform noise addition processing on the image area not occluded by the mask image to obtain a noise addition result, and fuse the noise addition result and the mask reference vector corresponding to the mask vector to obtain an unoccluded area vector; Perform a weighted summation on the occluded region vector and the unoccluded region vector to obtain the j-th layer mask prediction noise vector; Perform denoising processing based on the j-th layer mask prediction noise vector to obtain the j-th layer mask denoising processing result.
8. The method according to any one of claims 1 to 7, characterized in that, The feature encoding of the road network generation condition to obtain the road network feature vector includes: In the case where the road network generation condition includes a road network attribute condition image for indicating the distribution information of surface environment features, perform binarization processing on the road network attribute condition image to obtain a grayscale image; perform size scaling adjustment on the grayscale image to obtain an attribute condition feature vector, where the condition feature vector includes the attribute condition feature vector; In the case where the road network generation condition includes a density condition image for indicating the density of surface element distribution, map each pixel value in the density condition image to a density vector respectively; splice the density vectors to obtain a density condition feature vector, where the condition feature vector includes the density condition feature vector.
9. The method according to any one of claims 1 to 7, characterized in that, Before obtaining the road network generation condition, it further includes: Obtain sample road network generation conditions carrying different types of surface feature distribution information; Use the sample road network generation conditions to train a decoding network for generating road network images until a first convergence condition is reached, where the first convergence condition is used to indicate that the difference loss value between the input image and the output image of the decoding network is less than a first predetermined threshold; Use the sample road network generation conditions and the decoding network that reaches the first convergence condition to jointly train an initialized multi-layer iterative denoising processing network until a second convergence condition is reached to obtain the multi-layer iterative denoising processing network, where the second convergence condition is used to indicate that the difference value between the predicted noise and the label noise output by the multi-layer iterative denoising processing network during training is less than a second predetermined threshold.
10. An apparatus for generating a road network image, characterized in that, It includes: An acquisition unit for acquiring a road network generation condition, where the road network generation condition is used to indicate various different surface feature distribution information referred to when generating a road network image; An encoding unit for performing feature encoding on the road network generation condition to obtain a road network feature vector; A denoising processing unit for performing denoising iterative processing on the reference noise based on the road network feature vector to obtain a road network hidden vector; A decoding unit for decoding the road network hidden vector to generate a road network image matching the road network generation condition.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program, when run by a processor, executes the method described in any one of claims 1 to 9.
12. A computer program product, comprising a computer program / instructions, characterized in that, The computer program / instructions, when executed by a processor, implement the steps of the method described in any one of claims 1 to 9.
13. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 9 through the computer program.