Conditional smoke image generation method and system based on generative adversarial network
By adopting a conditional smoke image generation method based on generative adversarial networks, the problem of insufficient data in fire smoke detection models is solved, and reasonable filling of smoke texture in complex backgrounds is achieved, thereby improving the robustness of fire detection.
Patent Information
- Application Number
- CN202310742261.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-06-21
AI Technical Summary
Existing fire smoke detection models lack large-scale and diverse training data. Generative adversarial networks are unstable during training and generate smoke images that differ greatly from the deep features of real images, making it difficult to effectively extract smoke features and generate smoke images with a specified background.
By subtracting smoke regions to form a smoke-free background image and a smoke contour mask, and using adversarial training of a multi-scale extended fusion generative network and a dual-branch discriminative network, realistic smoke images are generated. Combined with fire numerical simulation to obtain the smoke contour mask, the smoke texture is reasonably filled in a complex background.
The robustness of the smoke detection model in new scenarios has been improved, and the generated smoke images are deeply fused with the features of real images, thus enhancing the detection capability of fire detection.
Smart Images

Figure CN116721324B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of video fire detection and deep learning, and particularly relates to a conditional smoke image generation method and system based on a generative adversarial network. BACKGROUND
[0002] A smoke detection model based on deep learning needs a large amount of image data for training. However, the currently disclosed fire smoke dataset is too small in size and lacks scene diversity, making it difficult to support the research of smoke detection technology based on deep learning. In addition, due to the lack of a unified scene-rich fire smoke test set, the performance of different smoke detection algorithms is difficult to accurately compare. As can be seen, the lack of data is an important factor limiting the development of image smoke detection technology. Although some scholars have tried to use image processing software or three-dimensional modeling software to create synthetic smoke video images, there are differences in deep image features between such synthetic smoke images and real images, and the use of such synthetic smoke images as training data does not significantly improve the performance of the smoke detection model. The generative adversarial network can learn deep image features from real smoke images in adversarial training. However, the generative adversarial network has problems such as unstable training, gradient disappearance, and mode collapse. The smoke in the image has a semi-transparent feature and no fixed shape, and the existing generative adversarial network is not suitable for generating diversified smoke images. Therefore, how to effectively extract and represent the static features of smoke, generate smoke images of a specified background, and construct a large-scale dataset for smoke detection model training has become a problem to be solved. SUMMARY
[0003] To solve the above technical problems, the present application provides a conditional smoke image generation method and system based on a generative adversarial network.
[0004] The technical solution of the present application is as follows: a conditional smoke image generation method based on a generative adversarial network, comprising:
[0005] Step S1: removing the smoke area in the smoke image in the training set to form a one-to-one corresponding smoke-free background image, smoke contour mask, and real smoke image;
[0006] Step S2: superimposing the smoke-free background image and the smoke contour mask and inputting them into a multi-scale expansion fusion generation network to fill the missing smoke area in the smoke-free background image, thereby obtaining a generated smoke image;
[0007] Step S3: inputting the generated smoke image and the corresponding real smoke image into a double-branch discriminant network to obtain various loss values through comparison of local and global features, thereby guiding the adversarial training of the multi-scale expansion fusion generation network and the double-branch discriminant network, and obtaining a trained multi-scale expansion fusion generation network.
[0008] Step S4: fire numerical simulation is carried out for the target scene, the fire initial smoke diffusion image of the target scene is obtained by calculation, the smoke contour mask is obtained through pixel binaryzation processing, the pixels in the mask area in the target scene image are deducted to obtain a background image with a missing area; the background image with the missing area and the smoke contour mask are input into the trained multi-scale expansion fusion generation network, reasonable smoke texture is filled in the background image with the missing area, and a realistic smoke generation image is obtained.
[0009] Compared with the prior art, the present application has the following advantages:
[0010] The application discloses a conditional smoke image generation method based on a generative adversarial network, which can learn the expression of smoke deep features in different backgrounds in game play, and realize reasonable filling of smoke texture in a complex background. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 A flowchart of the conditional smoke image generation method based on the generative adversarial network in the embodiment of the present application is shown in the figure.
[0012] Figure 2 A schematic diagram of a training sample in the embodiment of the present application is shown in the figure.
[0013] Figure 3 A structural schematic diagram of a multi-scale expansion fusion generation network in the embodiment of the present application is shown in the figure.
[0014] Figure 4 A structural schematic diagram of a multi-scale expansion fusion module in the embodiment of the present application is shown in the figure.
[0015] Figure 5 A double-branch discriminant network in the embodiment of the present application is shown in the figure.
[0016] Figure 6 A schematic diagram of a smoke generation image effect in the embodiment of the present application is shown in the figure.
[0017] Figure 7 A structural block diagram of a conditional smoke image generation system based on a generative adversarial network in the embodiment of the present application is shown in the figure. Detailed Implementation
[0018] This invention provides a conditional smoke image generation method based on generative adversarial networks, which solves the problem of insufficient training data in current video fire detection.
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below through specific implementations and in conjunction with the accompanying drawings.
[0020] Example 1
[0021] like Figure 1 As shown in the figure, an embodiment of the present invention provides a conditional smoke image generation method based on generative adversarial networks, which includes the following steps:
[0022] Step S1: Remove the smoke region from the smoke images in the training set to form a one-to-one correspondence of a smoke-free background image, a smoke contour mask, and a real smoke image;
[0023] Step S2: The smoke-free background image and the smoke contour mask are superimposed and then input into the multi-scale extended fusion generation network to fill the missing smoke areas in the smoke-free background image, thus obtaining the generated smoke image.
[0024] Step S3: Input the generated smoke image and the corresponding real smoke image into the dual-branch discriminant network. By comparing the local and global features, the loss values of each feature are obtained, thereby guiding the adversarial training of the multi-scale extended fusion generation network and the dual-branch discriminant network to obtain the trained multi-scale extended fusion generation network.
[0025] Step S4: Perform fire numerical simulation for the target scene. Calculate the initial smoke diffusion image of the target scene and obtain a smoke contour mask through pixel binarization. Subtract pixels in the mask area from the target scene image to obtain a background image with missing regions. Input the background image with missing regions and the smoke contour mask into a trained multi-scale extended fusion generation network to fill the background image with missing regions with reasonable smoke textures, thus obtaining a realistic smoke generation image.
[0026] In one embodiment, step S1 above, which involves subtracting smoke regions from smoke images in the training set to form a one-to-one correspondence of a smoke-free background image, a smoke contour mask, and a real smoke image, specifically includes:
[0027] Step S11: Collect real smoke images under various scenarios through ignition experiments;
[0028] The selected experimental sites in the examples of the present application include Huangshan, Daxing'anling, Zifengshan, countryside, campus, street, standard combustion chamber, etc. Experiments are carried out using various common combustibles as fuel, including cotton rope, smoke bomb, beech, straw, diesel oil, gasoline, n-heptane, etc. When conducting ignition experiments in indoor scenes, a network camera and a mobile phone are used for shooting at the same time. When conducting ignition experiments in outdoor scenes, a drone is used for long-distance shooting, and a mobile phone is used for close-range shooting. After screening, a total of 200 real smoke videos and 40 smoke-free background videos are obtained.
[0029] Step S12: The real smoke image is segmented by artificial recognition to obtain a marked image;
[0030] The smoke image generation network designed in the present application requires three kinds of data, namely smoke-free background image, smoke contour mask and real smoke image, for network training, and the three kinds of data are required to be one-to-one corresponding. Since matched samples are difficult to obtain, the matched smoke-free background image and smoke contour mask are obtained by processing the real smoke image. The real smoke image is segmented and marked, and the marking principle is to segment all the smoke areas that can be recognized by artificial recognition, so as to ensure that there is no recognizable smoke outside the segmented area.
[0031] Step S13: The marked image is binarized to obtain a corresponding smoke contour mask, and the smoke-free background image is obtained by subtracting the smoke contour mask corresponding area from the real smoke image.
[0032] As shown in the schematic diagram of part of the training sample. Figure 2
[0033] In one embodiment, the step S2 described above: after superimposing the smoke-free background image and the smoke contour mask, inputting the multi-scale expansion fusion generation network, filling the missing smoke area in the smoke-free background image to obtain the generated smoke image, specifically comprising:
[0034] Step S21: constructing a multi-scale expansion fusion generation network comprising an encoder, multiple deep feature extractors and a decoder;
[0035] As shown in the structure schematic diagram of the multi-scale expansion fusion generation network. Figure 3
[0036] Step S22: after superimposing the smoke-free background image and the smoke contour mask, inputting the multi-scale expansion fusion generation network, first passing through the four convolutional layers of the encoder, and performing twice down-sampling to obtain a medium-scale feature map and a small-scale feature map;
[0037] As shown in the structure schematic diagram of the multi-scale expansion fusion generation network. Figure 3 The shown encoder is composed of four convolutional layers for twice down-sampling of the input smoke-free background image and smoke contour mask overlay image. Batch normalization layers are set after the first convolutional layer. After the second convolutional layer, a mesoscale feature map is obtained, with a size of 128x128x128. After the fourth convolutional layer, a small-scale feature map is obtained, with a size of 64x64x256.
[0038] Step S23: input the small-scale feature map into a first deep feature extractor composed of 8 multi-scale expansion fusion modules in series, the multi-scale expansion fusion module uses four convolution kernels with different dilation rates to extract features at the same time, and adopts addition and splicing to fuse features, and outputs a small-scale global feature map;
[0039] For the smoke image generation task, appropriate features need to be extracted from the background outside the contour to generate reasonable image content, so the receptive field of the convolution kernel used to extract features should be large enough. However, directly increasing the size of the convolution kernel will result in too many model parameters and difficult training. Dilated convolution can increase the receptive field without changing the number of convolution kernel parameters, but this method is realized by constructing a sparse convolution kernel, so many pixels will be skipped when extracting features, causing information loss. A multi-scale expansion fusion module is designed in the present application to solve this problem.
[0040] The small-scale feature map is input into a first deep feature extractor composed of 8 multi-scale expansion fusion modules in series, and the extracted global features can be used to coarsely fill the hollow area, and output a small-scale global feature map.
[0041] As Figure 4 The structure diagram of a multi-scale expansion fusion module is shown, the input feature map is extracted by using four convolution kernels with different dilation rates, which can extract image features under different receptive fields and fully utilize the input information. For the feature maps extracted by the four different convolution kernels, addition and splicing are used for feature fusion.
[0042] Step S24: after one time up-sampling of the small-scale global feature map, the small-scale global feature map is overlaid with the mesoscale feature map and input into a second deep feature extractor, wherein the second deep feature extractor is composed of 4 multi-scale expansion fusion modules in series, and outputs a mesoscale local feature map.
[0043] In order to more fully extract local features to generate a reasonable translucent smoke area, the medium-scale feature map output by the second convolutional layer of the encoder is spliced with the feature map obtained by upsampling the small-scale local feature map once, and the size of the spliced fusion feature map is 128x128x128. The fusion feature map contains rich local background features. The fusion feature map is input into the second deep feature extractor composed of four multi-scale expansion convolution fusion modules, and the extracted local features are used to fill the hollow area more finely, and a medium-scale local feature map is output.
[0044] Step S25: The medium-scale local feature map is input into the two convolutional layers of the decoder, and after being upsampled once and superimposed with the smoke-free background image, a generated smoke image is obtained.
[0045] In one embodiment, the step S3 of inputting the generated smoke image and the corresponding real smoke image into the double-branch discriminant network to obtain each loss value by comparing the local and global features, thereby guiding the adversarial training of the multi-scale expansion fusion generation network and the double-branch discriminant network, and obtaining the trained multi-scale expansion fusion generation network, specifically includes:
[0046] Step S31: The double-branch discriminant network includes two branches with the same structure, each branch being composed of six convolutional layers in series; wherein the first branch processes the complete input image to obtain a global feature map to judge the rationality of the input image; the second branch only processes the smoke area in the input image to obtain a local feature map to judge the authenticity of the smoke detail texture; finally, the global feature map and the local feature map are connected and input into a classifier to obtain a discriminant result of the input image.
[0047] As shown in Figure 5 , it is a structural diagram of the double-branch discriminant network. The lower branch is the first branch, and the upper branch is the second branch. After the two branches are down-sampled by six convolutional layers, a vector with a length of 512 is obtained.
[0048] Step S32: The generated smoke image and the corresponding real smoke image are input into the double-branch discriminant network to obtain the corresponding local feature map and the discriminant result.
[0049] Step S33: The discriminant results of the generated smoke image and the real smoke image are input into the adversarial loss function to obtain an adversarial loss value to guide the training of the multi-scale expansion fusion generation network and the double-branch discriminant network, wherein the adversarial loss function of the multi-scale expansion fusion generation network and the double-branch discriminant network is as follows.
[0050]
[0051]
[0052] wherein, is the adversarial loss value of the multi-scale expansion fusion generation network, is the adversarial loss value of the double-branch discriminant network, r is the pixel value of the real smoke image, f is the pixel value of the generated smoke image, is the average value, relative monitor D Ra (I r ,I f ) and D Ra (I r ,I r ) are as follows:
[0053]
[0054]
[0055] wherein, sigma (·) is a sigmoid function, and C (·) is a discriminant result output by the double-branch discriminant network;
[0056] The loss function of the original GAN takes 0 and 1 as labels, and does not directly optimize the distance between the data distributions of real and false images in high-dimensional space as the target. In the case that the two data distributions do not completely coincide in high-dimensional space, it is still possible to meet the requirements of the loss function under the mapping of a certain dimension, so the training is unstable. The loss function in the present application optimizes the data distribution difference of real and false images in high-dimensional space as the target.
[0057] Step S34: comparing the generated smoke image with the local features of the real smoke image to obtain a feature loss to guide the texture detail completion performance of the multi-scale expansion fusion generation network, and the feature loss function is as follows:
[0058]
[0059]
[0060] wherein, is the feature loss value of the multi-scale expansion fusion generation network, l is the feature scale number, and in the embodiment of the present application, 5 scales are compared, N l is the element number of the feature map of each scale, is the local feature map of each scale on the second branch of the double-branch discriminant network; is the local feature of the real smoke image;
[0061] In order to ensure that the generated smoke image is consistent with the real smoke image in the high-dimensional space feature distribution, the feature loss under each scale is obtained by comparing the feature maps of each scale on the local feature extraction branch of the discrimination network.
[0062] Step S35: according to and The parameters of the multi-scale expansion fusion generation network and the double-branch discrimination network are updated until the double-branch discrimination network cannot distinguish between the real smoke image and the generated smoke image.
[0063] In the embodiment of the application, the iteration number is set to 100000, the batch size is set to 8, and the initial learning rate of the generation network and the discrimination network is 0.0002, and the learning rate is reduced by half every 20000 iterations.
[0064] In one embodiment, the above step S4: fire numerical simulation is carried out for the target scene, the fire initial smoke diffusion image of the target scene is obtained by calculation, the smoke contour mask is obtained through pixel binaryzation processing, the pixels in the mask area of the target scene image are deducted, and the background image with missing area is obtained; the background image with missing area and the smoke contour mask are input into the trained multi-scale expansion fusion generation network, the reasonable smoke texture is filled in the background image with missing area, and the realistic smoke generation image is obtained.
[0065] In order to obtain a large number of real and reliable smoke contour masks, the CFD simulation method is used to simulate the smoke diffusion of the early fire, and the numerical simulation software FDS is selected. The software can accurately calculate the smoke diffusion results of the target scene at each time within a period of time based on fluid mechanics. The factors affecting smoke diffusion mainly include the combustion characteristics of the fire source, the size of the fire source, the scene and the environmental wind. For the combustion material category, the common wood is selected as the fire source in the embodiment of the present application, and the combustion characteristics of oak wood are referred to. The mass ratio of carbon element, hydrogen element and oxygen element of the fuel is set to 1:1.7:0.72, the critical temperature of the flame is set to 1427℃, the smoke production is set to 0.015Ys, and the CO production is set to 0.004YCO. The size of the fire combustion surface is set to 0.2m*0.2m, and the unit area heat release rate is set to 500KW / m2. For the outdoor scene, the smoke of the fire in the open area is little affected by the surrounding buildings and trees in the early stage, so no obstacles are set in the present application, and the calculation area is set to 20m*20m*20m, and the boundaries on the four sides and the top are set to open surfaces. For the outdoor environmental wind speed, 0m / s, 0.5m / s and 1m / s are selected for simulation. For the indoor scene, the diffusion of the fire smoke is mainly affected by the building structure and the ventilation condition, and the standard combustion chamber of the laboratory is taken as an example for numerical simulation of the early fire smoke diffusion. The length of the room is 10m, the width is 7m, and the height is 4m. Two windows are installed side by side on a 7m*4m wall, the window size is 2m*2m, the distance between the two windows is 1.5m, and the window bottom is 1.5m away from the ground. In the case of closed combustion chamber entrance, the two windows are the only natural ventilation means. When simulating indoor fire, no environmental wind speed is considered, and only two cases of window closing and opening are considered. The FDS software can visually display the smoke diffusion results through three-dimensional animation. The texture details of the smoke in the numerical simulation visualization results are quite different from the real smoke, but the shape of the smoke is very close to the real situation. The smoke diffusion visualization results under a certain view angle can be processed to produce a reasonable smoke contour mask suitable for the target scene.
[0066] First, the smoke diffusion visualization results under a certain view angle at each time are output as image data V, and the three channels of the image are superimposed to obtain a single-channel image M. If the pixel value at (i, j) in M is not 0, it is set to 1, and finally a binary smoke contour mask is obtained. bg The input I of the multi-scale extended fusion generation network is obtained from the target scene background image I and the smoke contour mask M in At this time, in I bg =I (1-M).
[0067] I inInput the trained multi-scale expansion fusion generation network to obtain a generated smoke image, multiply the generated smoke image with a smoke contour mask, and add the generated smoke image and a target scene background image to obtain a final output, Figure 6 As shown in the effect diagram of the smoke generated image in the embodiment of the application.
[0068] The application discloses a conditional smoke image generation method based on a generative adversarial network, which can learn the expression of smoke deep features in different backgrounds in game play, and realize reasonable filling of smoke texture in a complex background.
[0069] Embodiment two
[0070] As Figure 7 shown, the embodiment of the application provides a conditional smoke image generation system based on a generative adversarial network, which comprises the following modules:
[0071] The preprocessing module 51 is used for deducting the smoke area in the smoke image in the training set to form a one-to-one corresponding smoke-free background image, a smoke contour mask, and a real smoke image;
[0072] The generated smoke image is input into the multi-scale expansion fusion generation network after being superimposed with the smoke contour mask, and the missing smoke area in the smoke-free background image is filled to obtain the generated smoke image;
[0073] The generated smoke image and the corresponding real smoke image are input into the double-branch discriminant network, and each loss value is obtained through comparison of local and global features, so as to guide the adversarial training of the multi-scale expansion fusion generation network and the double-branch discriminant network, and obtain the trained multi-scale expansion fusion generation network;
[0074] The smoke generation image module 54 is used for fire numerical simulation for a target scene, and a fire initial smoke diffusion image of the target scene is obtained by calculation. After pixel binaryzation processing, a smoke contour mask is obtained. The pixels in the mask area in the target scene image are deducted to obtain a background image with a missing area. The background image with the missing area and the smoke contour mask are input into the trained multi-scale expansion fusion generation network to fill reasonable smoke texture in the background image with the missing area, and a realistic smoke generation image is obtained.
[0075] The above embodiments are provided only for the purpose of describing the present application, and are not intended to limit the scope of the present application. The scope of the present application is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present application shall be encompassed within the scope of the present application.
Claims
1. A conditional smoke image generation method based on generative adversarial networks, characterized in that, include: Step S1: Remove the smoke region from the smoke images in the training set to form a one-to-one correspondence of a smoke-free background image, a smoke contour mask, and a real smoke image; Step S2: The smoke-free background image and the smoke contour mask are superimposed and then input into a multi-scale extended fusion generation network to fill in the missing smoke regions in the smoke-free background image, resulting in a generated smoke image. Specifically, this includes: Step S21: Constructing the multi-scale extended fusion generative network includes: an encoder, multiple deep feature extractors, and a decoder; Step S22: The smoke-free background image and the smoke contour mask are superimposed and then input into the multi-scale extended fusion generation network. First, the image passes through the four convolutional layers of the encoder and is downsampled twice to obtain the medium-scale feature map and the small-scale feature map respectively. Step S23: Input the small-scale feature map into the first deep feature extractor. The first deep feature extractor is composed of 8 multi-scale expansion and fusion modules connected in series. The multi-scale expansion and fusion module uses four convolutional kernels with different dilation rates to extract features simultaneously and uses two methods, addition and concatenation, to perform feature fusion and output a small-scale global feature map. Step S24: After upsampling the small-scale global feature map once, it is superimposed with the mesoscale feature map and then input into the second deep feature extractor. The second deep feature extractor is composed of four multi-scale expansion fusion modules connected in series, and outputs a mesoscale local feature map. Step S25: The mesoscale local feature map is passed through two convolutional layers of the decoder, upsampled once, and then superimposed on the smoke-free background image to obtain the generated smoke image; Step S3: Input the generated smoke image and the corresponding real smoke image into the dual-branch discriminant network. Obtain various loss values by comparing local and global features, thereby guiding the adversarial training between the multi-scale extended fusion generation network and the dual-branch discriminant network to obtain a trained multi-scale extended fusion generation network. Specifically, this includes: Step S31: Construct a dual-branch discriminant network comprising two branches with identical structures, each consisting of six concatenated convolutional layers; wherein, the first branch processes the complete input image to obtain a global feature map to determine the rationality of the input image; the second branch processes only the smoke region in the input image to obtain a local feature map to determine the authenticity of the smoke detail texture; finally, the global feature map and the local feature map are concatenated and input into the classifier to obtain the discrimination result of the input image; Step S32: Input the generated smoke image and its corresponding real smoke image into the dual-branch discrimination network to obtain the corresponding local feature map and discrimination result; Step S33: Input the discrimination result of the generated smoke image and the real smoke image into the adversarial loss function to obtain the adversarial loss value to guide the training of the multi-scale extended fusion generation network and the dual-branch discriminant network. The adversarial loss function of the multi-scale extended fusion generation network and the dual-branch discriminant network is as follows. (1) (2) in, The adversarial loss value of the multi-scale extended fusion generation network is given. The adversarial loss value of the dual-branch discrimination network is... These are the pixel values of a real smoke image. The pixel values of the generated smoke image. To calculate the average value, relative to the monitor and The specific structure is as follows: (3) (4) in, For the sigmoid function, The discrimination result output by the dual-branch discrimination network; Step S34: Compare the local features of the generated smoke image with those of the real smoke image to obtain feature loss, which guides the texture detail completion performance of the multi-scale extended fusion generation network. The feature loss function is shown below. (5) (6) in, Here, l represents the feature loss value of the multi-scale extended fusion generation network, and l is the feature scale index. This represents the number of elements in the feature map at each scale. These are local feature maps at various scales on the second branch of the dual-branch discriminant network. These are local features of the actual smoke image; Step S35: According to , and Update the parameters of the multi-scale extended fusion generation network and the dual-branch discrimination network until the dual-branch discrimination network can no longer distinguish between real smoke images and generated smoke images; Step S4: Perform fire numerical simulation for the target scene, calculate the initial smoke diffusion image of the target scene, obtain the smoke contour mask through pixel binarization, subtract the pixels in the mask area of the target scene image to obtain a background image with missing areas; input the background image with missing areas and the smoke contour mask into the trained multi-scale extended fusion generation network, fill the background image with missing areas with reasonable smoke texture, and obtain a realistic smoke generation image.
2. The conditional smoke image generation method based on generative adversarial networks according to claim 1, characterized in that, Step S1: Subtracting the smoke region from the smoke images in the training set to form a one-to-one corresponding smoke-free background image, smoke contour mask, and real smoke image, specifically including: Step S11: Collect real smoke images under various scenarios through ignition experiments; Step S12: The smoke area in the real smoke image is completely segmented by manual identification to obtain a marked image; Step S13: Binarize the marked image to obtain the corresponding smoke contour mask, and subtract the corresponding area of the smoke contour mask from the real smoke image to obtain a smoke-free background image.
3. A conditional smoke image generation system based on generative adversarial networks, characterized in that, Includes the following modules: The preprocessing module is used to remove the smoke region from the smoke images in the training set to form a one-to-one smoke-free background image, a smoke contour mask, and a real smoke image. A generative network module is constructed to overlay the smoke-free background image and the smoke contour mask, and then input the result into a multi-scale extended fusion generative network to fill in the missing smoke regions in the smoke-free background image, thereby obtaining a generated smoke image. Specifically, this includes: Step S21: Constructing the multi-scale extended fusion generative network includes: an encoder, multiple deep feature extractors, and a decoder; Step S22: The smoke-free background image and the smoke contour mask are superimposed and then input into the multi-scale extended fusion generation network. First, the image passes through the four convolutional layers of the encoder and is downsampled twice to obtain the medium-scale feature map and the small-scale feature map respectively. Step S23: Input the small-scale feature map into the first deep feature extractor. The first deep feature extractor is composed of 8 multi-scale expansion and fusion modules connected in series. The multi-scale expansion and fusion module uses four convolutional kernels with different dilation rates to extract features simultaneously and uses two methods, addition and concatenation, to perform feature fusion and output a small-scale global feature map. Step S24: After upsampling the small-scale global feature map once, it is superimposed with the mesoscale feature map and then input into the second deep feature extractor. The second deep feature extractor is composed of four multi-scale expansion fusion modules connected in series, and outputs a mesoscale local feature map. Step S25: The mesoscale local feature map is passed through two convolutional layers of the decoder, upsampled once, and then superimposed on the smoke-free background image to obtain the generated smoke image; A generative adversarial training module is used to input the generated smoke image and the corresponding real smoke image into a dual-branch discriminant network. By comparing local and global features, various loss values are obtained, thereby guiding the adversarial training between the multi-scale extended fusion generation network and the dual-branch discriminant network, resulting in a trained multi-scale extended fusion generation network. Specifically, this includes: Step S31: Construct a dual-branch discriminant network comprising two branches with identical structures, each consisting of six concatenated convolutional layers; wherein, the first branch processes the complete input image to obtain a global feature map to determine the rationality of the input image; the second branch processes only the smoke region in the input image to obtain a local feature map to determine the authenticity of the smoke detail texture; finally, the global feature map and the local feature map are concatenated and input into the classifier to obtain the discrimination result of the input image; Step S32: Input the generated smoke image and its corresponding real smoke image into the dual-branch discrimination network to obtain the corresponding local feature map and discrimination result; Step S33: Input the discrimination result of the generated smoke image and the real smoke image into the adversarial loss function to obtain the adversarial loss value to guide the training of the multi-scale extended fusion generation network and the dual-branch discriminant network. The adversarial loss function of the multi-scale extended fusion generation network and the dual-branch discriminant network is as follows. (1) (2) in, The adversarial loss value of the multi-scale extended fusion generation network is given. The adversarial loss value of the dual-branch discrimination network is... These are the pixel values of a real smoke image. The pixel values of the generated smoke image. To calculate the average value, relative to the monitor and The specific structure is as follows: (3) (4) in, For the sigmoid function, The discrimination result output by the dual-branch discrimination network; Step S34: Compare the local features of the generated smoke image with those of the real smoke image to obtain feature loss, which guides the texture detail completion performance of the multi-scale extended fusion generation network. The feature loss function is shown below. (5) (6) in, Here, l represents the feature loss value of the multi-scale extended fusion generation network, and l is the feature scale index. This represents the number of elements in the feature map at each scale. These are local feature maps at various scales on the second branch of the dual-branch discriminant network. These are local features of the actual smoke image; Step S35: According to , and Update the parameters of the multi-scale extended fusion generation network and the dual-branch discrimination network until the dual-branch discrimination network can no longer distinguish between real smoke images and generated smoke images; The smoke generation image module is used to perform fire numerical simulation for a target scene. It calculates the initial smoke diffusion image of the target scene, obtains a smoke contour mask through pixel binarization, subtracts pixels in the masked area from the target scene image to obtain a background image with missing regions, and inputs the background image with missing regions and the smoke contour mask into the trained multi-scale extended fusion generation network. The network fills the background image with missing regions with reasonable smoke texture to obtain a realistic smoke generation image.