Image sequence generation method and system for railway pedestrian intrusion based on posture migration
By acquiring pedestrian pose sequences in non-railway scenarios and using a pose transfer generation model to generate pedestrian intrusion images in railway scenarios, the problem of scarce pedestrian intrusion images in railway scenarios is solved. The generated images have high appearance and pose realism, and automatic annotation saves manpower.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2026-03-31
AI Technical Summary
The scarcity and difficulty in obtaining images of pedestrian intrusions in railway scenarios result in low recognition accuracy of deep learning models, while manual annotation is labor-intensive and may be inaccurate.
By acquiring pedestrian pose sequences in non-railway scenes, a pose transfer pedestrian generation model is established, including a generator and a discriminator, to generate pedestrian intrusion images in railway scenes. Pose transfer technology is used to generate pedestrian intrusion image sequences in railway scenes, and automatic annotation is performed.
The generated images of pedestrian intrusion in railway scenes have clear appearance and texture, distinct postures, and high realism, solving the problem of scarce pedestrian intrusion images and saving on manual annotation costs.
Smart Images

Figure CN115376064B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of railway operation safety detection technology, and in particular to an image sequence generation method and system for railway pedestrian intrusion based on posture transfer. Background Technology
[0002] With the continuous increase in railway train speeds, the intrusion of foreign objects into railway clearance gauges is a major factor contributing to frequent railway safety accidents, posing a significant threat to railway operational safety. Currently, illegal pedestrian intrusion into railway tracks remains a leading cause of injuries and fatalities in railway accidents. Pedestrians illegally crossing tracks seriously threaten railway transportation safety and cause enormous losses to the national economy and people's property.
[0003] Current methods for detecting foreign objects in railway clearance include both contact and non-contact methods. Common contact methods include dual-grid monitoring technology and fiber optic grating technology. Among these, the protective grid technology detects the presence of foreign objects by checking for impacts to the grid. It is accurate in detecting sudden foreign objects such as falling rocks, but it cannot detect other intrusive foreign objects. Furthermore, protective grids require large-scale installation and regular maintenance, which cannot meet the diverse detection requirements of foreign object intrusions into railway clearances. Non-contact detection methods include infrared, millimeter-wave radar, laser, and video analysis. Among these, intrusion target identification and detection methods based on video analysis provide relatively intuitive detection results and are therefore widely used in railway safety systems.
[0004] With the development of deep learning technology, deep learning-based methods for detecting intrusions have gradually matured. Video analytics-based intrusion target identification and detection methods utilize deep learning models to determine the category of intrusion targets with high accuracy. However, the accuracy of deep learning identification models largely depends on the richness and quantity of the intrusion image samples used for model training; a larger sample size generally results in higher identification accuracy.
[0005] Because high-speed railways operate in a closed environment, it is extremely difficult to obtain intrusion target image samples in high-speed rail scenarios. Even if such images are obtained, they represent only a small number of intrusion images within railway scenarios, and the intrusion scenarios are not diverse enough. Furthermore, constructing annotated datasets for deep learning requires extensive manual image annotation work, consuming a significant amount of manpower, and is also susceptible to inaccuracies due to human error. Summary of the Invention
[0006] The embodiments of the present invention provide a method and system for generating image sequences of railway pedestrian intrusion based on pose transfer, so as to overcome the defects of the prior art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution.
[0008] On one hand, the present invention provides a method for generating image sequences of pedestrian intrusion into railways based on pose transfer, comprising:
[0009] Obtain pedestrian pose sequences from non-railway scenes as target pose sequences;
[0010] A model is built and trained to obtain a pose transfer pedestrian generation model. The pose transfer pedestrian generation model includes a generator and a discriminator. The discriminator includes an appearance discriminator and a pose discriminator.
[0011] The predetermined pedestrian appearance image and target pose sequence are input into the pose transfer pedestrian generation model to obtain several target pedestrian images; and
[0012] The target pedestrian image is composited into the railway scene image to generate a railway pedestrian intrusion image.
[0013] Optionally, the method further includes:
[0014] Several images of pedestrian intrusions into railway lines are used to generate a video sequence of pedestrian intrusions into railway scenes, based on the target pose sequence.
[0015] Optionally, the method further includes:
[0016] Automatically annotate images of pedestrian intrusions into railway scenes to generate annotated images of pedestrian intrusions into railway scenes.
[0017] Optionally, after synthesizing the target pedestrian image into the railway scene image, the position coordinates of the pedestrians and the height and width data of the pedestrian bounding boxes in the railway pedestrian intrusion image are recorded to obtain an annotated image of the railway scene pedestrian intrusion.
[0018] Optionally, pedestrian pose data can be extracted from public datasets, and the corresponding pedestrian images and the extracted pedestrian pose data can be used as a dataset to train the model, thereby obtaining a pose transfer pedestrian generation model with optimal parameters.
[0019] Optionally, the training methods include:
[0020] Extract pedestrian image samples and corresponding pedestrian poses from the training set;
[0021] The extracted pedestrian image samples and the corresponding pedestrian poses are input into the generator. Based on the pre-set target pose and the appearance of the pedestrian image samples, a pedestrian pose transfer image with the input pedestrian image appearance and target pose is generated.
[0022] The discriminator distinguishes pedestrian pose transfer images; and
[0023] When the predetermined loss value is reached, the pedestrian generation model with optimal posture migration parameters is obtained.
[0024] Optionally, the appearance discriminator is used to determine whether the appearance of the generated pedestrian image is consistent with the appearance of the pedestrian in the input pedestrian image, and the pose discriminator is used to determine whether the generated pedestrian pose is consistent with the target pose.
[0025] The structure of the appearance discrimination device includes:
[0026] The concatenation layer is used to concatenate the input pedestrian pose transfer image and the target pedestrian pose, as well as the target image and the target pedestrian pose, to obtain a concatenation vector.
[0027] Multiple downsampling modules are used to downsample the concatenated vector to obtain the downsampled feature vector; and
[0028] Multiple residual modules are used to extract features from the downsampled feature vector to obtain discriminative features.
[0029] Among them, the appearance discriminator completes the discrimination based on the discriminative features.
[0030] Secondly, the present invention also provides an image sequence generation system for railway pedestrian intrusion based on pose transfer, comprising:
[0031] The pose extraction module is used to obtain pedestrian pose sequences in non-railway scenes, obtain target pose sequences, extract pedestrian pose data from the Deep Fashion public dataset, and generate a dataset from pedestrian images corresponding to the pedestrian pose data.
[0032] The pose transfer pedestrian generation module is used to build a model and train it using a dataset of pedestrian images corresponding to pedestrian pose data to obtain a pose transfer pedestrian generation model with optimal parameters. It inputs a predetermined pedestrian appearance image and a target pose sequence into the pose transfer pedestrian generation model to obtain several target pedestrian images; and
[0033] The railway scene intrusion pedestrian image synthesis module is used to insert the target pedestrian image into the railway scene image to generate a railway pedestrian intrusion image.
[0034] Optionally, the system also includes:
[0035] The railway scene intrusion video generation module is used to generate a railway scene intrusion video sequence by combining several railway scene intrusion pedestrian image synthesis modules according to the target pose sequence.
[0036] Optionally, the system also includes:
[0037] The automatic pedestrian intrusion image annotation module is used to automatically annotate images of pedestrian intrusions into railway scenes, generating annotated images of pedestrian intrusions into railway scenes.
[0038] The beneficial effects of this invention are as follows: By acquiring pedestrian posture sequences in non-railway scenes, a target posture sequence is obtained. Then, a model is established and trained to obtain a posture transfer pedestrian generation model. After that, a predetermined pedestrian appearance image and the target posture sequence are input into the posture transfer pedestrian generation model to obtain several target pedestrian images. Then, the target pedestrian images are inserted into railway scene images to generate railway pedestrian intrusion images. This obtains a rich sequence of railway scene pedestrian intrusion images with both pedestrian appearance and railway scene. In the generated railway pedestrian intrusion image data, the pedestrian appearance texture is clear, the posture is distinct, and the realism is high, which solves the problem of scarce and difficult-to-obtain railway scene pedestrian intrusion images.
[0039] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A structural block diagram of an image sequence generation system for railway pedestrian intrusion based on pose transfer, provided in an embodiment of the present invention;
[0042] Figure 2 A flowchart illustrating the image sequence generation method for railway pedestrian intrusion based on pose transfer provided in an embodiment of the present invention;
[0043] Figure 3 This is a block diagram of the generator structure provided in an embodiment of the present invention;
[0044] Figure 4 This is a schematic diagram of the spatial attention mechanism feature fusion module structure provided in an embodiment of the present invention;
[0045] Figure 5 This is a framework diagram of the pedestrian pose extraction module provided in an embodiment of the present invention;
[0046] Figure 6 The pedestrian image and the corresponding extracted pose sequence diagram provided in the embodiments of the present invention;
[0047] Figure 7 A schematic diagram of the framework for the pose transfer pedestrian generation process provided in an embodiment of the present invention;
[0048] Figure 8 This is a schematic diagram of the appearance discriminator provided in an embodiment of the present invention;
[0049] Figure 9 A comparison chart showing the image generation results of the GSGAN algorithm and PATN algorithm provided in the embodiments of the present invention;
[0050] Figure 10 This is a rendering of the generated railway pedestrian intrusion image sequence provided by an embodiment of the present invention;
[0051] Figure 11 This is a schematic diagram of annotated images of pedestrian intrusion in a railway scene provided in an embodiment of the present invention. Detailed Implementation
[0052] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0053] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.
[0054] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0055] Terminology Explanation:
[0056] Railway scenario: This refers to a scenario where, under closed operation, a high-speed train is passing through or may pass through, making it impossible to directly capture images of pedestrian intrusion using cameras.
[0057] Non-railway scenes: These can be any road or open area, where a camera can be used to capture video sequences of a pedestrian's various poses.
[0058] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0059] Example 1
[0060] like Figure 1 As shown, this embodiment 1 provides an image sequence generation system for railway pedestrian intrusion based on pose transfer, including:
[0061] The pose extraction module is used to obtain pedestrian pose sequences in non-railway scenes, obtain target pose sequences, extract pedestrian pose data from the Deep Fashion public dataset, and generate a dataset from pedestrian images corresponding to the pedestrian pose data.
[0062] The pose transfer pedestrian generation module is used to build a model and train it using a dataset of pedestrian images corresponding to pedestrian pose data to obtain a pose transfer pedestrian generation model with optimal parameters. It inputs a predetermined pedestrian appearance image and a target pose sequence into the pose transfer pedestrian generation model to obtain several target pedestrian images; and
[0063] The railway scene intrusion pedestrian image synthesis module is used to insert the target pedestrian image into the railway scene image to generate a railway pedestrian intrusion image.
[0064] In this embodiment 1, the above-mentioned image sequence generation system for railway pedestrian intrusion based on pose transfer is used to realize the image sequence generation method for railway pedestrian intrusion based on pose transfer, including:
[0065] On one hand, the present invention provides a method for generating image sequences of pedestrian intrusion into railways based on pose transfer, comprising:
[0066] Obtain the pedestrian pose sequence in non-railway scenes to obtain the target pose sequence;
[0067] A model is built and trained to obtain a pose transfer pedestrian generation model. The pose transfer pedestrian generation model includes a generator and a discriminator. The discriminator includes an appearance discriminator and a pose discriminator.
[0068] The predetermined pedestrian appearance image and target pose sequence are input into the pose transfer pedestrian generation model to obtain several target pedestrian images; and
[0069] The target pedestrian image is composited into the railway scene image to generate a railway pedestrian intrusion image.
[0070] Optionally, the method further includes:
[0071] Several images of pedestrian intrusions into railway lines are used to generate a video sequence of pedestrian intrusions into railway scenes, based on the target pose sequence.
[0072] Optionally, the method further includes:
[0073] Automatically annotate images of pedestrian intrusions into railway scenes to generate annotated images of pedestrian intrusions into railway scenes.
[0074] Optionally, after synthesizing the target pedestrian image into the railway scene image, the position coordinates of the pedestrians and the height and width data of the pedestrian bounding boxes in the railway pedestrian intrusion image are recorded to obtain an annotated image of the railway scene pedestrian intrusion.
[0075] Optionally, pedestrian pose data can be extracted from public datasets, and the corresponding pedestrian images and the extracted pedestrian pose data can be used as a dataset to train the model, thereby obtaining a pose transfer pedestrian generation model with optimal parameters.
[0076] Optionally, the training methods include:
[0077] Extract pedestrian image samples and corresponding pedestrian poses from the training set;
[0078] The extracted pedestrian image samples and the corresponding pedestrian poses are input into the generator. Based on the pre-set target pose and the appearance of the pedestrian image samples, a pedestrian pose transfer image with the input pedestrian image appearance and target pose is generated.
[0079] The discriminator distinguishes pedestrian pose transfer images; and
[0080] When the predetermined loss value is reached, the pedestrian generation model with optimal posture migration parameters is obtained.
[0081] Optionally, the appearance discriminator is used to determine whether the appearance of the generated pedestrian image is consistent with the appearance of the pedestrian in the input pedestrian image, and the pose discriminator is used to determine whether the generated pedestrian pose is consistent with the target pose.
[0082] The structure of the appearance discrimination device includes:
[0083] The concatenation layer is used to concatenate the input pedestrian pose transfer image and the target pedestrian pose, as well as the target image and the target pedestrian pose, to obtain a concatenation vector.
[0084] Multiple downsampling modules are used to downsample the concatenated vector to obtain the downsampled feature vector; and
[0085] Multiple residual modules are used to extract features from the downsampled feature vector to obtain discriminative features.
[0086] Among them, the appearance discriminator completes the discrimination based on the discriminative features.
[0087] Example 2
[0088] This embodiment was trained using the PyTorch deep learning platform on an Ubuntu 7.5.0-3ubuntu1 operating system and an NVIDIA Tesla V100S GPU. The processor was a Hygon C86 7165, and the batch size was set to 16 samples. Furthermore, this embodiment used the Adam (Adaptive Moment Estimation) optimization algorithm to update the model parameters. The initial learning rate was set to 0.0002, and starting from the 499th epoch, the learning rate was reduced by 16 for each subsequent epoch. The learning rate adjustment method is shown in equation (1), and the two exponential decay rates are set to 0.5 and 0.999, respectively. In addition, the weight parameters of the three loss functions are set as follows: , , and feature mapping channels .
[0089] (1)
[0090] In the formula, l r l is the current learning rate. r0 is the initial learning rate; max is the function to maximize the value; epoch is the number of training iterations.
[0091] like Figure 2 As shown, this embodiment 2 provides a method for generating image sequences of pedestrian intrusion into railways based on pose transfer, including the following steps:
[0092] Step 1: Obtain the pedestrian pose sequence in the non-railway scene as the target pose sequence.
[0093] Specifically, this embodiment takes railway scenario A and non-railway scenario B as examples for illustration. The pedestrian posture sequence in scenario B is extracted as the target posture, and a posture sequence is generated. This posture sequence can be used as the target posture sequence for posture migration from scenario B to scenario A.
[0094] Step 2: Build and train the model to obtain the pose transfer pedestrian generation model. The pose transfer pedestrian generation model includes a generator and a discriminator. The discriminator includes an appearance discriminator and a pose discriminator.
[0095] Specifically, the pose transfer pedestrian generation model is based on GSGAN. The GSGAN algorithm structure includes a generator and two discriminators. The generator's input includes pedestrian appearance and pedestrian pose. Its purpose is to learn the distribution from the original pedestrian pose to the target pedestrian pose, generating a pedestrian image with the target pose to deceive the discriminators. The input is a desired pedestrian appearance image S and a series of desired target pose sequences P. i (i=1,…,N, representing the number of pose sequences), extracted from N frames of video images in scenario B described above. The pose transfer pedestrian generation model internally performs adversarial analysis to generate N pedestrian images with appearance S and pose Pi. The discriminator includes an appearance discriminator and a pose discriminator. The input to the appearance discriminator consists of two parts: real sample pairs and generated sample pairs. The input to the pose discriminator consists of two parts: real pose pairs and generated pose pairs. The two discriminators calculate the loss between the predicted label value and the true label value of the input sample, and further improve the learning ability of the discriminator through gradient propagation, better judging the authenticity, and thus supervising the generator to improve the quality of pedestrian image samples.
[0096] In traditional techniques, pedestrian samples generated by GAN-based network models often suffer from unclear appearance textures and skeleton poses. The effectiveness of generating pedestrian appearance textures and pose transfer largely depends on the algorithm's network structure; therefore, the generator in this embodiment is as follows: Figure 3 The structure shown contains N layers (N=9 in this embodiment), consisting of a GCN module and a spatial attention feature fusion module, and can be used to generate pedestrian images with target poses. The generator in this embodiment solves the problems of appearance texture clarity and pose transfer accuracy, and can generate pedestrian image sample sequences with clear appearance textures and generate specified actions according to a specified pose sequence.
[0097] Specifically, the GCN module is a pose transfer module based on graph convolutional networks. It can aggregate the features of nodes near a given node and learn the node features through weighted aggregation, which facilitates prediction tasks. The pedestrian pose map contains 18 nodes, each with its own features. These features form a feature matrix, and the relationships between the nodes also form a matrix. When the target pose and the corresponding source pose of the input image are input to the GCN module, through layer-to-layer propagation, the source pose is modeled for node relationships via spatial mapping, and then mapped back to the image space to obtain an intermediate pose. This intermediate pose is then used as one input to the spatial attention mechanism feature fusion module.
[0098] Specifically, such as Figure 4As shown, the spatial attention mechanism feature fusion module has two inputs: one is the intermediate pose P generated by the previous GCN module, and the other is the image appearance feature I (the appearance feature of each layer is obtained by adding the appearance features input and output of the previous spatial attention mechanism feature fusion module). Specifically, the pose feature P is input through convolutional layers and activation layers before entering the spatial attention module SA to generate pose P'. P' is then fused with the appearance feature I generated through two convolutional layers to generate the output appearance feature I'. This structure enhances the fusion of appearance and pose features. By adding a spatial attention module SA, it can search for locations where important information is gathered, guiding the output of the appearance feature.
[0099] In this embodiment, the Deepfashion public dataset is used as the training set. The Deepfashion public dataset contains 48,674 clothing image samples, including 50 clothing categories: upper body clothing, lower body clothing, and full-body clothing. The pose skeletons of the 48,674 model clothing samples in the Deepfashion public dataset are extracted to generate 18 keypoint coordinates. In the final pose coordinate data, the pose features of each pedestrian image correspond to the X and Y coordinates of these 18 keypoints.
[0100] like Figure 5 The diagram shows the algorithm framework for pedestrian pose coordinate detection. This module is built upon a convolutional neural network. First, a vector field is used to represent the keypoint association region in the image. Then, based on the extracted association region, the keypoints of pedestrians in the image are determined. Next, the Hungarian algorithm is used to match keypoints of the same pedestrian, ultimately generating a pedestrian pose sequence from video frames. First, single-image pedestrian pose coordinate data from the DeepFashion public dataset is generated based on the algorithm, and this data, along with pedestrian image data, forms a database for training the pose transfer algorithm. Additionally, the generation of pedestrian pose sequences from other non-railway scenes is also based on the pose extraction algorithm, providing target poses for the pose transfer algorithm. The generated pedestrian pose sequences are shown in the diagram. Figure 6 As shown.
[0101] like Figure 7 The algorithm framework model is shown. In the training phase, the DeepFashion public dataset and its pose data are used to train the network. In the application phase, the optimal generator model is used to generate a sequence of pedestrian images with the desired appearance and target pose sequence.
[0102] The training steps for the pose transfer pedestrian generation model are as follows:
[0103] Step 2.1: Extract pedestrian image samples and corresponding pedestrian poses from the training set;
[0104] Step 2.2: Input the extracted pedestrian image samples and the corresponding pedestrian poses into the generator. Based on the pre-set target pose and the appearance of the pedestrian image samples, generate a pedestrian pose transfer image with the input pedestrian image appearance and target pose.
[0105] Step 2.3: The discriminator discriminates the pedestrian pose transfer images;
[0106] Step 2.4: When the predetermined loss value is reached, the pose transfer pedestrian generation model with optimal parameters is obtained.
[0107] In this embodiment, the generator is fed three parts as input: a person image sample from the Deepfashion dataset, the corresponding pose of the sample, and the target pose of the sample. The generator output is a person image corresponding to the target pose. The input person pose and the target pose are first concatenated and then input into the generator network. The concatenation channel includes the RGB channels of the person image sample. To increase the correlation between the source pose and the target pose of the image sample, this embodiment uses a graph convolutional network to extract pose features. Finally, a spatial attention mechanism is used to obtain key pose features and combine them with appearance features for output. After passing through multiple graph convolutional networks and feature fusion, the features are upsampled to obtain the final generated pedestrian image. A discriminator is used to calculate the loss and better supervise the generator. The input to the pose discriminator is a fake sample pair between the generated image and the target pose image, and a real sample pair between the target image and the target pose image; the input to the appearance discriminator is a fake sample pair between the generated image and the input image, and a real sample pair between the target image and the input image.
[0108] The network model of the appearance discriminator is as follows: Figure 8 As shown, firstly, the generated image and target pose pair are concatenated with the target image and target pose pair to obtain vectors. These vectors are then downsampled through two convolutions with a stride of 2, and further processed by six residual modules for feature extraction. The discriminator's task is achieved by judging these features. The network model of the pose discriminator is basically similar to that of the appearance discriminator. This embodiment uses the appearance discriminator as an example for explanation; the pose discriminator will not be described in detail here.
[0109] Three loss functions are employed: adversarial loss, pixel-level L1 loss, and perceptual loss. The adversarial loss comprises the losses from the appearance discriminator and the pose discriminator, aiming to determine the probability of identical appearances in the input images of the appearance discriminator and the degree of matching between the pedestrian pose in the generated image and the target pedestrian pose in the input image of the pose discriminator. The generator aims to generate an image that the discriminator will classify as true, so two pose image pairs are input into the pose discriminator, and the BCE loss function in the PyTorch framework is used to calculate the loss between the discriminator's output value and the label value. The appearance adversarial loss also uses the BCE loss function to calculate the loss between the appearance discriminator's output and the label, inputting two image pairs into the appearance discriminator. Pixel-level L1 loss is used to calculate the difference between the generated image and the target image.
[0110] Perceptual loss aims to reduce pose distortion, making the generated images appear more natural and smooth. The training process of a generative adversarial network (GAN) is an alternating optimization of the generator and discriminator. The generator is trained to minimize the objective function, making the generated data distribution approximate the real image. The discriminator, on the other hand, is trained to maximize the objective function.
[0111] During the training phase, the DeepFashion public dataset and its pose data are used to train the network. The input of GSGAN consists of two sets of pose and image data. Therefore, a target pose and target image are set for each set of data so that the network can learn the pose transfer while ensuring the consistency of the pedestrian's appearance. Then, the two sets of pose data and pedestrian images are input into the network in pairs for training.
[0112] Based on the above network structure and loss function, adversarial training is performed on the generator and discriminator. In the initial training phase, the generator's sample generation performance is poor; the appearance and pose discriminators output 0 for negative samples and 1 for positive samples. With the discriminator parameters fixed, the loss is fed back to the generator, which updates its network weights based on the loss and regenerates pedestrian images. This iteration is repeated until the discriminator can no longer distinguish between positive and negative samples. At this point, the generator has learned some features, and its weights are no longer updated. The generator parameters are fixed, and the discriminator network weights are updated. This iteration is repeated until the discriminator can correctly distinguish between positive and negative samples. At this point, the discriminator network weights are also no longer updated, and the discriminator parameters are fixed. Thus, the optimal generator and discriminator model parameters are obtained, serving as the optimal model for the application phase.
[0113] Step 3: Input the predetermined pedestrian appearance image and target pose sequence into the pose transfer pedestrian generation model to obtain several target pedestrian images.
[0114] Step 4: Composite the target pedestrian image into the railway scene image to generate a railway pedestrian intrusion image.
[0115] Specifically, during the generation process, the size of the pedestrian image needs to be adjusted according to the pedestrian's position in the railway scene to ensure that the generated pedestrian height is similar to that of a real person at that position. In this embodiment, the pedestrian height is generated by projecting parallel straight lines from the real scene onto the image space and comparing them to vanishing points. First, two straight lines parallel to the rails are extracted, and the vanishing points in the image are determined. Since the actual height of a person in the railway scene is fixed, the straight line formed by the head of the same pedestrian at different longitudinal positions on the rails should be parallel to the rails. Using this straight line and the vanishing points of the rail lines, the pedestrian image height at different positions in the image can be determined. Since the pose transfer pedestrian generation module can generate pedestrians with various appearances, by inputting them into different railway scenes, images of pedestrians with different appearances intruding into different railway scenes can be generated.
[0116] To verify the effectiveness of this embodiment, an ablation experiment was conducted, and the network model was evaluated using three evaluation metrics: IS, SSIM, and Pckh. Table 1 shows the test results of the ablation experiment.
[0117] Table 1 Test results of ablation experiment
[0118]
[0119] As shown in Table 1, this example uses the original GAN network as a baseline for comparative analysis. It can be seen that after adding the GCN module, the PCKh0.5 index improved from 0.9651 to 0.9702, indicating that the GCN module enhanced the connection between the original pose and the target pose, thus improving the pose transfer effect. When the SA module was added, the SSIM and IS indices used to evaluate image quality improved by 0.0557 and 0.14 respectively, and PCKh0.5 also improved by 0.0236, indicating that the SA module plays an important role in improving the pose transfer effect and the quality of the generated image.
[0120] Secondly, to verify the effectiveness of the proposed GSGAN algorithm, this embodiment compares it with the classic pedestrian pose transfer algorithm PATN. In this embodiment, the training datasets for both the GSGAN and PATN algorithms are the same, and the relevant parameters of PATN are set to their optimal values. Figure 9As shown, compared to the GSGAN algorithm in the second row, the PATN algorithm in the first row performs worse in terms of pedestrian clothing details and facial feature details. The lower body of the first three images generated by PATN is very blurry, while the legs of the first three images generated by the GSGAN algorithm are clearly visible. As shown in Table 2, the test results of the two algorithms are compared. The GSGAN algorithm achieved higher scores in all three metrics. PCKh0.5 of 0.9938 indicates that the pedestrian pose generated by the GSGAN algorithm in this embodiment is basically no different from the target pose, which is 0.0284 higher than that of PATN, proving that the method in this embodiment has a better pose transfer effect.
[0121] Table 2 Test results of the two algorithms
[0122]
[0123] The image synthesis algorithm in this example is based on determining the pedestrian size while maintaining the track gauge-pedestrian size ratio, and then compositing it with the railway background. First, images of empty railway scenes during normal operation were collected from surveillance videos of sections of the Yuanping Railway, the circular railway test field of the China Academy of Railway Sciences, the Guangzhou-Shenzhen High-Speed Railway, and the Baoji-Lanzhou High-Speed Railway. These images included different locations on the railway, such as tunnel entrances, switches, and cuttings, as well as different weather conditions, such as sunny, cloudy, and rainy days. The image resolution was 1920×1080. The images were preprocessed to filter out noise other than the tracks, such as trees. Hough transform was used to detect straight lines on the tracks and hidden points. Projection was used to determine the straight lines between the tracks, thus determining the sleeper pixels at different locations and ultimately estimating the pedestrian size. Finally, the images were composited with the railway background, as shown below. Figure 10 The image shown is the result of generating a sequence of pedestrian intrusion images along the railway. Finally, a total of 11 pedestrian appearances and 12 railway scenes were generated, each sequence consisting of 92 consecutive frames of pedestrian intrusion. Converting the image sequences into videos yields 132 videos, totaling 12,144 images. Furthermore, while deep learning algorithms are trained based on the intrusion images and their labeled data, labeling such a large amount of image data is labor-intensive. The method presented in this paper automatically obtains the labeled data while generating the intrusion images, saving both time and manpower.
[0124] To verify whether the generated images meet the requirements of the training samples, this example uses four widely used pre-trained network models—Yolov3, Faster R-CNN, SSD, and DSSD-321—to repeatedly test the image sequences, and uses Average Precision (AP) to evaluate the results. Table 3 shows the test results of the generated image sequences. Each detection network achieves high test scores on the generated sample database and the scores are close to those on the real database, indicating that the generated railway pedestrian intrusion image sequences are close to real pedestrian intrusion images, proving that the generated intrusion image sequences meet the requirements of the training samples.
[0125] Table 3. Evaluation of the object detection network on generated and real samples
[0126]
[0127] Furthermore, this example will ultimately provide training data for the foreign object intrusion detection network. To verify the effectiveness of using generated pedestrian intrusion images as training data, this example uses them as training data to verify their improvement effect on the detection network model. This paper sets up five datasets. The first dataset contains only 672 real railway pedestrian intrusion images, including 966 intruding pedestrians. Each subsequent dataset adds 1000 generated railway pedestrian intrusion images to the previous dataset, with 1000 more intruding pedestrians. Using these five datasets, a target detection network is trained, resulting in five target detection models. These models are then tested to verify the effectiveness of the generated data. This paper uses the YOLOv3 network model and evaluates it using AP. The results are shown in Table 4. YOLOv3 can be trained to a satisfactory state using only 672 real data images, with a detection accuracy of 71.6%. However, after adding 1000 generated data images, the detection accuracy of the model improved by 0.4%, Model 3 improved by 7.4%, Model 4 improved by 9.4%, and Model 5 improved by 12.7%. This shows that the intrusion data generated in this paper can improve the detection accuracy of the target detection network model. In other words, the generated image sequences can be used as training data for the research and testing of the detection network.
[0128] Table 4 Test results of models with different training sets
[0129]
[0130] Step 5: Generate a railway scene intrusion video sequence from several images of pedestrians intruding into the railway according to the target pose sequence.
[0131] Specifically, multiple generated images of pedestrians in railway scenes are used to generate a railway scene intrusion video sequence in railway scene A, which has the same posture as in non-railway scene B but a different appearance. This video sequence can generate a railway scene pedestrian intrusion video sequence with different postures, different appearances and continuous actions based on different pedestrian posture sequences in different scenes B, posture-migrating pedestrians and input desired pedestrian appearance.
[0132] Step 6: Automatically annotate the images of pedestrian intrusions into the railway to generate annotated images of pedestrian intrusions into the railway scene.
[0133] Specifically, the railway scene intrusion pedestrian image synthesis module can work automatically during its operation, generating labeled image samples of railway scene pedestrian intrusion according to the format of labeled data in different datasets. It can synthesize pedestrian images with different appearances and postures generated by the pose transfer pedestrian generation module into the railway scene. During the synthesis process, the size of the pedestrian in the image is related to its position in the railway scene. In reality, the annotation process of railway intrusion pedestrian images involves annotating the position, height, and width of the pedestrian bounding box in the image. Therefore, in this embodiment, the position and size of the pedestrian in the synthesized image are directly recorded during the railway scene intrusion pedestrian image synthesis process. Together with the synthesized image, this constitutes an automatic annotation dataset for railway pedestrian intrusion images, including the corresponding railway intrusion image, the position of the synthesized pedestrian target, and the height and width of the pedestrian bounding box.
[0134] like Figure 11 As shown in the figure, box 1 represents the railway scene image, with its upper left corner at pixel coordinates (0,0). Box 2 represents the minimum bounding rectangle of the pedestrian image after it has been synthesized into the railway scene image, with its upper left corner at pixel coordinates (x0,y0) in the railway scene image, a height of H, and a width of W. Based on these parameters, labeled images in the railway scene can be automatically generated. The automatically generated labeled data can be adjusted in content and format according to different data and required labeling formats to adapt to the usage requirements of different datasets. Depending on the position of the generated pedestrian image in the railway scene image, pedestrian images with different heights H and widths W will be generated according to the railway scene pedestrian intrusion image synthesis method, and different railway pedestrian intrusion image labeled data samples will be generated.
[0135] In summary, this embodiment of the invention obtains a target pose sequence by acquiring pedestrian pose sequences in non-railway scenes, then builds and trains a model to obtain a pose transfer pedestrian generation model. The predetermined pedestrian appearance image and target pose sequence are then input into the pose transfer pedestrian generation model to obtain several target pedestrian images. These target pedestrian images are then inserted into railway scene images to generate railway pedestrian intrusion images. This process yields a rich sequence of railway scene pedestrian intrusion images, featuring clear pedestrian appearance textures, distinct poses, and high realism. This solves the problem of scarce and difficult-to-obtain railway scene pedestrian intrusion images.
[0136] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.
[0137] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for method or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the description of the method embodiments. The method and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0138] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for generating an image sequence of railway pedestrian intrusion based on pose migration, characterized in that, The method comprises: acquiring a pedestrian posture sequence in a non-railway scene as a target posture sequence; establishing a model and training to obtain a posture transfer pedestrian generation model, the posture transfer pedestrian generation model comprising a generator and a discriminator, the discriminator comprising an appearance discriminator and a posture discriminator; inputting a predetermined pedestrian appearance image and the target posture sequence into the posture transfer pedestrian generation model to obtain a plurality of target pedestrian images; and synthesizing the target pedestrian images into a railway scene image to generate a railway pedestrian intrusion image; wherein the appearance discriminator is used to determine whether the appearance of the generated pedestrian image is consistent with the pedestrian appearance of the input pedestrian image, and the posture discriminator is used to determine whether the posture of the generated pedestrian is consistent with the target posture; the appearance discriminator comprises: a concatenation layer used to concatenate the input pedestrian posture transfer image and the target pedestrian posture, and the target image and the target pedestrian posture respectively to obtain a concatenation vector; a plurality of down-sampling modules used to down-sample the concatenation vector to obtain a down-sampled feature vector; and a plurality of residual modules used to extract features from the down-sampled feature vector to obtain a discrimination feature, wherein the appearance discriminator completes discrimination based on the discrimination feature; the synthesizing of the target pedestrian images into the railway scene image to generate the railway pedestrian intrusion image comprises: determining the size of the pedestrian image at different positions by using a method based on parallel lines having the same vanishing point, and adjusting the target pedestrian image based on the determined size of the pedestrian image; synthesizing the adjusted target pedestrian image into the railway scene image to obtain the railway pedestrian intrusion image; the generator comprises a GCN module and a spatial attention feature fusion module, one input of the spatial attention feature fusion module is an intermediate posture P generated by the last GCN module, and the other input is an image appearance feature I, the image appearance feature I of each layer is obtained by adding the image appearance feature input by the spatial attention feature fusion module of the last layer and the output appearance feature, the intermediate posture P input enters a spatial attention module SA through a convolution layer and an activation layer to generate a posture P', and P' is fused with the feature generated by two layers of convolution of the appearance feature I to generate an output appearance feature I'; generating a railway scene pedestrian intrusion video sequence according to the target posture sequence based on a plurality of the railway pedestrian intrusion images; automatically labeling the railway pedestrian intrusion images to generate labeled images of railway scene pedestrian intrusion.
2. The method of claim 1, wherein, After synthesizing the target pedestrian image into the railway scene image, record the position coordinates of the pedestrian in the railway pedestrian intrusion image and the height data and width data of the pedestrian box to obtain the labeled images of the railway scene pedestrian intrusion.
3. The method of claim 1, wherein, Extract pedestrian posture data of a public data set, and use the corresponding pedestrian image and the extracted pedestrian posture data as a training set to train the model to obtain the posture transfer pedestrian generation model with optimal parameters.
4. The method of claim 3, wherein, The training method comprises: extracting pedestrian image samples and corresponding pedestrian postures in the training set; inputting the extracted pedestrian image sample and a corresponding pedestrian posture of the pedestrian image sample and a preset target posture into a generator, generating a pedestrian posture migration image with a pedestrian image appearance and a target posture of the input according to the preset target posture and an appearance of the pedestrian image sample; the discriminator discriminates the pedestrian posture migration image; and when a predetermined loss value is reached, obtaining the posture migration pedestrian generation model with optimal parameters.
5. A system for generating a sequence of images for railway pedestrian intrusion based on pose migration, characterized by, The system comprises: a posture extraction module configured to obtain a pedestrian posture sequence in a non-railway scene, obtain a target posture sequence, and extract pedestrian posture data of a public data set, and generate a data set of pedestrian images corresponding to the pedestrian posture data; a posture migration pedestrian generation module configured to establish a model, train the model by using the data set generated by the pedestrian images corresponding to the pedestrian posture data, obtain a posture migration pedestrian generation model with optimal parameters, and input a predetermined pedestrian appearance image and the target posture sequence into the posture migration pedestrian generation model to obtain a plurality of target pedestrian images; and a railway scene intrusion pedestrian image synthesis module configured to insert the target pedestrian images into railway scene images to generate railway pedestrian intrusion images. The appearance discriminator is configured to determine whether the appearance of the generated pedestrian image is consistent with the pedestrian appearance of the input pedestrian image, and the posture discriminator is configured to determine whether the generated pedestrian posture is consistent with the target posture. The appearance discriminator comprises: a concatenation layer configured to concatenate the input pedestrian posture migration image and the target pedestrian posture, and concatenate the target image and the target pedestrian posture respectively to obtain a concatenation vector; a plurality of down-sampling modules configured to down-sample the concatenation vector to obtain a down-sampled feature vector; and a plurality of residual modules configured to extract features from the down-sampled feature vector to obtain discrimination features. The appearance discriminator completes discrimination based on the discrimination features. The railway scene intrusion pedestrian image synthesis module is specifically configured to: determine the size of the pedestrian image at different positions by using a method based on parallel lines having the same vanishing point, and adjust the target pedestrian image based on the determined size of the pedestrian image; synthesizes the adjusted target pedestrian image into the railway scene image to obtain the railway pedestrian intrusion image. The generator comprises a GCN module and a spatial attention feature fusion module. One input of the spatial attention feature fusion module is an intermediate posture P generated by the last GCN module, and the other input is an image appearance feature I. The image appearance feature I of each layer is obtained by adding the image appearance feature input by the spatial attention feature fusion module of the last layer and the output appearance feature. The intermediate posture P input passes through a convolution layer and an activation layer to enter a spatial attention module SA to generate a posture P'. P' is fused with the feature generated by two layers of convolution of the appearance feature I to generate the output appearance feature I'. The railway scene intrusion video generation module is configured to generate a railway scene intrusion video sequence according to the target posture sequence by combining a plurality of railway scene intrusion pedestrian images generated by the railway scene intrusion pedestrian image synthesis module. The pedestrian intrusion image automatic labeling module is configured to label the railway pedestrian intrusion image to generate a labeled image of the railway scene pedestrian intrusion.
Citation Information
Patent Citations
Method for generating picture of pedestrian with arbitrary pose
CN108564119A
Human body posture migration method based on attention mechanism
CN111161200A