Training method of vehicle detection model, traffic state perception method and system

By extracting features and generating rotating anchor frames from single-temporal remote sensing images, a vehicle detection model is trained, which solves the problems of poor accuracy and efficiency in vehicle detection in traditional methods, achieves efficient traffic condition perception, and reduces costs.

CN119723477BActive Publication Date: 2026-01-16CHINA ENERGY ENG GRP GUANGDONG ELECTRIC POWER DESIGN INST CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411702575.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2026-01-16
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Traditional traffic condition perception methods based on remote sensing images have poor detection accuracy and efficiency when detecting vehicles, and the acquisition of multi-temporal remote sensing images is difficult and costly, affecting their practicality.

Method used

By extracting features from panchromatic images in single-temporal remote sensing images and generating rotated anchor boxes, a vehicle detection model is trained to identify vehicles of any orientation, position, and size. Combined with single-temporal remote sensing images, vehicle target detection and speed analysis are performed, reducing the cost of traffic condition perception.

Benefits of technology

It improves the accuracy and efficiency of vehicle detection, reduces the cost of traffic condition perception, and enhances practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723477B_ABST
    Figure CN119723477B_ABST
Patent Text Reader

Abstract

The application discloses a kind of training method of vehicle detection model, traffic state perception method and system, wherein the training method obtains panchromatic image in single-phase remote sensing image, and the panchromatic image is preprocessed, to obtain panchromatic training image;Image feature extraction is carried out to the panchromatic training image, to obtain panchromatic image feature;Panchromatic image feature is extracted to candidate region, to obtain target rotating anchor frame set, target rotating anchor frame set includes several target rotating anchor frames corresponding to target vehicle;Rotating region pooling is carried out to target rotating anchor frame set, to obtain anchor frame feature map;According to anchor frame feature map, the parameter of initialized vehicle detection model is updated, to obtain trained vehicle detection model.The method can improve the detection accuracy and detection efficiency of vehicle detection, and reduce the cost of traffic state perception, improve the practicability of traffic state perception by providing a kind of vehicle detection model.The application relates to the technical field of vehicle detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle detection, and in particular to a vehicle detection model training method, a traffic state perception method and system. BACKGROUND

[0002] At present, the traditional traffic state perception method based on remote sensing images usually realizes vehicle detection based on multi-temporal remote sensing images, and then realizes traffic state perception based on the detected multi-temporal vehicle data. However, the detection accuracy and efficiency of this method are not good when performing vehicle detection. Moreover, due to the difficulty and high cost of obtaining multi-temporal remote sensing images, the practicability of this method for traffic state perception is not satisfactory.

[0003] Therefore, the problems of the prior art still need to be solved and optimized. SUMMARY

[0004] The present application aims to at least partly solve one of the problems in the related art.

[0005] To this end, one object of the present application is to provide a vehicle detection model training method, a traffic state perception method and system, wherein the method can improve the detection accuracy and efficiency of vehicle detection, and reduce the cost of traffic state perception, and improve the practicability of traffic state perception.

[0006] To achieve the above technical purpose, the technical solutions adopted by the embodiments of the present application include:

[0007] In a first aspect, the present application provides a vehicle detection model training method, a traffic state perception method and system, comprising:

[0008] obtaining a panchromatic image in a single-temporal remote sensing image about a target vehicle, and pre-processing the panchromatic image to obtain a panchromatic training image;

[0009] performing image feature extraction on the panchromatic training image to obtain a panchromatic image feature;

[0010] performing candidate region extraction on the panchromatic image feature to obtain a target rotated anchor frame set, the target rotated anchor frame set comprising a plurality of target rotated anchor frames corresponding to the target vehicle, and each target rotated anchor frame corresponding to at least one of a different bounding box scale, aspect ratio or rotation angle;

[0011] performing rotated region pooling on the target rotated anchor frame set to obtain an anchor frame feature map;

[0012] performing parameter updating on an initialized vehicle detection model according to the anchor frame feature map to obtain a trained vehicle detection model.

[0013] In addition, the method according to the above-mentioned embodiments of the present application can further have the following additional technical features:

[0014] Further, in an embodiment of the present application, the visual scene recognition model comprises a plurality of cascaded dynamic pooling network frameworks, and the multi-layer dynamic pooling aggregation of the local features comprises:

[0015] The local features are input into the plurality of cascaded dynamic pooling network frameworks to obtain a plurality of image features, each of which corresponds to one of the dynamic pooling network frameworks;

[0016] The feature flattening is performed on all the image features to obtain flattened features corresponding to different feature levels;

[0017] The feature flattening is performed on all the image features to obtain flattened features corresponding to different feature levels;

[0018] Further, in an embodiment of the present application, the candidate region extraction of the panchromatic image features to obtain the target set of rotated anchor boxes corresponding to the panchromatic image features comprises:

[0019] The panchromatic image features are subjected to rotated anchor box generation processing to obtain a first intermediate anchor box set, the first intermediate anchor box set comprising a plurality of intermediate rotated anchor boxes corresponding to the panchromatic image features, each of the intermediate rotated anchor boxes having at least one of a different region scale, aspect ratio or rotation angle;

[0020] The anchor box feature mapping is performed on the first intermediate anchor box set to obtain a second intermediate anchor box set;

[0021] The second intermediate anchor box set is subjected to bounding box classification regression to obtain the target set of rotated anchor boxes.

[0022] Further, in an embodiment of the present application, the bounding box classification regression of the second intermediate anchor box set to obtain the target set of rotated anchor boxes comprises:

[0023] The anchor box classification is performed on the second intermediate anchor box set to obtain a third intermediate anchor box set, the third intermediate anchor box set being a collection of intermediate rotated anchor boxes corresponding to the target vehicle;

[0024] The bounding box regression offset update is performed on the third intermediate anchor box set to obtain the target set of rotated anchor boxes.

[0025] Further, in an embodiment of the present application, the bounding box regression offset update of the third intermediate anchor box set to obtain the target set of rotated anchor boxes comprises:

[0026] performing bounding box offset regression prediction on the third intermediate anchor box set to obtain bounding box offset prediction data;

[0027] performing bounding box offset update on the third intermediate anchor box set according to the bounding box offset prediction data to obtain the target rotated anchor box set.

[0028] Further, in an embodiment of the present application, the performing rotated region pooling on the target rotated anchor box set to obtain an anchor box feature map comprises:

[0029] performing rotated coordinate transformation on the target rotated anchor box set to obtain a fourth rotated anchor box set, each target rotated anchor box in the fourth rotated anchor box set corresponding to the same rotation angle;

[0030] performing region scale cropping on the fourth rotated anchor box set to obtain a fifth rotated anchor box set;

[0031] performing feature map pooling mapping on the fifth rotated anchor box set to obtain the anchor box feature map.

[0032] In a second aspect, embodiments of the present application provide a traffic state perception method of a vehicle detection model, comprising:

[0033] obtaining traffic network data and a target single-time-phase remote sensing image, the target single-time-phase remote sensing image comprising a target panchromatic image and a target multispectral image;

[0034] inputting the target panchromatic image into the trained vehicle detection model to perform vehicle target detection, to obtain a vehicle target detection result;

[0035] performing vehicle speed analysis processing on the target multispectral image according to the vehicle target detection result, to obtain vehicle speed data;

[0036] performing traffic state perception on the traffic network data according to the vehicle speed data and the vehicle target detection result, to obtain a traffic state perception result.

[0037] Further, in embodiments of the present application, the performing vehicle speed analysis processing on the target multispectral image according to the vehicle target detection result, to obtain vehicle speed data, comprises:

[0038] obtaining a first waveband spectral image and a second waveband spectral image according to the target multispectral image;

[0039] performing first candidate region extraction on the first waveband spectral image according to the vehicle target detection result, to obtain a first waveband candidate region, and performing second candidate region extraction on the second waveband spectral image according to the vehicle target detection result, to obtain a second waveband candidate region;

[0040] According to the first waveband candidate region, the second waveband candidate region is subjected to region displacement matching to obtain the vehicle moving speed data.

[0041] Further, in the embodiment of the present application, the vehicle moving speed data is obtained by subjecting the second waveband candidate region to region displacement matching according to the first waveband candidate region, which comprises:

[0042] A spectral imaging time difference of the target multispectral image is obtained.

[0043] According to the first waveband candidate region, the second waveband candidate region is subjected to template matching to obtain first vehicle pixel coordinate data and second vehicle pixel coordinate data, the first vehicle pixel coordinate data being pixel coordinate data of a traveling vehicle in the first waveband candidate region, and the second vehicle pixel coordinate data being pixel coordinate data of the traveling vehicle in the second waveband candidate region.

[0044] According to the first vehicle pixel coordinate data, the second vehicle pixel coordinate data is subjected to vehicle displacement calculation to obtain vehicle displacement data.

[0045] According to the spectral imaging time difference, the vehicle displacement data is subjected to vehicle moving speed calculation to obtain the vehicle moving speed data.

[0046] Further, in the embodiment of the present application, the traffic state perception result is obtained by subjecting the traffic network data to traffic state perception according to the vehicle moving speed data and the vehicle target detection result, which comprises:

[0047] According to the vehicle target detection result, the traffic network data is subjected to vehicle space matching to obtain vehicle space matching data.

[0048] According to the vehicle moving speed data, the traffic state perception result is obtained by subjecting the vehicle space matching data to traffic state evaluation.

[0049] In a third aspect, an embodiment of the present application provides a training system of a vehicle detection model, comprising:

[0050] A first processing unit is configured to obtain a panchromatic image in a single-phase remote sensing image of a target vehicle, and to pre-process the panchromatic image to obtain a panchromatic training image.

[0051] A second processing unit is configured to extract image features from the panchromatic training image to obtain panchromatic image features.

[0052] a third processing unit configured to perform candidate region extraction on the full-color image features to obtain a target rotated anchor box set, the target rotated anchor box set including a plurality of target rotated anchor boxes corresponding to the target vehicle, each target rotated anchor box having at least one of a different bounding box scale, aspect ratio, or rotation angle;

[0053] a fourth processing unit configured to perform rotated region pooling on the target rotated anchor box set to obtain an anchor box feature map;

[0054] a fifth processing unit configured to perform parameter updating on an initialized vehicle detection model according to the anchor box feature map to obtain a trained vehicle detection model.

[0055] In a fourth aspect, an embodiment of the present application provides a traffic state perception system of a vehicle detection model, including:

[0056] a sixth processing unit configured to obtain traffic network data and a target single-time-phase remote sensing image, the target single-time-phase remote sensing image including a target full-color image and a target multi-spectral image;

[0057] a seventh processing unit configured to input the target full-color image into the trained vehicle detection model to perform vehicle target detection, and obtain a vehicle target detection result;

[0058] an eighth processing unit configured to perform vehicle speed analysis processing on the target multi-spectral image according to the vehicle target detection result, and obtain vehicle speed data;

[0059] a ninth unit configured to perform traffic state perception on the traffic network data according to the vehicle speed data and the vehicle target detection result, and obtain a traffic state perception result.

[0060] In a fifth aspect, an embodiment of the present application further provides an electronic device, including:

[0061] at least one processor;

[0062] at least one memory configured to store at least one program;

[0063] when the at least one program is executed by the at least one processor, the at least one processor implements the method described above.

[0064] In a sixth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a processor executable program, and the processor executable program is used to implement the method described above when executed by the processor.

[0065] Advantages and beneficial effects of the present application will be partially given in the following description, partially will become apparent from the following description, or will be learned by practicing the present application:

[0066] The training method of the vehicle detection model, the traffic state perception method and the system disclosed by the embodiments of the present application, wherein the training method obtains a panchromatic image in a single-phase remote sensing image about a target vehicle, and pre-processes the panchromatic image to obtain a panchromatic training image; image feature extraction is performed on the panchromatic training image to obtain panchromatic image features; candidate region extraction is performed on the panchromatic image features to obtain a target rotating anchor frame set, the target rotating anchor frame set includes a plurality of target rotating anchor frames corresponding to the target vehicle, and at least one of the frame scale, aspect ratio or rotation angle corresponding to each target rotating anchor frame is different; the anchor frame feature map is obtained by performing rotating region pooling on the target rotating anchor frame set; and the initialized vehicle detection model is updated in parameters according to the anchor frame feature map to obtain a trained vehicle detection model. The training method extracts a plurality of target rotating anchor frames with at least one of the frame scale, aspect ratio or rotation angle being different by performing candidate region extraction on the panchromatic image features in the single-phase remote sensing image, and performs rotating region pooling on all the extracted target rotating anchor frames, which can make the trained vehicle detection model be able to identify vehicles of any orientation, position and size, and can effectively improve the detection accuracy and efficiency of vehicle detection. In addition, the traffic state perception method uses the single-phase remote sensing image to perform vehicle target detection by the vehicle detection model, and then performs vehicle speed analysis and traffic state perception according to the obtained vehicle target detection result, which can effectively reduce the cost of traffic state perception and improve the practicality of traffic state perception. BRIEF DESCRIPTION OF DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly expressing part of the embodiments of the technical solutions of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the premise of the drawings.

[0068] Figure 1 A flowchart of a training method of a vehicle detection model provided by the embodiments of the present application is shown in the figure.

[0069] Figure 2 A flowchart of a traffic state perception method of a vehicle detection model provided by the embodiments of the present application is shown in the figure.

[0070] Figure 3 A structural framework diagram of a training system of a vehicle detection model provided by the embodiments of the present application is shown in the figure.

[0071] Figure 4 A structural schematic diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0072] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar components have the same or similar designations and functions throughout the various figures and embodiments. The embodiments described below are examples in which the present application is applied, and are merely intended to explain the present application, and should not be understood as limiting the present application. The step numbers in the following embodiments are merely set for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the description herein is for describing the embodiments of the present application only and is not intended to limit the present application.

[0074] First, the technical terms involved in the embodiments of the present application are explained as follows:

[0075] Single-time-phase remote sensing image: refers to a remote sensing image acquired by a high spatial resolution satellite at a specific time point.

[0076] Multi-time-phase remote sensing image: refers to a plurality of remote sensing images acquired by a high spatial resolution satellite at different time points.

[0077] Panchromatic image (PAN): a black and white image acquired only in the visible light band, having the highest spatial resolution.

[0078] Multispectral image (MS): a color image acquired in multiple spectral bands, usually having a lower spatial resolution than the panchromatic image.

[0079] Ghost phenomenon: refers to a significant difference in the imaging position of a high-speed moving object on images in different spectral bands, resulting in repeated or blurred images in the synthesized image. Specifically, since a high spatial resolution satellite often carries multiple sensors, the images in different spectral bands acquired by the high spatial resolution satellite at the same time phase will have a slight imaging time difference between the sensors, causing the imaging position of a high-speed moving vehicle to shift on the images in two different spectral bands.

[0080] Currently, a traditional traffic state perception method based on remote sensing images usually detects vehicles based on multi-temporal remote sensing images, and then perceives the traffic state based on the detected multi-temporal vehicle data. However, since this method uses multi-temporal remote sensing images, it needs to perform vehicle detection on remote sensing images at different times respectively, and needs more computing resources. The generated candidate regions are usually horizontal target rectangular frames, which cannot well meet the detection requirements of vehicles in any direction, and the detection accuracy and efficiency of vehicle detection are not good. In addition, when obtaining multi-temporal remote sensing images, high spatial resolution satellites are needed to collect remote sensing images at different time points. However, satellite resources of high spatial resolution satellites are scarce, and it is difficult and costly to obtain multi-temporal remote sensing images, and the practicability of traffic state perception is not satisfactory.

[0081] Therefore, the embodiments of the present application provide a vehicle detection model training method, a traffic state perception method and system, wherein the training method extracts candidate regions from full-color image features in single-temporal remote sensing images to extract a plurality of target rotated anchor frames with at least one of different frame sizes, aspect ratios or rotation angles, and performs rotated region pooling on all the extracted target rotated anchor frames. The vehicle detection model trained in this way can generate vehicle target detection results in any direction, position and size based on single-temporal remote sensing images, and can effectively improve the detection accuracy and efficiency of vehicle detection. In addition, the traffic state perception method uses the vehicle detection model to perform vehicle target detection using single-temporal remote sensing images, and then performs vehicle speed analysis and traffic state perception based on the obtained vehicle target detection results, which can effectively reduce the cost of traffic state perception and improve the practicability of traffic state perception.

[0082] Reference Figure 1 In the embodiments of the present application, a vehicle detection model training method comprises:

[0083] In step 110, a full-color image in single-temporal remote sensing images about a target vehicle is obtained, and the full-color image is preprocessed to obtain a full-color training image.

[0084] In the embodiments of the present application, single-temporal remote sensing images about a target vehicle for training can be obtained, which can be collected by a high spatial resolution satellite or obtained from a publicly available remote sensing image dataset (such as a WorldView-3 remote sensing image dataset or a ZY-3 remote sensing image dataset) on the Internet. In addition, the obtained single-temporal remote sensing images include a full-color image and a multi-spectral image. In step 110, the full-color image can be cut according to a preset pixel size (such as 512x512 pixels), and then the rotated frame of a vehicle target in each cut full-color image block is labeled to obtain a full-color training image.

[0085] Step 120, image feature extraction is performed on the panchromatic training image to obtain a panchromatic image feature;

[0086] In the embodiments of the present application, image feature extraction is used to extract image features in the panchromatic training image and generate a corresponding feature map. Specifically, step 120 can be inputting the panchromatic training image into a pre-trained deep convolutional neural network, such as a feature pyramid network ResNet50-FPN, capturing different vehicle target features in the panchromatic training image through multi-layer convolution and pooling operations, and determining the generated high-level feature map as the panchromatic image feature.

[0087] Step 130, candidate region extraction is performed on the panchromatic image feature to obtain a target rotation anchor box set, the target rotation anchor box set including a plurality of target rotation anchor boxes corresponding to the target vehicle, each target rotation anchor box corresponding to at least one different bounding box scale, aspect ratio or rotation angle;

[0088] In the embodiments of the present application, step 130 can extract target rotation anchor boxes related to the target vehicle in the feature map of the panchromatic image feature, each target rotation anchor box corresponding to at least one different bounding box scale, aspect ratio or rotation angle, and determine all obtained target rotation anchor boxes as the target rotation anchor box set.

[0089] In some embodiments, the step 130, candidate region extraction is performed on the panchromatic image feature to obtain a target rotation anchor box set corresponding to the panchromatic image feature, including:

[0090] A1, performing a rotation anchor box generation process on the panchromatic image feature to obtain a first intermediate anchor box set, the first intermediate anchor box set including a plurality of intermediate rotation anchor boxes corresponding to the panchromatic image feature, each intermediate rotation anchor box corresponding to at least one different region scale, aspect ratio or rotation angle;

[0091] A2, performing anchor box feature mapping on the first intermediate anchor box set to obtain a second intermediate anchor box set;

[0092] In the embodiments of the present application, step A1 can be to generate a series of intermediate rotating anchor boxes with different bounding box scales, aspect ratios and rotation angles at each feature position in the feature map of the full-color image features. Specifically, for a feature position in the feature map of the full-color image features, step A1 can generate a plurality of intermediate rotating anchor boxes based on the feature position, wherein the bounding box scale of the intermediate rotating anchor boxes can be any one of 32 pixels, 64 pixels, 128 pixels, 256 pixels, 512 pixels, etc., the aspect ratio can be any one of 1:1, 1:2, 2:1, etc., and the rotation angle can be any one of 0°, 30°, 60°, 90°, 120°, 150°, etc. By combining different bounding box scales, aspect ratios and rotation angles, a plurality of intermediate rotating anchor boxes are obtained, each of which has at least one different bounding box scale, aspect ratio or rotation angle. The examples in the present application are for illustration only and do not limit the bounding box scale, aspect ratio and rotation angle. The specific rotation angle can also be any one of 15°, 45°, etc., and the bounding box scale and aspect ratio are the same. In addition, for the remaining feature positions in the feature map of the full-color image features, similar to the foregoing, a simple analogy can be made, which will not be described here in detail.

[0093] It can be understood that the anchor box feature mapping in step A2 is used to improve the feature expression capability of each intermediate rotating anchor box, which can specifically input each intermediate rotating anchor box in the first intermediate anchor box set to a 3x3 size convolution layer, and then convert the features in the intermediate rotating anchor box through a nonlinear activation function to obtain the second intermediate anchor box set.

[0094] A3, performing bounding box classification regression on the second intermediate anchor box set to obtain the target rotating anchor box set.

[0095] Further, the step A3, the bounding box classification regression on the second intermediate anchor box set to obtain the target rotating anchor box set, comprises:

[0096] A31, performing anchor box classification on the second intermediate anchor box set to obtain a third intermediate anchor box set, the third intermediate anchor box set being a set of intermediate rotating anchor boxes corresponding to the target vehicle;

[0097] In the embodiments of the present application, step A31 is used to judge whether the intermediate rotating anchor boxes in the second intermediate anchor box set are foreground candidate regions containing the target vehicle or background candidate regions not containing the target vehicle, and then all intermediate rotating anchor boxes containing the target vehicle are determined as the third intermediate anchor box set.

[0098] Exemplarily, the anchor box classification in the step A31 can be inputting the second intermediate anchor box set into a classification layer for binary classification, specifically, a 1x1 convolution operation can be adopted to process all the intermediate rotated anchor boxes in the second intermediate anchor box set, and output the foreground candidate region probability and the background candidate region probability of each intermediate rotated anchor box; then, the output is mapped to a probability value through a Softmax function, and the intermediate rotated anchor boxes possibly containing the target vehicle are screened based on a preset probability threshold, to obtain a third intermediate anchor box set.

[0099] A32, performing anchor box regression offset update on the third intermediate anchor box set to obtain the target rotated anchor box set.

[0100] Further, the step A32 of performing anchor box regression offset update on the third intermediate anchor box set to obtain the target rotated anchor box set comprises:

[0101] A321, performing anchor box offset regression prediction on the third intermediate anchor box set to obtain anchor box offset prediction data;

[0102] A322, performing anchor box offset update on the third intermediate anchor box set according to the anchor box offset prediction data to obtain the target rotated anchor box set.

[0103] In the embodiments of the present application, the step A32 is used to generate target rotated anchor boxes with more accurate positioning on the basis of the intermediate rotated anchor boxes. Specifically, the step A321 can be performing regression prediction on all the intermediate rotated anchor boxes in the third intermediate anchor box set through a 1x1 convolution operation, outputting five offset parameters, which are the horizontal and vertical offsets (Δx, Δy) of the center coordinates of the intermediate rotated anchor boxes, the width and height offsets (Δw, Δh) and the rotation angle offset Δθ, and determining the five offset parameters as the anchor box offset prediction data.

[0104] It can be understood that the step A322 can be adjusting the corresponding parameters of the intermediate rotated anchor boxes respectively based on the obtained anchor box offset prediction data, so as to generate target rotated anchor boxes with accurate positioning, and then determining all the generated target rotated anchor boxes as the target rotated anchor box set.

[0105] Step 140, performing rotated region pooling on the target rotated anchor box set to obtain an anchor box feature map;

[0106] In the embodiments of the present application, the rotated region pooling is used to extract the feature blocks containing the target vehicle in each target rotated anchor box in the target rotated anchor box set, and pool map the feature blocks containing the target vehicle to a fixed-size feature map, so as to obtain the anchor box feature map.

[0107] In some embodiments, the step 140 of performing the rotated region pooling on the target rotated anchor box set to obtain an anchor box feature map comprises:

[0108] B1, performing a rotated coordinate transformation on the target rotated anchor box set to obtain a fourth rotated anchor box set, each target rotated anchor box in the fourth rotated anchor box set corresponding to a same rotation angle;

[0109] B2, performing region scale cropping on the fourth rotated anchor box set to obtain a fifth rotated anchor box set;

[0110] B3, performing feature map pooling mapping on the fifth rotated anchor box set to obtain the anchor box feature map.

[0111] In the embodiments of the present application, the step B1 can be inputting the target rotated anchor box set into a rotation coordinate transformation layer, and aligning the rotation angle of each target rotated anchor box in the target rotated anchor box set to a unified direction through the rotation coordinate transformation layer. Specifically, an affine transformation matrix corresponding to each target rotated anchor box can be calculated based on the rotation angle of each target rotated anchor box, and then the rotation axis of each target rotated anchor box is rotated to a unified direction based on the calculated affine transformation matrix, so as to obtain the fourth rotated anchor box set.

[0112] It can be understood that the step B2 can be inputting the fourth rotated anchor box set into a region cropping layer, and cropping a corresponding feature region based on the frame size (such as frame width, frame height) of the target rotated anchor box through the region cropping layer, so as to cut out a feature block containing target vehicle information from the feature map to obtain the fifth rotated anchor box set.

[0113] It should be noted that the step B3 can be dividing the cropped feature region into fixed grid units (such as 7x7 or 14x14), and performing a max-pooling or average-pooling operation in each grid unit, so as to map the target rotated anchor boxes of different sizes into a fixed-size feature map to obtain the anchor box feature map.

[0114] The step 150 of updating the initialized vehicle detection model according to the anchor box feature map to obtain a trained vehicle detection model.

[0115] In the embodiment of the present application, step 150 can be identifying the target vehicle detection result corresponding to the anchor frame feature map through the target classifier, then determining a target loss value based on the target candidate anchor frame and the predicted category in the target vehicle detection result, and the real candidate frame and the real category in the full-color training image, then updating the parameters of the model through the back propagation algorithm based on the target loss value, and after several iterations, the trained vehicle detection model can be obtained. The specific number of iterations can be pre-set, or the training is considered to be completed when the test set reaches the accuracy requirement.

[0116] Specifically, the embodiment of the present application can first input the anchor frame feature map into the target classifier, which includes a full connection layer, a target classification layer and a frame regression layer. Specifically, the anchor frame feature map can be further feature extracted and information fused through the full connection layer; then the features output by the full connection layer are classified by the target classification layer to obtain the category of each target candidate anchor frame; then the target candidate anchor frame is regressed and offset predicted by the frame regression layer to obtain the target vehicle detection result. The regression and offset prediction is similar to the frame regression offset update in the foregoing step A32, which can be simply analogized.

[0117] Referring to Figure 2 The traffic state perception method of the vehicle detection model proposed in the embodiment of the present application comprises:

[0118] Step 210, acquiring traffic road network data and target single-time-phase remote sensing image, the target single-time-phase remote sensing image comprising a target panchromatic image and a target multispectral image;

[0119] In the embodiment of the present application, the traffic road network data can be the data related to the road on which the vehicle travels, which can specifically comprise road center, number of lanes, road type, speed limit information, etc. The target single-time-phase remote sensing image can be the remote sensing image corresponding to the position region of the traffic road network data, which can comprise a target panchromatic image and a target multispectral image.

[0120] Step 220, inputting the target panchromatic image into the trained vehicle detection model to perform vehicle target detection, and obtaining a vehicle target detection result;

[0121] In the embodiment of the present application, after obtaining the target panchromatic image about the traveling vehicle, the target panchromatic image can be input into the trained vehicle detection model to perform vehicle target detection, so as to obtain the vehicle target detection result output by the vehicle detection model, which records the candidate frame and the category of the several traveling vehicles in the target panchromatic image.

[0122] Step 230, performing vehicle moving speed analysis processing on the target multispectral image according to the vehicle target detection result, to obtain vehicle moving speed data;

[0123] In the embodiments of the present application, the candidate regions of the driving vehicle in different spectral bands of the target multispectral image can be extracted based on the candidate bounding boxes in the vehicle target detection result, and the moving speed of the driving vehicle (i.e., the vehicle moving speed data) can be determined based on the candidate regions in different spectral bands.

[0124] In some embodiments, the step 230 of performing vehicle moving speed analysis processing on the target multispectral image according to the vehicle target detection result to obtain vehicle moving speed data comprises:

[0125] C1, obtaining a first-band spectral image and a second-band spectral image according to the target multispectral image;

[0126] C2, performing first candidate region extraction on the first-band spectral image according to the vehicle target detection result to obtain a first-band candidate region, and performing second candidate region extraction on the second-band spectral image according to the vehicle target detection result to obtain a second-band candidate region;

[0127] In the embodiments of the present application, step C1 can be to obtain two images of different spectral bands of the target multispectral image, and they are denoted as the first-band spectral image and the second-band spectral image. Specifically, for a target single-time-phase remote sensing image collected by a WorldView-3 satellite, the first-band spectral image can be the imaging result of a multispectral sensor in a blue light band (0.450 μm-0.510 μm), and the second-band spectral image can be the imaging result of the multispectral sensor in a green light band (0.510 μm-0.580 μm). The present application examples are only for illustration, and the first-band spectral image and the second-band spectral image can also be images of other spectral bands, which are not limited to the present application.

[0128] It can be understood that for a driving vehicle, the first candidate region extraction in step C2 can be to extract the candidate region of the driving vehicle in the first-band spectral image based on the candidate bounding box of the driving vehicle in the vehicle target detection result, and the specific candidate region size can be determined by the vehicle size and the possible displacement range of the driving vehicle, so as to obtain the first-band candidate region. The content of the second-band candidate region is similar to that of the first-band candidate region, which can be simply inferred.

[0129] C3, performing region displacement matching on the second-band candidate region according to the first-band candidate region to obtain the vehicle moving speed data.

[0130] Further, the step C3, according to the first waveband candidate region, performs region displacement matching on the second waveband candidate region to obtain the vehicle moving speed data, includes:

[0131] C31, acquiring a spectral imaging time difference of the target multispectral image;

[0132] C32, according to the first waveband candidate region, performing template matching on the second waveband candidate region to obtain first vehicle pixel coordinate data and second vehicle pixel coordinate data, the first vehicle pixel coordinate data being pixel coordinate data of a traveling vehicle in the first waveband candidate region, and the second vehicle pixel coordinate data being pixel coordinate data of the traveling vehicle in the second waveband candidate region;

[0133] C33, according to the first vehicle pixel coordinate data, performing vehicle displacement calculation on the second vehicle pixel coordinate data to obtain vehicle displacement data;

[0134] C34, according to the spectral imaging time difference, performing vehicle moving speed calculation on the vehicle displacement data to obtain the vehicle moving speed data.

[0135] In the embodiments of the present application, first, a spectral imaging time difference of a target multispectral image can be acquired, specifically, a spectral imaging time difference between a first waveband spectral image and a second waveband spectral image can be acquired, and the spectral imaging time difference can be acquired based on satellite parameters selected.

[0136] It can be understood that, in the embodiments of the present application, the first waveband candidate region and the second waveband candidate region can be feature extracted and matched based on a template matching algorithm, so as to obtain the first vehicle pixel coordinate data and the second vehicle pixel coordinate data. Specifically, the step C32 can be that the first waveband candidate region is used as a template image, the second waveband candidate region is used as a to-be-matched image, a normalized cross-correlation (NCC) is used for template matching, then a matching position with the maximum NCC value is determined as the final second vehicle pixel coordinate data, and pixel coordinate data corresponding to the first waveband candidate region is directly determined as the final first vehicle pixel coordinate data, the first vehicle pixel coordinate data and the second vehicle pixel coordinate data corresponding to the same traveling vehicle.

[0137] It should be noted that, the step C33 can be that, based on a spatial resolution corresponding to the target multispectral image, pixel displacement of the traveling vehicle in the first waveband spectral image and the second waveband spectral image is converted into actual displacement of the vehicle on the ground (i.e., vehicle displacement data), wherein the spatial resolution corresponding to the target multispectral image can be acquired from parameters of a high spatial resolution satellite collecting the target multispectral image.

[0138] Specifically, the equivalent expression of the vehicle displacement data in the embodiment of the present application can be:

[0139]

[0140] wherein d i is the vehicle displacement data; r is the spatial resolution corresponding to the target multi-spectral image; Δx i is the pixel coordinate difference of the first vehicle pixel coordinate data and the second vehicle pixel coordinate data in the horizontal coordinate; Δy i is the pixel coordinate difference of the first vehicle pixel coordinate data and the second vehicle pixel coordinate data in the vertical coordinate.

[0141] It is worth mentioning that the step C34 can be based on the spectral imaging time difference between the first band spectral image and the second band spectral image and the vehicle displacement data to determine the vehicle speed data of the driving vehicle, specifically, it can be calculated based on the ratio of the vehicle displacement data and the spectral imaging time difference.

[0142] It needs to be supplemented that since there is a small imaging time difference in single-time-phase remote sensing image, for high-speed moving targets (such as vehicles), there will be a position shift (i.e. ghost phenomenon) in different spectral bands of multi-spectral image. In order to avoid this ghost phenomenon, the traditional traffic state perception method based on remote sensing image usually realizes vehicle detection based on multi-time-phase remote sensing image, and then realizes traffic state perception based on the detected multi-time-phase vehicle data. The embodiment of the present application realizes the analysis of the instantaneous speed of the driving vehicle based on the position shift caused by the spectral imaging time difference, which can reduce the cost of traffic state perception and improve the practicability of traffic state perception.

[0143] Step 240, according to the vehicle speed data and the vehicle target detection result, the traffic state of the traffic road network data is perceived, and the traffic state perception result is obtained.

[0144] In the embodiment of the present application, the traffic state of the traffic road network can be determined based on the vehicle position of each driving vehicle in the vehicle target detection result and the vehicle speed data of each driving vehicle, so as to obtain the traffic state perception result.

[0145] In some embodiments, the step 240, according to the vehicle speed data and the vehicle target detection result, the traffic state of the traffic road network data is perceived, and the traffic state perception result is obtained, including:

[0146] D1, according to the vehicle target detection result, the vehicle spatial matching of the traffic road network data is performed, and the vehicle spatial matching data is obtained;

[0147] D2, according to the vehicle moving speed data, traffic state evaluation is performed on the vehicle space matching data, and the traffic state perception result is obtained.

[0148] In the embodiment of the present application, step D1 can first generate a planar layer of each lane based on the road center line, the number of lanes, the road type, the speed limit information and other traffic road network information in the traffic road network data, so as to form a lane planar layer; then, based on the GIS spatial analysis technology, the vehicle position of each driving vehicle in the vehicle target detection result is superimposed on the corresponding spatial position of the lane planar layer, so as to determine the lane to which each driving vehicle belongs in the traffic road network, and the vehicle space matching data is obtained.

[0149] It can be understood that for a lane in the vehicle space matching data, step D2 can first calculate the average speed on the lane based on the vehicle moving speed data corresponding to each driving vehicle on the lane, and the equivalent expression of the average speed on the lane can be:

[0150]

[0151] wherein, is the average speed of the driving vehicle on the jth lane in the vehicle space matching data; n j is the total number of driving vehicles on the jth lane in the vehicle space matching data; v i is the vehicle moving speed data corresponding to the ith driving vehicle on the jth lane in the vehicle space matching data.

[0152] It is worth mentioning that the average speeds of the remaining lanes are the same, and after obtaining the average speed of each lane in the vehicle space matching data, the congestion level of each lane in the vehicle space matching data can be determined based on the speed limit of each lane and the preset congestion threshold, so as to obtain the traffic state perception result. Specifically, for a lane, the congestion coefficient of the lane can be determined based on the average speed of the lane and the corresponding speed limit, and the equivalent expression of the congestion coefficient can be:

[0153]

[0154] wherein, C j is the congestion coefficient of the jth lane; v limit,j is the speed limit of the jth lane.

[0155] It should be noted that after the congestion coefficients of each lane are determined, the size relationship between each lane coefficient and a preset congestion threshold can be compared respectively to obtain the congestion level of each lane, so as to obtain the traffic state perception result. The congestion threshold in the embodiment of the application can include 0.25, 0.5, 0.75, etc. The specific value of the congestion threshold can be set according to the actual situation. Specifically, for a certain lane, if the congestion coefficient is less than 0.25, it can be considered that the congestion level of the lane is smooth; or, if the congestion coefficient is greater than or equal to 0.25 and less than 0.5, it can be considered that the congestion level of the lane is basically smooth; or, if the congestion coefficient is greater than or equal to 0.25 and less than 0.75, it can be considered that the congestion level of the lane is congested; or, if the congestion coefficient is greater than or equal to 0.75, it can be considered that the congestion level of the lane is severely congested, and the rest of the lanes are the same. The traffic state perception result can be obtained by integrating the congestion levels corresponding to each lane.

[0156] A training system of a vehicle detection model according to an embodiment of the application is described in detail below with reference to the accompanying drawings.

[0157] With reference to Figure 3 The training system of the vehicle detection model according to the embodiment of the application comprises:

[0158] The first processing unit 101 is configured to obtain a panchromatic image in a single-phase remote sensing image about a target vehicle, and pre-process the panchromatic image to obtain a panchromatic training image.

[0159] The second processing unit 102 is configured to perform image feature extraction on the panchromatic training image to obtain panchromatic image features.

[0160] The third processing unit 103 is configured to perform candidate region extraction on the panchromatic image features to obtain a target rotated anchor box set, wherein the target rotated anchor box set comprises a plurality of target rotated anchor boxes corresponding to the target vehicle, and each target rotated anchor box has at least one different corresponding bounding box scale, aspect ratio or rotation angle.

[0161] The fourth processing unit 104 is configured to perform rotated region pooling on the target rotated anchor box set to obtain an anchor box feature map.

[0162] The fifth processing unit 105 is configured to perform parameter updating on an initialized vehicle detection model according to the anchor box feature map to obtain a trained vehicle detection model.

[0163] With reference to Figure 4 The embodiment of the application further provides an electronic device comprising:

[0164] at least one processor 201;

[0165] at least one memory 202, configured to store the at least one program;

[0166] The at least one program is executed by the at least one processor 201, so that the at least one processor 201 implements the method embodiments described above.

[0167] Similarly, it can be understood that the contents in the above method embodiments are all applicable to the present device embodiments, the present device embodiments specifically implement the functions same as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0168] The present application also provides a computer readable storage medium, which stores a program executable by the processor 201, and the program executable by the processor 201, when executed by the processor 201, is used for implementing the above method embodiments.

[0169] Similarly, the contents in the above method embodiments are all applicable to the present computer readable storage medium embodiments, the present computer readable storage medium embodiments specifically implement the functions same as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0170] In some alternative embodiments, the functions / operations mentioned in the block diagram can not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two blocks shown in succession can actually be executed substantially simultaneously or the blocks can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example, with the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of larger operations are independently executed.

[0171] Furthermore, although the present application is described in the context of functional modules, it is understood that one or more of the functions and / or features can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary to an understanding of the present application. Rather, the actual implementation is within the routine skill of those in the art, given the nature of the property, function and internal relationships of the various functional modules disclosed herein. Therefore, the present application is not limited to the specific embodiments described herein, but only by the claims, with equivalents of the scope of these claims being permitted. It is also understood that the specific concepts disclosed herein are intended to be illustrative only and not limiting of the scope of the application, which is to be given the full breadth of the appended claims and any equivalents thereof.

[0172] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0173] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer readable medium for use by an instruction execution system, device or equipment (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, device or equipment) or in conjunction with these instructions execution system, device or equipment. For the purpose of the present specification, "computer readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by an instruction execution system, device or equipment or in conjunction with these instruction execution system, device or equipment.

[0174] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.

[0175] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above described embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, known in the art, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), and / or the like.

[0176] In the above description of the present specification, the description referring to the terms "one embodiment", "another embodiment", or "certain embodiments" or the like means that a specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present specification. The illustrative expressions of the above terms do not necessarily refer to the same embodiment or example in the present specification. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.

[0177] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that changes, modifications, alternatives and variations to these embodiments can be made without departing from the principles and spirit of the application, the scope of which is defined in the appended claims and their equivalents.

[0178] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present application, and these equivalent modifications or replacements are included in the scope defined by the claims of the present application.

Claims

1. A traffic state perception method of a vehicle detection model, characterized by, The method comprises the following steps: acquiring traffic network data and target single-time remote sensing images, wherein the target single-time remote sensing images comprise a target panchromatic image and a target multispectral image; inputting the target panchromatic image into a trained vehicle detection model to perform vehicle target detection, and obtaining a vehicle target detection result; performing vehicle speed analysis processing on the target multispectral image according to the vehicle target detection result, and obtaining vehicle speed data; performing traffic state perception on the traffic network data according to the vehicle speed data and the vehicle target detection result, and obtaining a traffic state perception result; the step of performing vehicle speed analysis processing on the target multispectral image according to the vehicle target detection result, and obtaining vehicle speed data, comprises the following steps: acquiring a first waveband spectral image and a second waveband spectral image from the target multispectral image, wherein the spectral waveband of the first waveband spectral image is different from the spectral waveband of the second waveband spectral image; performing first candidate region extraction on the first waveband spectral image according to the vehicle target detection result, and obtaining a first waveband candidate region, and performing second candidate region extraction on the second waveband spectral image according to the vehicle target detection result, and obtaining a second waveband candidate region; performing region displacement matching on the second waveband candidate region according to the first waveband candidate region, and obtaining the vehicle speed data; wherein the trained vehicle detection model is obtained by the following steps: acquiring a panchromatic image in single-time remote sensing images of a target vehicle, and performing preprocessing on the panchromatic image to obtain a panchromatic training image; performing image feature extraction on the panchromatic training image to obtain panchromatic image features; performing candidate region extraction on the panchromatic image features to obtain a target rotating anchor box set, wherein the target rotating anchor box set comprises a plurality of target rotating anchor boxes corresponding to the target vehicle, and at least one of the following is different for each target rotating anchor box: a bounding box scale, an aspect ratio, or a rotation angle; performing rotating region pooling on the target rotating anchor box set to obtain an anchor box feature map; performing parameter updating on an initialized vehicle detection model according to the anchor box feature map to obtain the trained vehicle detection model.

2. The method of claim 1, wherein, the step of performing candidate region extraction on the panchromatic image features to obtain a target rotating anchor box set corresponding to the panchromatic image features, comprises the following steps: performing rotating anchor box generation processing on the panchromatic image features to obtain a first intermediate anchor box set, wherein the first intermediate anchor box set comprises a plurality of intermediate rotating anchor boxes corresponding to the panchromatic image features, and at least one of the following is different for each intermediate rotating anchor box: a region scale, an aspect ratio, or a rotation angle; performing anchor box feature mapping on the first intermediate anchor box set to obtain a second intermediate anchor box set; performing bounding box classification regression on the second intermediate anchor box set to obtain the target rotating anchor box set.

3. The method of claim 2, wherein, the step of performing bounding box classification regression on the second intermediate anchor box set to obtain the target rotating anchor box set, comprises the following steps: performing anchor box classification on the second intermediate anchor box set to obtain a third intermediate anchor box set, wherein the third intermediate anchor box set is a set of intermediate rotating anchor boxes corresponding to the target vehicle; The third intermediate anchor frame set is subjected to a bounding box regression offset update to obtain the target rotating anchor frame set.

4. The method of claim 3, wherein, The bounding box regression offset update of the third intermediate anchor frame set to obtain the target rotating anchor frame set comprises: The third intermediate anchor frame set is subjected to a bounding box offset regression prediction to obtain bounding box offset prediction data. The third intermediate anchor frame set is subjected to a bounding box offset update according to the bounding box offset prediction data to obtain the target rotating anchor frame set.

5. The method of claim 1, wherein, The target rotating anchor frame set is subjected to a rotating region pooling to obtain an anchor frame feature map, which comprises: The target rotating anchor frame set is subjected to a rotating coordinate transformation to obtain a fourth rotating anchor frame set, each target rotating anchor frame in the fourth rotating anchor frame set corresponding to the same rotating angle; The fourth rotating anchor frame set is subjected to a region scale cropping to obtain a fifth rotating anchor frame set; The fifth rotating anchor frame set is subjected to a feature map pooling mapping to obtain the anchor frame feature map.

6. The method of claim 1, wherein, The first waveband candidate region is subjected to a region displacement matching on the second waveband candidate region to obtain the vehicle moving speed data, which comprises: An imaging time difference of the target multispectral image is obtained; The first waveband candidate region is subjected to a template matching on the second waveband candidate region to obtain first vehicle pixel coordinate data and second vehicle pixel coordinate data, the first vehicle pixel coordinate data being pixel coordinate data of a traveling vehicle in the first waveband candidate region, and the second vehicle pixel coordinate data being pixel coordinate data of the traveling vehicle in the second waveband candidate region; The second vehicle pixel coordinate data is subjected to a vehicle displacement calculation according to the first vehicle pixel coordinate data to obtain vehicle displacement data; The vehicle displacement data is subjected to a vehicle moving speed calculation according to the imaging time difference to obtain the vehicle moving speed data.

7. The method of claim 1, wherein, The traffic state perception result is obtained by performing traffic state perception on the traffic road network data according to the vehicle moving speed data and the vehicle target detection result, which comprises: The vehicle spatial matching data is obtained by performing vehicle spatial matching on the traffic road network data according to the vehicle target detection result; The traffic state perception result is obtained by performing traffic state evaluation on the vehicle spatial matching data according to the vehicle moving speed data.

Citation Information

Patent Citations

  • Automatic collecting method of high-resolution satellite remote sensing traffic flow information

    CN102855759A

  • Pooling method and system for target rotation box detection based on deep learning

    CN112633265A

  • Image data feature extraction and defect identification method, device and system

    CN114049620A