A vehicle position detection device, method and storage medium

Through the combination of a deep aggregation network and a resolution adjustment module, the image resolution is adaptively adjusted, which solves the problem of information loss caused by image size uniformity in the prior art, and improves the accuracy of vehicle position detection and feature extraction capabilities.

CN114170428BActive Publication Date: 2025-07-22WUYI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111345997.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-12
Publication Date
2025-07-22
Estimated Expiration
2041-11-12

AI Technical Summary

Technical Problem

In vehicle detection and tracking, the existing neural networks have lost local information due to the unified image size processing, which affects the feature extraction ability and reduces the accuracy of classification results.

Method used

Deep aggregation network is adopted, and the resolution adjustment module of bilinear interpolation and bicubital interpolation algorithm is combined to adaptively adjust the image resolution, fusing spatial information and semantic information at different resolutions and levels.

Benefits of technology

It improves the accuracy of vehicle positioning and tracking, solves the shortcomings of image resolution adjustment in the existing network during upsampling and downsampling, and enhances feature extraction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170428B_ABST
    Figure CN114170428B_ABST
Patent Text Reader

Abstract

The present invention discloses a vehicle position detection device, which includes an image input unit, a feature extraction unit, and a classification unit; the feature extraction unit extracts image features of a vehicle image, the feature extraction unit includes a depth aggregation network, and the backbone network of the depth aggregation network includes a plurality of first feature extraction modules. The first feature extraction module includes a resolution adjustment module, and the resolution adjustment module adjusts the resolution of the image based on the bilinear interpolation algorithm and the bicubic interpolation algorithm; the classification unit obtains a classification result corresponding to the vehicle position according to the image features; fuses the spatial information and semantic information of vehicle images with different resolutions and different levels, and adaptively adjusts the resolution of the input image according to the task requirements, and can adjust the image resolution during the upsampling and downsampling training processes of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent detection, and particularly to a vehicle position detection device, method and storage medium. Background Art

[0002] Vehicle detection and tracking technologies play an important role in intelligent transportation systems and the field of driverless. Vehicle detection and tracking technologies often use neural networks for recognition. Currently, the feature extraction network in the neural network requires unified-size input for the input image. There are methods such as scaling and cropping to unify the size, but this may lose some local information of the vehicle image. These methods are applied before the feature extraction network and cannot be applied during the training process of the feature extraction network, which easily leads to the fact that the scaled size of the image is not the optimal size. And the resolution problem will affect the feature extraction ability of the network, thereby reducing the accuracy of the classification result. Summary of the Invention

[0003] An object of the present invention is to solve at least one of the technical problems existing in the prior art, and to provide a vehicle position detection device, method and storage medium.

[0004] The technical solution adopted by the present invention to solve its problems is as follows:

[0005] In a first aspect of the present invention, a vehicle position detection device includes:

[0006] An image input unit for inputting a vehicle image;

[0007] A feature extraction unit for extracting image features of the vehicle image. The feature extraction unit includes a depth aggregation network, and the backbone network of the depth aggregation network includes a plurality of first feature extraction modules connected in sequence. The output images of the plurality of first feature extraction modules have the same resolution. The first feature extraction module includes a resolution adjustment module, and the resolution adjustment module adjusts the resolution of the image based on the bilinear interpolation algorithm and the bicubic interpolation algorithm;

[0008] A classification unit for obtaining a classification result corresponding to the vehicle position according to the image features.

[0009] According to the first aspect of the present invention, the resolution adjustment module includes a first resolution adjustment module. The first feature extraction module includes a plurality of dimensionality reduction modules connected in sequence. The dimensionality reduction module includes a first dimensionality reduction module. The first dimensionality reduction module includes a max pooling layer, the first resolution adjustment module and a first convolutional layer. The max pooling layer and the first resolution adjustment module are connected in parallel to the first convolutional layer, and the output of the max pooling layer and the output of the first resolution adjustment module have the same size.

[0010] According to the first aspect of the present invention, the resolution adjustment module includes a second resolution adjustment module, the dimensionality reduction module further includes a second dimensionality reduction module, the second dimensionality reduction module includes a downsampling layer and a batch normalization layer, and the downsampling layer, the second resolution adjustment module, and the batch normalization layer are connected in sequence.

[0011] According to the first aspect of the present invention, the dimensionality reduction module further includes a third dimensionality reduction module and a fourth dimensionality reduction module. The output of the first dimensionality reduction module serves as the input of the second dimensionality reduction module, the output of the second dimensionality reduction module serves as the input of the third dimensionality reduction module, and the output of the second dimensionality reduction module and the output of the third dimensionality reduction module serve as the input of the fourth dimensionality reduction module.

[0012] According to the first aspect of the present invention, the resolution adjustment module includes a first adjustment branch, a second adjustment branch, and a third adjustment branch; the first adjustment branch includes a first adjustment module based on the bicubic interpolation algorithm, the second adjustment branch includes a second adjustment module based on the bilinear interpolation algorithm, and the third adjustment branch includes a third adjustment module based on the bilinear interpolation algorithm; the inputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are the same, and the outputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are aggregated.

[0013] According to the first aspect of the present invention, the second adjustment branch further includes a first convolutional block, a residual block, and a second convolutional block; the first convolutional block, the second adjustment module, the residual block, and the second convolutional block are connected in sequence.

[0014] According to the first aspect of the present invention, the resolution adjustment module includes a first aggregation layer and a second aggregation layer. The input of the first aggregation layer includes the output of the third adjustment branch, the output of the second adjustment module, and the output of the second convolutional block; the input of the second aggregation layer includes the output of the first aggregation layer and the output of the first adjustment branch, and the output of the second aggregation layer is the output of the resolution adjustment module.

[0015] According to the first aspect of the present invention, the depth aggregation network includes multiple feature extraction network layers, and the resolutions of the features corresponding to the multiple feature extraction network layers decrease layer by layer; each feature extraction network layer is connected to the backbone network.

[0016] In the second aspect of the present invention, a vehicle position detection method, the vehicle position detection method is applied to the vehicle position detection device as described in the first aspect of the present invention, and the vehicle position detection method includes the following steps:

[0017] Input a vehicle image through an image input unit;

[0018] Extract the image features of the vehicle image through the feature extraction unit;

[0019] Obtain a classification result corresponding to the vehicle position according to the image features through the classification unit.

[0020] In a third aspect of the present invention, a storage medium stores executable instructions, and when the executable instructions are executed by a processor, the vehicle position detection method described in the first aspect of the present invention is implemented.

[0021] The above solution has at least the following beneficial effects: Through the deep aggregation network, the vehicle images with different resolutions and different levels are fused in terms of spatial information and semantic information; The resolution adjustment module based on the bilinear interpolation algorithm and the bicubic interpolation algorithm adaptively adjusts the resolution of the input image according to the task requirements, solving the problem that the existing network can only adjust the image resolution during the upsampling and downsampling processes that cannot be trained in the deep network by cropping the image. This makes the vehicle position positioning and tracking results more accurate.

[0022] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention will be further described below in conjunction with the drawings and examples.

[0024] Figure 1 is a structural diagram of a vehicle position detection device according to an embodiment of the present invention;

[0025] Figure 2 is a structural diagram of the deep aggregation network;

[0026] Figure 3 is a structural diagram of the first feature extraction module;

[0027] Figure 4 is a structural diagram of the second feature extraction module;

[0028] Figure 5 is a structural diagram of the resolution adjustment module;

[0029] Figure 6 is a flowchart of a vehicle position detection method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0030] This part will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the drawings. The role of the drawings is to supplement the description of the text part of the specification, enabling people to intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it cannot be understood as a limitation on the protection scope of the present invention.

[0031] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0032] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, above, below, within, etc. are understood as including the present number. If the first and second are described only for the purpose of distinguishing technical features, they should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0033] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.

[0034] Referring to Figure 1 and Figure 2 , an embodiment of the first aspect of the present invention provides a vehicle position detection device.

[0035] The vehicle position detection device includes an image input unit, a feature extraction unit, and a classification unit.

[0036] Among them, the image input unit is used to input vehicle images; the feature extraction unit is used to extract the image features of the vehicle images. The feature extraction unit includes a depth aggregation network. The backbone network of the depth aggregation network includes a plurality of first feature extraction modules connected in sequence. The output images of the plurality of first feature extraction modules have the same resolution. The first feature extraction module includes a resolution adjustment module, and the resolution adjustment module adjusts the resolution of the image based on the bilinear interpolation algorithm and the bicubic interpolation algorithm; the classification unit is used to obtain a classification result corresponding to the vehicle position according to the image features.

[0037] In this embodiment, through the deep aggregation network, the vehicle images with different resolutions and different levels are fused in terms of spatial information and semantic information; through the resolution adjustment module based on the bilinear interpolation algorithm and the bicubic interpolation algorithm, the resolution of the input image is adaptively adjusted according to the task requirements, solving the problem that the existing network can only adjust the resolution of the image by cropping the image and cannot adjust the resolution in the upsampling and downsampling processes of the deep network training. This makes the vehicle position positioning and tracking results more accurate.

[0038] Referring to Figure 2 In some embodiments of the first aspect of the present invention, the deep aggregation network includes multiple layers of feature extraction network layers, and the resolution of the features corresponding to the multiple layers of feature extraction network layers decreases layer by layer; each layer of feature extraction network layer is connected to the backbone network.

[0039] The deep aggregation network merges the feature hierarchy in an iterative and hierarchical manner, making the network have higher accuracy and fewer parameters.

[0040] For the deep aggregation network, A1, A2, A3, A4, A5 are all the first feature extraction modules, and A1, A2, A3, A4, A5, B1, B2, C1, E are all aggregation nodes. A, A1, A2, A3, A4, A5, and E form the backbone network of the deep aggregation network. B, B1, and B2 are the first-layer feature extraction networks, C and C1 are the second-layer feature extraction networks, and D is the third-layer feature extraction network. The input image is input into A for 1 / 4 downsampling, and the output of A is the feature map after 1 / 4 downsampling. The resolutions of the feature maps from A to A1, A1 to A2, A2 to A3, A3 to A4, A4 to A5, and A5 to E are all the same, and are all 1 / 4 of the input image. The output of A is downsampled by 1 / 2 in B to obtain a 1 / 8 feature map, the output of B is downsampled by 1 / 2 in C to obtain a 1 / 16 feature map, and the output of C is downsampled by 1 / 2 in D to obtain a 1 / 32 feature map. The output of B is upsampled and input into A1 for feature fusion with the output of A, and the resolution of the upsampled feature map of the output of B is the same as the resolution of the output of A. And so on, in fact, E fuses the features of the feature maps downsampled by 1 / 4, 1 / 8, 1 / 16, and 1 / 32.

[0041] In the aggregation node, a combination of a convolutional layer, a batch normalization layer, and a ReLU activation function is adopted to simplify the network. The basic aggregation function is defined as follows: where σ is a non-linear activation function, and W and b are convolutional kernel parameters.

[0042] Residual connections are adopted between the aggregation nodes to ensure that the gradient does not vanish or explode in the deep network. Then the aggregation function including the residual connection is defined as follows:

[0043] Referring to Figure 5In certain embodiments of the first aspect of the present invention, the resolution adjustment module includes a first adjustment branch, a second adjustment branch, and a third adjustment branch; the first adjustment branch includes a first adjustment module based on the bicubic interpolation algorithm, the second adjustment branch includes a second adjustment module based on the bilinear interpolation algorithm, and the third adjustment branch includes a third adjustment module based on the bilinear interpolation algorithm; the inputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are the same, and the outputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are aggregated.

[0044] The bilinear interpolation algorithm is described below.

[0045] First, linear interpolation is performed in the x direction, and we have:

[0046]

[0047]

[0048] Then, linear interpolation is performed in the y direction, and we have:

[0049]

[0050] Then the result of the bilinear interpolation algorithm is expressed as:

[0051]

[0052] The bicubic interpolation algorithm is described below.

[0053] The bicubic interpolation algorithm can be expressed as:

[0054] In certain embodiments of the first aspect of the present invention, the second adjustment branch further includes a first convolutional block, a residual block, and a second convolutional block; the first convolutional block, the second adjustment module, the residual block, and the second convolutional block are connected in sequence.

[0055] The first convolutional block includes a convolutional layer, a ReLu activation function layer, a convolutional layer, a ReLu activation function layer, and a batch normalization layer connected in sequence.

[0056] The structure of the first convolutional layer of the first convolutional block is: the size of the convolutional kernel is set to 3*3, the stride is 1, the number of convolutional kernels is 16, and the padding is 1; the resolution of the output image is 1088*608.

[0057] The structure of the second convolutional layer of the first convolutional block is: the size of the convolutional kernel is set to 5*5, the stride is 1, the number of convolutional kernels is 32, and there is no padding at the edge; the resolution of the output image is 1080*600.

[0058] There are two residual blocks, and the structure of the residual block is: a convolutional layer, a batch normalization layer, a ReLu activation function layer, a convolutional layer, and a batch normalization layer. The convolutional kernel size of the convolutional layer is set to 3*3, the stride is 1, the number of convolutional kernels is 32, and the padding is 1.

[0059] The second convolutional block includes a convolutional layer and a ReLU activation function layer.

[0060] The structure of the convolutional layer of the second convolutional block is: the size of the convolutional kernel is set to 3*3, the stride is 1, the number of convolutional kernels is 32, and the padding is 1; the resolution of the output image is 13*10.

[0061] In certain embodiments of the first aspect of the present invention, the resolution adjustment module includes a first aggregation layer and a second aggregation layer. The input of the first aggregation layer includes the output of the third adjustment branch, the output of the second adjustment module, and the output of the second convolutional block; the input of the second aggregation layer includes the output of the first aggregation layer and the output of the first adjustment branch, and the output of the second aggregation layer is the output of the resolution adjustment module. The output of the first aggregation layer can be input to the second aggregation layer after passing through a depthwise separable convolutional layer. The structure of this depthwise separable convolutional layer is: using depthwise separable convolution, the size of the convolutional kernel is set to 3*3, the stride is 1, the number of convolutional kernels is 32, and the padding is 1; the resolution of the output image is 512*288.

[0062] After being processed by this resolution adjustment module, the image size can be effectively changed, and the accuracy of the trained task can be improved under the condition of resolution change.

[0063] Refer to Figure 3 , in certain embodiments of the first aspect of the present invention, the resolution adjustment module includes a first resolution adjustment module. The first feature extraction module includes a plurality of dimensionality reduction modules connected in sequence. The dimensionality reduction module includes a first dimensionality reduction module. The first dimensionality reduction module includes a max pooling layer, the first resolution adjustment module, and a first convolutional layer. The max pooling layer and the first resolution adjustment module are connected in parallel to the first convolutional layer, and the output sizes of the max pooling layer and the first resolution adjustment module are the same; a batch normalization layer is connected after the first convolutional layer. By the method of parallel superposition of the max pooling layer and the resolution adjustment module, the effect of enhancing the image can be achieved. The first convolutional layer plays a role in adjusting the number of output channels.

[0064] In certain embodiments of the first aspect of the present invention, the resolution adjustment module includes a second resolution adjustment module, the dimensionality reduction module further includes a second dimensionality reduction module, the second dimensionality reduction module includes a downsampling layer and a batch normalization layer, and the downsampling layer, the second resolution adjustment module, and the batch normalization layer are connected in sequence. After the batch normalization layer, there are also connected a ReLU activation function layer, a convolutional layer, and a batch normalization layer.

[0065] It should be noted that the structures of the first resolution adjustment module and the second resolution module are the same as the structure and function of the above-mentioned resolution adjustment module.

[0066] In certain embodiments of the first aspect of the present invention, the dimensionality reduction module further includes a third dimensionality reduction module and a fourth dimensionality reduction module. The output of the first dimensionality reduction module serves as the input of the second dimensionality reduction module, the output of the second dimensionality reduction module serves as the input of the third dimensionality reduction module, and the output of the second dimensionality reduction module and the output of the third dimensionality reduction module serve as the input of the fourth dimensionality reduction module.

[0067] Specifically, the third dimensionality reduction module includes a downsampling layer, a convolutional layer, a batch normalization layer, a ReLU activation function layer, a convolutional layer, and a batch normalization layer connected in sequence. The fourth dimensionality reduction module includes a root module, a convolutional layer, a batch normalization layer, and a ReLU activation function layer connected in sequence.

[0068] In this embodiment, the resolution adjustment module can be trained and learned in the deep network. Therefore, the first feature extraction module can better extract the feature information of the vehicle in the image, and finally enable the classification unit to generate accurate results. In vehicle detection and tracking, the network can better detect and locate a single vehicle and increase the possibility of trajectory matching.

[0069] Refer to Figure 4 , in another aspect, as an equivalent solution of the first feature extraction module, the structure of the second feature extraction module is as follows: a resolution adjustment module, a fifth dimensionality reduction module, a sixth dimensionality reduction module, a seventh dimensionality reduction module, and an eighth dimensionality reduction module connected in sequence.

[0070] The fifth dimensionality reduction module includes a resolution adjustment module, a convolutional layer, and a batch normalization layer connected in sequence.

[0071] The sixth dimensionality reduction module has the same structure as the second dimensionality reduction module, the seventh dimensionality reduction module has the same structure as the third dimensionality reduction module, and the eighth dimensionality reduction module has the same structure as the fourth dimensionality reduction module.

[0072] The input of the separate resolution adjustment module is the input of the second feature extraction module; the input of the fifth dimensionality reduction module is the output of the separate resolution adjustment module; the input of the sixth dimensionality reduction module is the output of the separate resolution adjustment module and the output of the fifth dimensionality reduction module; the input of the seventh dimensionality reduction module is the output of the sixth dimensionality reduction module, and the input of the eighth dimensionality reduction module is the output of the sixth dimensionality reduction module and the output of the seventh dimensionality reduction module. The output of the eighth dimensionality reduction module is the output of the second feature extraction module.

[0073] Referring to Figure 6 , an embodiment of the second aspect of the present invention provides a vehicle position detection method. The vehicle position detection method is applied to the vehicle position detection device described in the embodiment of the first aspect of the present invention.

[0074] The vehicle position detection method includes the following steps:

[0075] Step S100: Input a vehicle image through an image input unit;

[0076] Step S200: Extract image features of the vehicle image through a feature extraction unit;

[0077] Step S300: Obtain a classification result corresponding to the vehicle position according to the image features through a classification unit.

[0078] Through a feature extraction unit including a deep aggregation network, the vehicle images with different resolutions and different levels are fused in terms of spatial information and semantic information; through a resolution adjustment module based on the bilinear interpolation algorithm and the bicubic interpolation algorithm, the resolution of the input image is adaptively adjusted according to task requirements, solving the problem that the existing network can only adjust the image resolution during the upsampling and downsampling processes that cannot be trained in the deep network by cropping the image. This makes the vehicle position positioning and tracking results more accurate.

[0079] An embodiment of the third aspect of the present invention provides a storage medium. An executable instruction is stored in the storage medium, and when the executable instruction is executed by a processor, the vehicle position detection method described in the embodiment of the first aspect of the present invention is implemented.

[0080] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program units, or other data. The computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that the communication medium typically includes computer-readable instructions, data structures, program units, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0081] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, they should fall within the protection scope of the present invention.

Claims

1. A vehicle position detection device, characterized in that, Comprising: An image input unit for inputting vehicle images; A feature extraction unit for extracting image features of the vehicle images. The feature extraction unit includes a depth aggregation network, and the backbone network of the depth aggregation network includes a plurality of first feature extraction modules connected in sequence. The output images of the plurality of first feature extraction modules have the same resolution. The first feature extraction module includes a resolution adjustment module, and the resolution adjustment module adjusts the resolution of the image based on the bilinear interpolation algorithm and the bicubic interpolation algorithm; A classification unit for obtaining a classification result corresponding to the vehicle position according to the image features; Wherein, the resolution adjustment module includes a first resolution adjustment module. The first feature extraction module includes a plurality of dimensionality reduction modules connected in sequence. The dimensionality reduction module includes a first dimensionality reduction module, and the first dimensionality reduction module includes a max pooling layer, the first resolution adjustment module, and a first convolutional layer. The max pooling layer and the first resolution adjustment module are connected in parallel to the first convolutional layer, and the outputs of the max pooling layer and the first resolution adjustment module have the same size.

2. The vehicle position detection device according to claim 1, wherein The resolution adjustment module includes a second resolution adjustment module. The dimensionality reduction module further includes a second dimensionality reduction module, and the second dimensionality reduction module includes a downsampling layer and a batch normalization layer. The downsampling layer, the second resolution adjustment module, and the batch normalization layer are connected in sequence.

3. The vehicle position detection device according to claim 2, characterized in that, The dimensionality reduction module further includes a third dimensionality reduction module and a fourth dimensionality reduction module. The output of the first dimensionality reduction module serves as the input of the second dimensionality reduction module, the output of the second dimensionality reduction module serves as the input of the third dimensionality reduction module, and the outputs of the second dimensionality reduction module and the third dimensionality reduction module serve as the input of the fourth dimensionality reduction module.

4. A vehicle position detection device according to claim 2, characterized in that, The first resolution adjustment module and the second resolution adjustment module include a first adjustment branch, a second adjustment branch, and a third adjustment branch; the first adjustment branch includes a first adjustment module based on the bicubic interpolation algorithm, the second adjustment branch includes a second adjustment module based on the bilinear interpolation algorithm, and the third adjustment branch includes a third adjustment module based on the bilinear interpolation algorithm; the inputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are the same, and the outputs of the first adjustment branch, the second adjustment branch, and the third adjustment branch are aggregated.

5. A vehicle position detection device according to claim 4, characterized in that, The second adjustment branch further includes a first convolutional block, a residual block, and a second convolutional block; the first convolutional block, the second adjustment module, the residual block, and the second convolutional block are connected in sequence.

6. The vehicle position detection device according to claim 5, characterized in that The first resolution adjustment module and the second resolution adjustment module include a first aggregation layer and a second aggregation layer. The input of the first aggregation layer includes the output of the third adjustment branch, the output of the second adjustment module, and the output of the second convolutional block; the input of the second aggregation layer includes the output of the first aggregation layer and the output of the first adjustment branch, and the output of the second aggregation layer is the output of the resolution adjustment module.

7. A vehicle position detection device according to claim 1, wherein The deep aggregation network includes multiple feature extraction network layers, and the resolutions of the features corresponding to the multiple feature extraction network layers decrease layer by layer; each feature extraction network layer is connected to the backbone network.

8. A vehicle position detection method, characterized in that, The vehicle position detection method is applied to the vehicle position detection device according to any one of claims 1 to 7, and the vehicle position detection method includes the following steps: Input a vehicle image through an image input unit; Extract the image features of the vehicle image through a feature extraction unit; Obtain a classification result corresponding to the vehicle position according to the image features through a classification unit.

9. A storage medium, characterized in that, An executable instruction is stored in the storage medium, and when the executable instruction is executed by a processor, the vehicle position detection method according to claim 8 is implemented.