Target vehicle rotating frame detection and identification method and system

By configuring a feature extraction network and attention fusion unit with a spatial state model, combined with a feature pyramid network and a rotating target detection network, the accuracy and robustness of target vehicle detection in the prior art are solved, and the rapid and accurate identification of target vehicles of different angles and rotating states is achieved.

CN120014224AActive Publication Date: 2025-05-16TECH & ENG CENT FOR SPACE UTILIZATION CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510177109.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Existing target vehicle detection methods face challenges in accuracy and robustness, especially when dealing with scenarios with diverse vehicles, small target sizes, complex features and varied poses.

Method used

A feature extraction network configured with a spatial state model is used to generate the category and rectangular rotation frame of the target vehicle through multiple feature extraction of optical images and processing of the feature pyramid network, combined with the attention fusion unit and the rotation target detection network.

Benefits of technology

It improves the detection accuracy and robustness of the target vehicle, and can be suitable for detection scenarios of complex textures and rotating targets, meeting the needs of different usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014224A_ABST
    Figure CN120014224A_ABST
Patent Text Reader

Abstract

The invention provides a target vehicle rotating frame detection and identification method and system, and relates to the technical field of vehicle detection. The method provided by the invention comprises the following steps: acquiring an optical image, wherein the optical image comprises a to-be-detected target vehicle; inputting the optical image into a feature extraction network configured with a spatial state model, and outputting a plurality of first feature maps corresponding to the optical image; inputting the plurality of first feature maps into a feature pyramid network, and outputting a plurality of second feature maps; and inputting the plurality of second feature maps into a rotating target detection network, and outputting the category of the target vehicle and a rectangular rotating frame of the target vehicle on the optical image, the rectangular rotating frame being used for representing the position and the rotating angle of the target vehicle. Compared with a target detection method in the related technology, the method provided by the invention can improve the feature extraction capability as required, thereby improving the detection precision, and meeting the use requirements of a user in different use scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle detection, and in particular to a method and system for detecting and identifying a rotating frame of a target vehicle. Background Art

[0002] With the rapid development of intelligent technology, the detection of small rotating targets in optical images has gained wide attention in fields such as traffic monitoring and autonomous driving. It can also be solved by detecting target vehicles in images to determine the category of the target vehicle and its position information in the optical image. However, due to factors such as the diversity of vehicles, small target size, complex fine-grained features, and diverse posture changes, existing detection methods face many challenges in accuracy and robustness.

[0003] Traditional target vehicle detection methods mainly rely on manual feature extraction technology, which is usually targeted at specific detection tasks. The feature extraction process relies heavily on manual design and experience, which directly affects the detection effect and algorithm performance. These traditional methods have significant limitations: for example, the amount of data is limited, the manually extracted features are difficult to generalize, the algorithm has poor portability, the time complexity is high, and the robustness is insufficient in complex environments. Therefore, the scope of application is limited, and the detection accuracy is poor. Especially in the scenario of dealing with complex and changing environments, it is difficult to meet the user's usage needs, which limits its large-scale practical application.

[0004] Therefore, there is an urgent need for a target vehicle rotation frame detection and recognition method and system that can quickly and accurately identify target vehicles at different angles and rotation states to meet the user's usage needs in different usage scenarios. Summary of the invention

[0005] The embodiments of the present invention provide a method and system for detecting and identifying a target vehicle rotation frame, which can quickly and accurately identify target vehicles at different angles and rotation states, improve detection accuracy, and meet the user's usage needs in different usage scenarios.

[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0007] In a first aspect, a method for detecting and identifying a rotating frame of a target vehicle is provided, the method comprising: acquiring an optical image, the optical image comprising a target vehicle to be detected; inputting the optical image into a feature extraction network configured with a spatial state model, and outputting a plurality of first feature maps corresponding to the optical image, the plurality of first feature maps respectively having different numbers of channels and sizes; inputting the plurality of first feature maps into a feature pyramid network, and outputting a plurality of second feature maps, the plurality of second feature maps respectively having the same number of channels and different sizes; inputting the plurality of second feature maps into a rotating target detection network, and outputting a category of the target vehicle and a rectangular rotating frame of the target vehicle on the optical image, the rectangular rotating frame being used to characterize a position and a rotation angle of the target vehicle.

[0008] In a possible implementation of the first aspect, the feature extraction network includes an image processing module, a first feature extraction module, and multiple second feature extraction modules connected in series in sequence, the image processing module is used to process the optical image to obtain multiple image blocks corresponding to the optical image, the first feature extraction module is used to determine a first feature map corresponding to the optical image according to the multiple image blocks corresponding to the optical image, the first second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the first feature extraction module, and the nth second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the n-1th second feature extraction module, wherein each second feature extraction module and the first feature map determined by the first feature extraction module have different channel numbers and sizes, respectively, and n is an integer greater than or equal to 2.

[0009] In a possible implementation of the first aspect, the first feature extraction module includes a visual state space unit and an attention fusion unit, and the visual state space unit and the attention fusion unit are connected; the second feature extraction module includes a downsampling processing unit, a visual state space unit, and an attention fusion unit; the downsampling processing unit is connected to the visual state space unit, and the visual state space unit is connected to the attention fusion unit; the downsampling processing unit is used to: perform downsampling processing on the feature map; the visual state space unit is used to: perform cross-scan processing on the image block or the feature map to obtain feature information, map the feature information based on the hidden state variable to obtain mapping information corresponding to the feature information, and perform cross-merging processing on the mapping information to obtain the feature map;

[0010] The formula for mapping feature information based on hidden state variables is:

[0011] h t =e ΔA h t-1 +Δ B x t ;

[0012] yt =Ch t-1 +Dx t ;

[0013] Among them, h t is the hidden state variable of variable t; x t is the characteristic information of variable t; y t is the mapping information corresponding to the feature information of variable t; A∈R N×N ,B∈R N×1 ,C∈R 1×N and D∈R 1 is the weight matrix;

[0014] The attention fusion unit is used to: perform weighted processing on the feature map based on the channel attention weight to obtain a channel attention weighted feature map, and perform weighted processing on the channel attention weighted feature map based on the spatial attention weight to obtain a first feature map;

[0015] The formula for determining the channel attention weight is:

[0016]

[0017] Among them, M c F is the channel attention weight, σ() is the sigmoid activation function; W 0 and W 1 is the shared weight of the fully connected layer; is the feature map processed by average pooling, is the feature map processed by maximum pooling;

[0018] The formula for determining the spatial attention weight is:

[0019]

[0020] Among them, M s F is the spatial attention weight; f 7×7 It is a convolution layer with a convolution kernel size of 7×7; is the channel attention weighted feature map after average pooling, It is the channel attention weighted feature map after maximum pooling.

[0021] In a possible implementation manner of the first aspect, the rotation object detection network includes a candidate box generation module and a candidate box classification regression module;

[0022] The candidate frame generation module is used to: determine multiple horizontal anchor frames with different aspect ratios corresponding to each second feature map based on multiple second feature maps; determine a horizontal anchor frame where a target vehicle exists from the multiple horizontal anchor frames with different aspect ratios corresponding to each second feature map; determine multiple offset values ​​corresponding to each horizontal anchor frame where a target vehicle exists; and determine a tilted candidate frame corresponding to each horizontal anchor frame where a target vehicle exists based on the multiple offset values ​​corresponding to each horizontal anchor frame where a target vehicle exists.

[0023] In a possible implementation manner of the first aspect, the candidate box classification regression module includes a first fully connected layer and a second fully connected layer connected in parallel;

[0024] The first fully connected layer is used to determine the category of the target vehicle based on the inclined candidate box corresponding to the horizontal anchor box where the target vehicle exists;

[0025] The second fully connected layer is used to: determine a rectangular rotation box corresponding to a tilted candidate box corresponding to each horizontal anchor box where a target vehicle exists; based on a preset step size, map a rectangular rotation box corresponding to a tilted candidate box corresponding to each horizontal anchor box where a target vehicle exists to a second feature map, and obtain coordinate values ​​corresponding to a target vehicle region included in the second feature map; grid the target vehicle region included in the second feature map to obtain a feature map of a target size, determine a grid value of the feature map of the target size at each position in each channel, and generate a rectangular rotation box of the target vehicle on the optical image according to the grid value of the feature map of the target size at each position in each channel;

[0026] Among them, the second feature map includes the coordinate value (x r ,y r ,w r ,h r ,θ) is determined by:

[0027]

[0028] Among them, (x, y, w, h, θ) is the coordinate value corresponding to the rectangular rotation box; s is the preset step size;

[0029]

[0030] The grid value F of the feature map of the target size at the grid (i, j) of the cth channel c The formula for determining ′(i,j) is:

[0031]

[0032] F cis the feature map F of the target size in the cth channel, n is the number of samples in the grid, area(i,j) is the coordinate set in the grid (i,j), and R() is the rotation transformation.

[0033] In a possible implementation of the first aspect, before inputting multiple second feature maps into a rotating target detection network and outputting the category of the target vehicle and the rectangular rotation frame of the target vehicle on the optical image, the method further includes: acquiring a training image, the training image including multiple training anchor frames and multiple real anchor frames; determining a label corresponding to each training anchor frame, the label including a positive sample, a negative sample or an invalid sample; wherein the intersection-and-union ratio of the training anchor frame labeled as a positive sample with any real anchor frame is greater than or equal to a first threshold or the intersection-and-union ratio of the training anchor frame with the target real anchor frame is greater than the intersection-and-union ratio of other training anchor frames with the target real anchor frame; the target real anchor frame is one of the multiple real anchor frames; the intersection-and-union ratio of the training anchor frame labeled as a negative sample with any real anchor frame is less than a second threshold; the intersection-and-union ratio of the training anchor frame labeled as an invalid sample with any real anchor frame is greater than or equal to the second threshold and less than the first threshold; constructing a loss function; based on the loss function, iteratively training the rotating target detection network according to the training anchor frames labeled as positive samples and negative samples to obtain a trained rotating target detection network;

[0034] Loss function L 1 for:

[0035]

[0036] Among them, i is the index of the training anchor box; N is the total number of samples; F cls is the cross entropy loss for classification; F reg For the return of L 1 Smoothing losses; represents the true label of the i-th training anchor box; p i is the foreground probability value output by the first fully connected layer; It is the offset of the i-th training anchor box relative to the true anchor box under the midpoint offset representation.

[0037] The beneficial effects of the present invention are as follows: the method provided by the present invention extracts features from optical images through a feature extraction network configured with a spatial state model, and can obtain more fine-grained visual information, thereby effectively and accurately extracting key features in optical images, and can be applicable to detection scenes of complex textures and rotating targets, thereby improving the detection accuracy and robustness of target vehicles. In addition, the method provided by the present invention improves the ability of the feature extraction network to focus on key information areas in optical images through an attention fusion unit, reducing interference from background noise and irrelevant areas. The attention fusion unit can effectively improve the discriminability of feature representation through the fusion of channel attention mechanism and spatial attention mechanism, further improve the detection accuracy of target vehicles, and meet the use requirements of the use scenarios of rotating target positioning and boundary refinement. Finally, the method provided by the present invention realizes the generation and classification regression of candidate boxes respectively through a candidate box generation module and a candidate box classification regression module, which can effectively improve the accuracy. In general, compared with the target detection method in the related art, the method provided by the present invention can improve the feature extraction capability as needed, thereby improving the detection accuracy, and meeting the use requirements of users in different use scenarios.

[0038] In a second aspect, the present invention provides a target vehicle rotation frame detection and recognition system, the system comprising: an acquisition module, used to acquire an optical image, the optical image comprising a target vehicle to be detected; a feature extraction network, used to output a plurality of first feature maps corresponding to the optical image according to the input optical image, the plurality of first feature maps respectively having different numbers of channels and sizes; the feature extraction network is configured with a spatial state model; a feature pyramid network, used to output a plurality of second feature maps according to the input plurality of first feature maps, the plurality of second feature maps respectively having the same number of channels and different sizes; a rotating target detection network, used to output the category of the target vehicle and a rectangular rotation frame of the target vehicle on the optical image according to the input plurality of second feature maps, the rectangular rotation frame being used to characterize the position and rotation angle of the target vehicle.

[0039] In a possible implementation of the second aspect, the feature extraction network includes an image processing module, a first feature extraction module and multiple second feature extraction modules connected in series in sequence, the image processing module is used to process the optical image to obtain multiple image blocks corresponding to the optical image, the first feature extraction module is used to determine a first feature map corresponding to the optical image according to the multiple image blocks corresponding to the optical image, the first second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the first feature extraction module, and the nth second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the n-1th second feature extraction module, wherein each second feature extraction module and the first feature map determined by the first feature extraction module have different channel numbers and sizes, respectively, and n is an integer greater than or equal to 2.

[0040] In a possible implementation of the second aspect, the first feature extraction module includes a visual state space unit and an attention fusion unit, and the visual state space unit and the attention fusion unit are connected; the second feature extraction module includes a downsampling processing unit, a visual state space unit, and an attention fusion unit; the downsampling processing unit is connected to the visual state space unit, and the visual state space unit is connected to the attention fusion unit; the downsampling processing unit is used to: perform downsampling processing on the feature map; the visual state space unit is used to: perform cross-scan processing on the image block or the feature map to obtain feature information, map the feature information based on the hidden state variable to obtain mapping information corresponding to the feature information, and perform cross-merging processing on the mapping information to obtain the feature map;

[0041] The formula for mapping feature information based on hidden state variables is:

[0042]

[0043] y t =Ch t-1 +Dx t ;

[0044] Among them, h t is the hidden state variable of variable t; x t is the characteristic information of variable t; y t is the mapping information corresponding to the feature information of variable t; A∈R N×N ,B∈R N×1 ,C∈R 1×N and D∈R 1 is the weight matrix;

[0045] The attention fusion unit is used to: perform weighted processing on the feature map based on the channel attention weight to obtain a channel attention weighted feature map, and perform weighted processing on the channel attention weighted feature map based on the spatial attention weight to obtain a first feature map;

[0046] The formula for determining the channel attention weight is:

[0047]

[0048] Among them, M c F is the channel attention weight, σ() is the sigmoid activation function; W 0 and W 1 is the shared weight of the fully connected layer; is the feature map processed by average pooling, is the feature map processed by maximum pooling;

[0049] The formula for determining the spatial attention weight is:

[0050]

[0051] Among them, M s F is the spatial attention weight; f 7×7 It is a convolution layer with a convolution kernel size of 7×7; is the channel attention weighted feature map after average pooling, It is the channel attention weighted feature map after maximum pooling.

[0052] According to a third aspect, an electronic device is provided, comprising a memory and one or more processors; the memory is coupled to the processor; wherein the memory stores computer program code, the computer program code comprises computer instructions, and when the computer instructions are executed by the processor, the electronic device executes a method as in any implementation of the first aspect.

[0053] According to a fourth aspect, a computer-readable storage medium is provided, comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method in any implementation of the first aspect.

[0054] According to a fifth aspect, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to execute the method in any implementation of the first aspect.

[0055] It can be understood that the beneficial effects that can be achieved by the system of the second aspect, the electronic device of the third aspect, the computer-readable storage medium of the fourth aspect, and the computer program product of the fifth aspect provided above can be referred to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A schematic diagram of the hardware structure of an electronic device shown in an embodiment of the present invention;

[0057] Figure 2 A flow chart of a method for detecting and identifying a rotating frame of a target vehicle shown in an embodiment of the present invention;

[0058] Figure 3 A schematic diagram of the structure of a feature extraction network shown in an embodiment of the present invention;

[0059] Figure 4 A schematic diagram of the structure of a first feature extraction module and a second feature extraction module shown in an embodiment of the present invention;

[0060] Figure 5 A schematic diagram of the structure of a visual state space unit shown in an embodiment of the present invention;

[0061] Figure 6 A schematic diagram of the structure of a 2D selective scanning module shown in an embodiment of the present invention;

[0062] Figure 7 A schematic diagram of a midpoint offset representation method for a rotating target provided by an embodiment of the present invention;

[0063] Figure 8 The figure is a schematic diagram of the hardware structure of a detection system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The technical solution in the embodiment of the present invention will be described below in conjunction with the accompanying drawings in the embodiment of the present invention. Among them, in the description of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; the "or" in the present invention is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. And, in the description of the present invention, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items.

[0065] In addition, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not necessarily limit the difference.

[0066] Meanwhile, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being better or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0067] With the rapid development of intelligent technology, the detection of small rotating targets in optical images has gained wide attention in fields such as traffic monitoring and autonomous driving. It can also be solved by detecting target vehicles in images to determine the category of the target vehicle and its position information in the optical image. However, due to factors such as the diversity of vehicles, small target size, complex fine-grained features, and diverse posture changes, existing detection methods face many challenges in accuracy and robustness.

[0068] Traditional target vehicle detection methods mainly rely on manual feature extraction technology, which is usually targeted at specific detection tasks. The feature extraction process relies heavily on manual design and experience, which directly affects the detection effect and algorithm performance. These traditional methods have significant limitations: for example, the amount of data is limited, the manually extracted features are difficult to generalize, the algorithm has poor portability, the time complexity is high, and the robustness is insufficient in complex environments. Therefore, the scope of application is limited, and the detection accuracy is poor. Especially in the scenario of dealing with complex and changing environments, it is difficult to meet the user's usage needs, which limits its large-scale practical application.

[0069] Therefore, there is an urgent need for a target vehicle rotation frame detection and recognition method and system that can quickly and accurately identify target vehicles at different angles and rotation states to meet the user's usage needs in different usage scenarios.

[0070] In view of this, an embodiment of the present invention provides a method for detecting and identifying a rotating frame of a target vehicle, the method comprising: acquiring an optical image, the optical image comprising a target vehicle to be detected; inputting the optical image into a feature extraction network configured with a spatial state model, outputting a plurality of first feature maps corresponding to the optical image, the plurality of first feature maps respectively having different numbers of channels and sizes; inputting the plurality of first feature maps into a feature pyramid network, outputting a plurality of second feature maps, the plurality of second feature maps respectively having the same number of channels and different sizes; inputting the plurality of second feature maps into a rotating target detection network, outputting a category of the target vehicle and a rectangular rotating frame of the target vehicle on the optical image, the rectangular rotating frame being used to characterize a position and a rotation angle of the target vehicle.

[0071] The method provided by the present invention extracts features from optical images through a feature extraction network configured with a spatial state model, and can obtain more fine-grained visual information, thereby effectively and accurately extracting key features in optical images, and can be applicable to detection scenes of complex textures and rotating targets, thereby improving the detection accuracy and robustness of target vehicles. In addition, the method provided by the present invention improves the ability of the feature extraction network to focus on key information areas in optical images through an attention fusion unit, reducing the interference of background noise and irrelevant areas. The attention fusion unit can effectively improve the discriminability of feature representation through the fusion of channel attention mechanism and spatial attention mechanism, further improve the detection accuracy of target vehicles, and meet the use requirements of the use scenarios of rotating target positioning and boundary refinement. Finally, the method provided by the present invention realizes the generation and classification regression of candidate boxes respectively through a candidate box generation module and a candidate box classification regression module, which can effectively improve the accuracy. In general, compared with the target detection method in the related art, the method provided by the present invention can improve the feature extraction capability as needed, thereby improving the detection accuracy, and meeting the use requirements of users in different use scenarios.

[0072] In some embodiments, a target vehicle rotating frame detection and identification method provided by an embodiment of the present invention can be performed by a target vehicle rotating frame detection and identification system 100 (hereinafter referred to as the detection system 100). As an example, the detection system 100 can be any electronic device 200 with data processing capabilities, such as a general-purpose computer, a personal computer, a laptop computer, a switch or a tablet computer, etc. The specific implementation of the detection system 100 is not limited here.

[0073] Figure 1 The hardware structure diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device 200 includes a processor 210, a memory 220 and a communication interface 230.

[0074] The processor 210 may include one or more processing cores. The processor 210 uses various interfaces and lines to connect various parts in the electronic device 200, and executes various functions and processes data of the electronic device 200 by running or executing instructions, programs, code sets or instruction sets stored in the memory 220, and calling data stored in the memory 220. Optionally, the processor 210 can be implemented in at least one hardware form of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA).

[0075] The memory 220 may include a random access memory (RAl) or a read-only memory (ROL). Optionally, the memory 220 includes a non-transitory computer-readable storage medium (non-transitory colputer-readable storage lediul). The memory 220 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 220 may include a program storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as an image acquisition function, a data processing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.

[0076] The communication interface 230 is used to communicate with other devices, equipment or communication networks, such as data storage devices, image processing equipment or Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0077] In physical implementation, the above-mentioned components (such as processor 210, memory 220 and communication interface 230) can be components in the same device (such as a laptop). Alternatively, at least two of the components can be set in the same device, that is, as different components in one device, such as a deployment method similar to devices or components in a distributed system.

[0078] It is to be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 200. In other embodiments of the present invention, the electronic device 200 may include more or fewer components than those illustrated, or combine certain components, or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0079] A method for detecting and identifying a rotating frame of a target vehicle provided by an embodiment of the present invention is described below in conjunction with the accompanying drawings.

[0080] Figure 2 The present invention provides a flow chart of a method for detecting and identifying a rotating frame of a target vehicle. Optionally, the method may be Figure 1 The electronic device 200 shown is executed, that is, executed by the detection system 100. The method may include the following steps:

[0081] S1. Acquire an optical image, where the optical image includes a target vehicle to be detected.

[0082] S2. Input the optical image into a feature extraction network configured with a spatial state model, and output a plurality of first feature maps corresponding to the optical image.

[0083] Among them, the multiple first feature maps have different channel numbers and sizes respectively.

[0084] In some embodiments, see Figure 3 , Figure 3 The structural schematic diagram of a feature extraction network shown in an embodiment of the present invention, the feature extraction network 120 includes an image processing module 121, a first feature extraction module 122 and a plurality of second feature extraction modules 123 connected in series in sequence, the image processing module 121 is used to process the optical image to obtain a plurality of image blocks corresponding to the optical image, the first feature extraction module 122 is used to determine a first feature map (feature map C1) corresponding to the optical image according to the plurality of image blocks corresponding to the optical image, the first second feature extraction module 123 (which can also be understood as a second feature extraction module 123 connected to the first feature extraction module 122) is used to determine a first feature map (feature map C2) corresponding to the optical image according to the first feature map (feature map C1) determined by the first feature extraction module 122, the nth second feature extraction module 123 is used to determine a first feature map (feature map Cn) corresponding to the optical image according to the first feature map (feature map Cn-1) determined by the n-1th second feature extraction module 123, Figure 3In the embodiment, the second second feature extraction module 123 is used to determine the first feature map (feature map C3) corresponding to the optical image according to the first feature map (feature map C2) determined by the first second feature extraction module 123, and the third second feature extraction module 123 is used to determine the first feature map (feature map C4) corresponding to the optical image according to the first feature map (feature map C3) determined by the second second feature extraction module 123. The first feature map determined by each second feature extraction module 123 and the first feature extraction module 122 respectively have different channel numbers and sizes, and n is an integer greater than or equal to 2, that is, the feature maps C1-C4 respectively have different channel numbers and sizes.

[0085] In one possible implementation, see Figure 4 , Figure 4 A structural schematic diagram of a first feature extraction module and a second feature extraction module shown in an embodiment of the present invention, the first feature extraction module 122 includes a visual state space (Visual State Space, VSS) unit 1221 and an attention fusion unit 1222, and the visual state space unit 1221 and the attention fusion unit 1222 are connected; the second feature extraction module 123 includes a downsampling processing unit 1223, a visual state space unit 1221 and an attention fusion unit 1222; the downsampling processing unit 1223 is connected to the visual state space unit 1221, and the visual state space unit 1221 is connected to the attention fusion unit 1222; the downsampling processing unit 1223 is used to: downsample the feature map; the visual state space unit 1221 is used to: cross-scan the image block or the feature map to obtain feature information, map the feature information based on the hidden state variable to obtain the mapping information corresponding to the feature information, and cross-merge the mapping information to obtain the feature map.

[0086] The formula for mapping feature information based on hidden state variables is:

[0087]

[0088] y t =Ch t-1 +Dx t ;

[0089] Among them, h t is the hidden state variable of variable t; x t is the characteristic information of variable t; y t is the mapping information corresponding to the feature information of variable t; A∈R N×N ,B∈R N×1 ,C∈R 1×N and D∈R 1 is the weight matrix.

[0090] In one example, see Figure 5 , Figure 5 The schematic diagram of the structure of a visual state space unit shown in an embodiment of the present invention, the visual state space unit (hereinafter referred to as VSS unit) includes a normalization layer, a 2D selective scanning module, a normalization layer and a feedforward neural network, wherein the normalization layer normalizes the input of each layer to make the output distribution of each layer more stable, thereby alleviating the problem of gradient vanishing or gradient exploding, and the feedforward neural network is composed of a fully connected layer and a nonlinear activation function. Its main function is to transform and enhance the input features to extract higher-level feature representations. For the input data, residual connections are used to facilitate direct transmission of information and reduce the risk of gradient vanishing. For the 2D selective scanning module, it should be noted that in the fine-grained detection and recognition tasks of rotating targets, the targets usually have complex texture features, and the rotation angles are variable and the target size is small, which poses a great challenge to traditional visual feature extraction methods. To solve the above problems, the present invention introduces a state space model (State Space Model, SSM). SSM is derived from the Kalman filter and can be regarded as a linear time invariant (Linear Time Invariant, LTI) system. By hiding the state variable h(t)∈R N The input signal u(t)∈R is mapped to the output response y(t)∈R. Specifically, the continuous-time SSM can be represented by the following linear ordinary differential equation (ODE):

[0091] h′(t)=Ah(t)+Bu(t)

[0092] y(t)=Ch(t)+Du(t)

[0093] In order to integrate the continuous-time SSM into the deep learning model, it must be discretized first. Specifically, the time interval [t a ,t b ], the hidden state variable h(t) at t = t b The analytical solution at can be expressed as:

[0094]

[0095] By sampling with a time scale parameter Δ, that is, h(t b ) can be written as:

[0096]

[0097] Among them, [a, b] represents the corresponding discrete time step interval. It is worth noting that this formula is approximate to the result of the Zero Order Hold (ZOH) method. Although the discretization operation enables the SSM model to be applied to deep learning models of non-continuous data, its application in visual data faces great challenges. Unlike natural language data that contains time series information, visual data is inherently non-sequential and contains rich spatial information such as local texture and global structure, which is particularly prominent in the task of rotating target detection. In order to solve this problem, the method provided in an embodiment of the present invention reconstructs the SSM through a convolution operation, expands the convolution kernel from 1D to 2D, and implements it in the form of an outer product. And the input processing is performed through the Selective Scan (SS) method, and a 2D selective scanning module is used. The hardware structure of the 2D selective scanning module is as follows Figure 6 As shown, the 2D selective scanning module includes a cross scanning unit 610, a discretization unit 620 and a cross merging unit 630. In this way, the SSM can adapt to the feature extraction requirements of visual data and rotating targets while retaining its high efficiency and structural advantages. It can also be understood that by using the 2D selective scanning module, the method provided in the embodiment of the present invention is a spatial state model for feature extraction network configuration.

[0098] The attention fusion unit 1222 is used to: perform weighted processing on the feature map based on the channel attention weight to obtain a channel attention weighted feature map, and perform weighted processing on the channel attention weighted feature map based on the spatial attention weight to obtain a first feature map;

[0099] The formula for determining the channel attention weight is:

[0100]

[0101] Among them, M c F is the channel attention weight, σ() is the sigmoid activation function; W 0 and W 1 is the shared weight of the fully connected layer; is the feature map processed by average pooling, is the feature map processed by maximum pooling;

[0102] The formula for determining the spatial attention weight is:

[0103]

[0104] Among them, M s F is the spatial attention weight; f 7×7 It is a convolution layer with a convolution kernel size of 7×7; is the channel attention weighted feature map after average pooling, It is the channel attention weighted feature map after maximum pooling.

[0105] Specifically, the specific implementation of the attention fusion unit 1222 is explained below through an example. The attention fusion unit 1222 includes two parts: a channel attention fusion unit and a spatial attention fusion unit.

[0106] For the channel attention fusion unit, in the channel attention module, the feature map is first compressed in the spatial dimension to obtain a one-dimensional vector for subsequent processing. When compressing in the spatial dimension, the channel attention fusion unit not only uses the average pooling operation, but also introduces the maximum pooling operation to more comprehensively aggregate the spatial information of the feature map. The results of average pooling and maximum pooling are input into a shared network to further compress the spatial dimension and sum element by element to produce the final channel attention map. For a picture, the channel attention mainly focuses on the important content in the image: average pooling provides feedback to each pixel of the feature map, while maximum pooling mainly provides gradient feedback to the area with the strongest response in the feature map during back propagation. The channel attention fusion unit can effectively enhance the channel features of key information while suppressing secondary features, thereby improving the discriminability of features.

[0107] For the spatial attention fusion unit, the spatial attention module compresses the features F′ extracted by the channel attention module in the channel dimension, and performs average pooling and maximum pooling operations respectively to generate a spatial attention map. Specifically, the maximum pooling extracts the maximum value of each spatial position on the channel, and extracts a total of h×w times; the average pooling averages each spatial position, and the number of extractions is also h×w. Then, the above-extracted feature maps (the number of channels is 1) are merged to generate a feature map containing 2 channels. This method enables the model to focus more on the key information areas in the image, reduce the interference of background or irrelevant areas, and thus improve the overall detection accuracy.

[0108] S3. Input the multiple first feature maps into a feature pyramid network, and output the multiple second feature maps.

[0109] The multiple second feature maps have the same number of channels and different sizes.

[0110] In one example, the number of first feature maps is 4, and the number of second feature maps is 5.

[0111] S4. Input the plurality of second feature maps into a rotating target detection network, and output the category of the target vehicle and a rectangular rotation box of the target vehicle on the optical image, where the rectangular rotation box is used to characterize the position and rotation angle of the target vehicle.

[0112] In some embodiments, the rotated target detection network includes a candidate box generation module and a candidate box classification and regression module; the candidate box generation module is used to: determine multiple horizontal anchor boxes with different aspect ratios corresponding to each second feature map based on multiple second feature maps; determine a horizontal anchor box where a target vehicle exists from the multiple horizontal anchor boxes with different aspect ratios corresponding to each second feature map; determine multiple offset values ​​corresponding to each horizontal anchor box where a target vehicle exists; determine a tilted candidate box corresponding to each horizontal anchor box where a target vehicle exists based on the multiple offset values ​​corresponding to each horizontal anchor box where a target vehicle exists.

[0113] Optionally, the candidate box classification regression module includes a first fully connected layer and a second fully connected layer connected in parallel; the first fully connected layer is used to determine the category of the target vehicle according to the tilted candidate box corresponding to the horizontal anchor box where the target vehicle exists; the second fully connected layer is used to: determine the rectangular rotation box corresponding to each tilted candidate box corresponding to the horizontal anchor box where the target vehicle exists; based on a preset step size, map the rectangular rotation box corresponding to each tilted candidate box corresponding to the horizontal anchor box where the target vehicle exists to the second feature map, and obtain the coordinate value corresponding to the target vehicle area included in the second feature map; grid the target vehicle area included in the second feature map to obtain a feature map of the target size, determine the grid value of the feature map of the target size at each position of each channel, and generate a rectangular rotation box of the target vehicle on the optical image according to the grid value of the feature map of the target size at each position of each channel;

[0114] Among them, the second feature map includes the coordinate value (x r ,y r ,w r ,h r ,θ) is determined by:

[0115]

[0116] Among them, (x, y, w, h, θ) is the coordinate value corresponding to the rectangular rotation box; s is the preset step size;

[0117]

[0118] The grid value F of the feature map of the target size at the grid (i, j) of the cth channel c The formula for determining ′(i,j) is:

[0119]

[0120] Fc is the feature map F of the target size in the cth channel, n is the number of samples in the grid, area(i,j) is the coordinate set in the grid (i,j), and R() is the rotation transformation.

[0121] To facilitate understanding of the present solution, the specific implementation of the rotating target detection network provided by an embodiment of the present invention is explained below through an example.

[0122] Exemplarily, for the candidate box generation module, the traditional directed candidate box generation network (such as Rotated RPN and RoI Transformer) has a high computational cost and low efficiency when generating directed candidate boxes, especially when processing tilted targets, showing obvious performance bottlenecks. In response to these problems, the present invention proposes a tilted target representation method based on midpoint offset representation, and adopts a simplified designed directed candidate box generation network-Oriented RPN. Oriented RPN is a lightweight fully convolutional network that increases the output parameters of the RPN regression branch from the traditional 4 to 6 to adapt to the midpoint offset representation of tilted targets. This structure significantly improves the efficiency and speed of candidate box generation while reducing the number of parameters. For input images of any size, Oriented RPN can efficiently generate a series of tilted candidate boxes through a lightweight fully convolutional network. The specific implementation of the candidate box generation module is as follows:

[0123] The 5 second feature maps {P 2 ,P 3 ,P 4 ,P 5 ,P 6} as input, and add an identical head design (3×3 convolution layer and two 1×1 parallel convolution layers) to each layer. At each spatial position of all layers of each second feature map, three horizontal anchor boxes with different aspect ratios {1:2, 1:1, 2:1} are assigned. The horizontal anchor box is in {P 2 ,P 3 ,P 4 ,P 5 ,P 6 The corresponding image areas on the feature layer are {32 2 ,64 2 ,128 2 ,256 2 ,512 2}. Each horizontal anchor box is represented by a 4-dimensional vector a(a x ,a y ,a w ,a h ), where (a x,a y ) is the center coordinate of the anchor, a w and a w Respectively represent the width and height of the anchor.

[0124] One of the two 1×1 parallel convolutional layers is a classification branch, which is used to score the current horizontal anchor frame and determine whether the horizontal anchor frame contains the target vehicle to be detected. The other is a regression branch, which outputs the offset of the candidate frame relative to the horizontal anchor frame δ = (δ x ,δ y ,δ w ,δ h ,δ α ,δ β ). The candidate box generation module generates 3 candidate boxes at each position of each second feature map, so the regression branch at one position has 6×3 output values. For each candidate box, the tilted candidate box can be obtained by decoding the regression output. The decoding steps are as follows:

[0125]

[0126] Among them, (x, y) is the center coordinate of the predicted candidate box, w and h are the width and height of the bounding rectangle of the predicted candidate box, Δα and Δβ are the deviation values ​​relative to the midpoint of the upper and right boundaries of the bounding rectangle, respectively. Finally, the midpoint offset representation can be used to generate the four-vertex coordinate set v = (v 1 ,v 2 ,v 3 ,v 4 ). Among them, Δα is v 1 Relative to the upper midpoint of the horizontal bounding box The deviation value of v 2 Relative to the upper midpoint of the horizontal bounding box According to symmetry, -Δα and -Δβ are v 3 and v 4 Deviation value about the bottom edge and the midpoint of the coordinate. In summary, the coordinates of the four vertices of the tilted bounding box can be expressed as:

[0127]

[0128] Then, the candidate box generation module predicts the parameter values ​​(x, y, w, h) of the bounding rectangle to achieve the regression of each tilted candidate box and the determination of the midpoint deviation parameters (Δα, Δβ), see Figure 7 , a schematic diagram of a midpoint offset representation method for a rotating target provided by an embodiment of the present invention.

[0129] Exemplarily, for the candidate box classification regression module, the input of the candidate box classification regression module is 5 second feature maps {P 2 ,P 3 ,P 4 ,P 5 ,P 6} and the tilted candidate boxes corresponding to each horizontal anchor box with a target vehicle. For each tilted candidate box, a fixed-size feature vector is extracted from the corresponding feature map using the Rotated Roi Align algorithm. Each feature vector is input into the first fully connected layer and the second fully connected layer. The first fully connected layer outputs K+1 category probabilities including the background. The second fully connected layer regresses the candidate boxes of the K target categories and outputs the offset.

[0130] Specifically, the Rotated RoI Align algorithm is used to extract rotation-invariant features from each tilted candidate box. Since the candidate boxes generated by the directed candidate box generation network are usually parallelograms, four vertices v = {v 1 ,v 2 ,v 3 ,v 4}, and then adjust it to a rectangle by extending the short diagonal to the same length as the long diagonal. After that, we can get the (x, y, w, h, θ) of the inclined rectangle, where is the angle between the horizontal axis and the long side of the rectangle. Next, use the step size s to map the tilted rectangle (x, y, w, h, θ) to the feature map F and obtain (x r ,y r ,w r ,h r ,θ) represents the Rotated Roi.

[0131] In some embodiments, before the above S4, the method provided by the embodiment of the present invention further includes:

[0132] Acquire a training image, wherein the training image includes a plurality of training anchor frames and a plurality of real anchor frames;

[0133] Determine a label corresponding to each training anchor frame, wherein the label includes a positive sample, a negative sample or an invalid sample; wherein the intersection and union ratio of the training anchor frame with the label of the positive sample and any real anchor frame is greater than or equal to a first threshold or the intersection and union ratio of the training anchor frame with the target real anchor frame is greater than the intersection and union ratio of other training anchor frames and the target real anchor frame; the target real anchor frame is one of the multiple real anchor frames; the intersection and union ratio of the training anchor frame with the label of the negative sample and any real anchor frame is less than a second threshold; the intersection and union ratio of the training anchor frame with the label of the invalid sample and any real anchor frame is greater than or equal to the second threshold and less than the first threshold;

[0134] In one example, the first threshold is 0.7 and the second threshold is 0.3.

[0135] Construct loss function;

[0136] Based on the loss function, the rotation object detection network is iteratively trained according to the training anchor frames labeled as positive samples and negative samples to obtain a trained rotation object detection network;

[0137] The loss function L 1 for:

[0138]

[0139] Among them, i is the index of the training anchor box; N is the total number of samples; F cls is the cross entropy loss for classification; F reg For the return of L 1 Smoothing losses; represents the true label of the i-th training anchor box; p i is the foreground probability value output by the first fully connected layer; It is the offset of the i-th training anchor box relative to the true anchor box under the midpoint offset representation.

[0140] It should be understood that is the offset of the i-th anchor box relative to the true box under the midpoint offset representation, expressed as a parameterized 6-dimensional vector This vector comes from the candidate box regression branch, and the specific calculation formula is as follows:

[0141]

[0142] As can be seen from the above S1-S4, the method provided by the embodiment of the present invention can extract features from optical images through a feature extraction network configured with a spatial state model, so as to obtain more fine-grained visual information, and then effectively and accurately extract key features in optical images, and can be applicable to detection scenes of complex textures and rotating targets, thereby improving the detection accuracy and robustness of target vehicles. In addition, the method provided by the present invention improves the ability of the feature extraction network to focus on key information areas in optical images through an attention fusion unit, reducing the interference of background noise and irrelevant areas. The attention fusion unit can effectively improve the discriminability of feature representation through the fusion of channel attention mechanism and spatial attention mechanism, further improve the detection accuracy of target vehicles, and meet the use requirements of the use scenarios of rotating target positioning and boundary refinement. Finally, the method provided by the present invention realizes the generation and classification regression of candidate boxes respectively through a candidate box generation module and a candidate box classification regression module, which can effectively improve the accuracy. In general, compared with the target detection method in the related art, the method provided by the present invention can improve the feature extraction capability as needed, thereby improving the detection accuracy, and meeting the use requirements of users in different use scenarios.

[0143] In order to facilitate understanding of the beneficial effects of the present solution, the beneficial effects of the method provided by the embodiment of the present invention are explained below based on comparative experiments. In one example, the detection system 100 first obtains 331 simulated images of vehicles of 22 categories, totaling 9505 instances. Among them, in order to improve the similarity between the simulated image and the real image, the present invention comprehensively simulates different scenes and weather conditions during the simulation process. Specifically, the generation of the simulated image takes into account various environmental factors, including but not limited to changes in illumination, weather influences, background complexity, etc., so that the simulated data can more realistically reflect the characteristics of the rotating target under different conditions, thereby enhancing the robustness and generalization ability of the model in practical applications. The detection device then uses the average recall rate (mRecall) and the average precision (mAP, mean Average Precision) as evaluation indicators of the method provided by the embodiment of the present invention.

[0144] Specifically, the average recall rate is a key indicator for evaluating the model's missed detection in the target detection task. The recall rate measures the proportion of the real targets detected by the model to all actual targets. Its formula is:

[0145]

[0146] Among them, True Positives refers to the objects correctly detected by the model, that is, the model predicts that they are objects, and these predictions completely match the real annotations (ground truth). False Negatives refers to the model failing to detect the actual objects, that is, the model missed an object. That is, the object exists but the model did not make a correct prediction.

[0147] In multi-category detection tasks, the average recall rate (mRecall) is the average of the recall rates of all categories. The level of mRecall reflects the model's coverage of target detection, especially when the background is complex, the target size is small, or the target is rotated, whether the model can effectively identify and recall various types of targets. If mRecall is high, it means that the model can detect targets in multiple categories better and avoid missed detection.

[0148] Average precision is one of the most commonly used evaluation indicators in the field of object detection, and is used to comprehensively evaluate the performance of the model in multi-category tasks. For each category, precision and recall are calculated at different detection thresholds. Precision measures the proportion of the boxes that are actually objects in the boxes that are judged as objects by the model. The calculation formula is:

[0149]

[0150] False Positives refers to the model mistakenly predicting that the object is a target, but there is no object at that location. In simple terms, the model mistakenly generates a detection box in the background or non-target area.

[0151] For each category, we first calculate the corresponding precision values ​​under different recall rates, and then calculate the area under the precision-recall curve (AP) based on these values. AP is a comprehensive measure of the balance between precision and recall, which comprehensively considers different types of errors that may occur in target detection (such as missed detection and false detection).

[0152] On this basis, mAP is the average of the AP values ​​of all categories, indicating the detection accuracy of the model on all categories. The higher the mAP value, the stronger the model's ability to identify targets of each category in the entire detection task. For fine-grained rotation target detection tasks, mAP can better reflect the comprehensive performance of the model in dealing with different rotation angles, complex backgrounds, and target size differences.

[0153] The present invention conducts an experimental comparison of the optical fine-grained rotation target detection accuracy of five rotation target detection schemes of technology 1, technology 2, technology 3, technology 4, and technology 5 in the related art. Technology 1 converts the angle prediction problem of the rotation box into a probability distribution prediction problem. The angle distribution of the predicted box and the real box is represented by Gaussian distribution, and the model is optimized by minimizing the KL divergence (Kullback-Leibler Divergence) between the two, thereby avoiding the angle periodicity problem. Technology 2 is a target detection method based on pixel-by-pixel prediction, which is used for horizontal box detection, supports rotation target detection, and uses angle encoding to avoid the angle periodicity problem. Technology 3 represents the rotation rectangle box by vertex. Each rotation rectangle of the target box is represented as the position coordinates of 4 vertices. Technology 4 is an improved candidate region generation method that can generate rotation candidate regions. On the basis of the ordinary RPN, a rotation transformer module (Rotation Transformer) is added to convert the horizontal candidate region into a rotation candidate region, so that the generated candidate region is more in line with the rotation target and improves the detection accuracy. Technology 5 is a rotation equivariant network for rotating target detection. Rotational convolution is used in feature extraction. Through this convolution operation, the convolution kernel can adapt to the rotation of the target in the image, so that the network can still extract effective features under the rotation of the target. The specific experimental results are shown in Tables 1, 2 and 3. Table 1 is a data table comparing the optical fine-grained rotation target detection accuracy of different rotating target detection networks, Table 2 is a data table of the recall of each category of each experiment of the optical fine-grained rotation target detection accuracy of different rotating target detection networks, and Table 3 is a data table of the AP of each category of each experiment of the optical fine-grained rotation target detection accuracy of different rotating target detection networks.

[0154] Table 1

[0155] technology mRecall mAP Technology 1 0.424 0.281 Technique 2 0.468 0.282 Technique 3 0.413 0.368 Technique 4 0.557 0.448 Technique 5 0.590 0.540 The present invention 0.712 0.649

[0156] Table 2

[0157]

[0158] Table 3

[0159]

[0160] It can be seen from Tables 1 to 3 above that the method provided by the embodiment of the present invention is significantly superior to the other five rotating target detection schemes in the related art in terms of the accuracy index of the optical fine-grained rotation simulation data set compared with the five methods in the related art. This shows that the feature extraction network configured with the spatial state model can process images more efficiently in the feature extraction stage and fully mine the key features in the image. At the same time, the two-stage rotating target detection network can accurately generate a rotating detection box (rectangular rotating box) based on the extracted high-quality features, thereby improving the ability to accurately identify fine-grained vehicle targets, especially in complex backgrounds and with more details.

[0161] Furthermore, in order to verify the role of each step and module in the technical solution of the present invention, a set of ablation experiments was designed, and the influence of each step on the accuracy of optical fine-grained rotation target detection was analyzed by replacing the feature extraction network, adding an attention fusion module, etc. Tables 4, 5 and 6 show the overall and category accuracy indicators on the optical fine-grained rotation simulation data set when different feature extraction networks (ResNet, Swin Transformer, VMamba and the feature extraction network configured with a spatial state model provided by the present invention) are used while keeping the rotation target detection network unchanged. Among them, Table 4 is the total result of the optical fine-grained rotation target detection accuracy ablation experiment, Table 5 is the recall of each category of the optical fine-grained rotation target detection accuracy experiment of different feature extraction networks, and Table 6 is the AP of each category of the optical fine-grained rotation target detection accuracy experiment of different feature extraction networks.

[0162] Table 4

[0163] Feature extraction network mRecall mAP ResNet 0.622 0.543 Swin Transformer 0.696 0.599 VMamba 0.661 0.607 The present invention 0.712 0.649

[0164] Table 5

[0165]

[0166]

[0167] Table 6

[0168]

[0169] As can be seen from Tables 4 to 6 above, the feature extraction network configured with a spatial state model provided by the embodiment of the present invention can extract richer and more detailed features compared to ResNet and Swin Transformer commonly used in existing rotation target detection, thereby showing better performance in the rotation target detection network. This shows that the feature extraction network configured with a spatial state model can more effectively capture key visual information when processing fine-grained rotation targets, thereby improving detection accuracy and robustness.

[0170] The above mainly introduces the scheme of the embodiment of the present invention from the perspective of the method. It is understandable that, in order to realize the above functions, the detection system 100 includes at least one of the hardware structure and software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiment of the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiment of the present invention.

[0171] The embodiment of the present invention can divide the detection system 100 into functional units according to the above method example. For example, the detection system 100 can be divided into functional units corresponding to various functions, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0172] For example, Figure 8 The hardware structure diagram of a detection system provided by an embodiment of the present invention is shown. The detection system 100 includes: an acquisition module 110, which is used to acquire an optical image, and the optical image includes a target vehicle to be detected; a feature extraction network 120, which is used to output multiple first feature maps corresponding to the optical image according to the input optical image, and the multiple first feature maps have different channel numbers and sizes; the feature extraction network 120 is configured with a spatial state model; a feature pyramid network 130, which is used to output multiple second feature maps according to the input multiple first feature maps, and the multiple second feature maps have the same channel number and different sizes; a rotation target detection network 140, which is used to output the category of the target vehicle and the rectangular rotation box of the target vehicle on the optical image according to the input multiple second feature maps, and the rectangular rotation box is used to characterize the position and rotation angle of the target vehicle.

[0173] It should be understood that the specific description of the above optional methods can refer to the above method embodiments, which will not be repeated here. In addition, the explanation of any detection system 100 provided above and the description of the beneficial effects can refer to the above corresponding method embodiments, which will not be repeated here.

[0174] The embodiment of the present invention further provides a computer-readable storage medium, in which at least one computer instruction is stored, and the at least one computer instruction is loaded and executed by a processor to implement the methods of the above embodiments. For the explanation of the relevant contents and the description of the beneficial effects in any of the above-mentioned computer-readable storage media, reference can be made to the above-mentioned corresponding embodiments, which will not be repeated here.

[0175] The embodiment of the present invention further provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above detection system 100. Optionally, the functions supported by the chip can be referred to above and will not be described in detail here.

[0176] Those skilled in the art will appreciate that all or part of the steps of the above embodiments can be implemented by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, a specific circuit structure (application specific integrated circuit, ASIC), a microprocessor (digital signal processor, DSP), a field programmable gate array (field prograllable gatearray, FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.

[0177] The embodiment of the present invention also provides a computer program product including instructions, when the instructions are run on a computer, the computer executes any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD), etc.

[0178] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiments of the present invention, such as but not limited to the above-mentioned memory, computer-readable storage medium and communication chip, etc., all have non-transitory. Those skilled in the art should be aware that in one or more of the above-mentioned examples, the functions described in the embodiments of the present invention can be implemented with hardware, software, firmware or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. Computer-readable storage media include computer storage media and communication media, wherein the communication medium includes any medium that is convenient for transmitting a computer program from one place to another. The storage medium can be any available medium that a general or special-purpose computer can access.

[0179] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A method for detecting and identifying a rotating frame of a target vehicle, characterized in that: The method comprises: Acquiring an optical image, wherein the optical image includes a target vehicle to be detected; Inputting the optical image into a feature extraction network configured with a spatial state model, and outputting a plurality of first feature maps corresponding to the optical image, wherein the plurality of first feature maps respectively have different numbers of channels and sizes; Inputting the plurality of first feature maps into a feature pyramid network, and outputting a plurality of second feature maps, wherein the plurality of second feature maps respectively have the same number of channels and different sizes; The multiple second feature maps are input into a rotating target detection network, and the category of the target vehicle and a rectangular rotation box of the target vehicle on the optical image are output, where the rectangular rotation box is used to characterize the position and rotation angle of the target vehicle.

2. The method according to claim 1, characterized in that The feature extraction network includes an image processing module, a first feature extraction module and multiple second feature extraction modules connected in series in sequence, the image processing module is used to process the optical image to obtain multiple image blocks corresponding to the optical image, the first feature extraction module is used to determine a first feature map corresponding to the optical image according to the multiple image blocks corresponding to the optical image, the first second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the first feature extraction module, and the nth second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the n-1th second feature extraction module, wherein each second feature extraction module and the first feature map determined by the first feature extraction module have different channel numbers and sizes, respectively, and n is an integer greater than or equal to 2.

3. The method according to claim 2, characterized in that The first feature extraction module includes a visual state space unit and an attention fusion unit, and the visual state space unit is connected to the attention fusion unit; the second feature extraction module includes a downsampling processing unit, a visual state space unit and an attention fusion unit; the downsampling processing unit is connected to the visual state space unit, and the visual state space unit is connected to the attention fusion unit; the downsampling processing unit is used to: perform downsampling processing on the feature map; the visual state space unit is used to: perform cross-scan processing on the image block or the feature map to obtain feature information, map the feature information based on the hidden state variable to obtain mapping information corresponding to the feature information, and perform cross-merging processing on the mapping information to obtain the feature map; The formula for mapping feature information based on hidden state variables is: y t =Ch t-1 +Dx t ; Among them, h t is the hidden state variable of variable t; x t is the characteristic information of variable t; y t is the mapping information corresponding to the feature information of variable t; A∈R N×N ,B∈R N×1 ,C∈R 1×N and D∈R 1 is the weight matrix; The attention fusion unit is used to: perform weighted processing on the feature map based on the channel attention weight to obtain a channel attention weighted feature map, and perform weighted processing on the channel attention weighted feature map based on the spatial attention weight to obtain a first feature map; The formula for determining the channel attention weight is: Among them, M c F is the channel attention weight, σ() is the sigmoid activation function; W0 and W1 are the shared weights of the fully connected layer; is the feature map processed by average pooling, is the feature map processed by maximum pooling; The formula for determining the spatial attention weight is: Among them, M s F is the spatial attention weight; f 7×7 It is a convolution layer with a convolution kernel size of 7×7; is the channel attention weighted feature map after average pooling, It is the channel attention weighted feature map after maximum pooling.

4. The method according to claim 3, characterized in that: The rotating object detection network includes a candidate box generation module and a candidate box classification regression module; The candidate frame generation module is used to: determine a plurality of horizontal anchor frames with different aspect ratios corresponding to each second feature map according to the plurality of second feature maps; Determine a horizontal anchor frame where the target vehicle exists from a plurality of horizontal anchor frames with different aspect ratios corresponding to each second feature map; Determine multiple offset values ​​corresponding to each horizontal anchor frame where the target vehicle exists; and determine a tilted candidate frame corresponding to each horizontal anchor frame where the target vehicle exists according to the multiple offset values ​​corresponding to each horizontal anchor frame where the target vehicle exists.

5. The method according to claim 4, characterized in that The candidate frame classification regression module includes a first fully connected layer and a second fully connected layer connected in parallel; The first fully connected layer is used to determine the category of the target vehicle according to the inclined candidate box corresponding to the horizontal anchor box where the target vehicle exists; The second fully connected layer is used to: determine a rectangular rotation box corresponding to a tilted candidate box corresponding to each horizontal anchor box where a target vehicle exists; based on a preset step size, map a rectangular rotation box corresponding to a tilted candidate box corresponding to each horizontal anchor box where a target vehicle exists to a second feature map, and obtain coordinate values ​​corresponding to a target vehicle region included in the second feature map; grid the target vehicle region included in the second feature map to obtain a feature map of a target size, determine a grid value of the feature map of the target size at each position in each channel, and generate a rectangular rotation box of the target vehicle on the optical image according to the grid value of the feature map of the target size at each position in each channel; Among them, the second feature map includes the coordinate value (x r ,y r ,w r ,h r ,θ) is determined by: Among them, (x, y, w, h, θ) is the coordinate value corresponding to the rectangular rotation box; s is the preset step size; The grid value F of the feature map of the target size at the grid (i, j) of the cth channel c ′ The formula for determining (i,j) is: F c is the feature map F of the target size in the cth channel, n is the number of samples in the grid, area(i,j) is the coordinate set in the grid (i,j), and R() is the rotation transformation.

6. The method according to claim 5, characterized in that Before inputting the plurality of second feature maps into a rotating target detection network and outputting the category of the target vehicle and a rectangular rotation box of the target vehicle on the optical image, the method further includes: Acquire a training image, wherein the training image includes a plurality of training anchor frames and a plurality of real anchor frames; Determine a label corresponding to each training anchor frame, wherein the label includes a positive sample, a negative sample or an invalid sample; wherein the intersection and union ratio of the training anchor frame with the label of the positive sample and any real anchor frame is greater than or equal to a first threshold or the intersection and union ratio of the training anchor frame with the target real anchor frame is greater than the intersection and union ratio of other training anchor frames and the target real anchor frame; the target real anchor frame is one of the multiple real anchor frames; the intersection and union ratio of the training anchor frame with the label of the negative sample and any real anchor frame is less than a second threshold; the intersection and union ratio of the training anchor frame with the label of the invalid sample and any real anchor frame is greater than or equal to the second threshold and less than the first threshold; Construct loss function; Based on the loss function, the rotation object detection network is iteratively trained according to the training anchor frames labeled as positive samples and negative samples to obtain a trained rotation object detection network; The loss function L1 is: Among them, i is the index of the training anchor box; N is the total number of samples; F cls is the cross entropy loss for classification; F reg is the L1 smoothing loss for regression; represents the true label of the i-th training anchor box; p i is the foreground probability value output by the first fully connected layer; It is the offset of the i-th training anchor box relative to the true anchor box under the midpoint offset representation.

7. A target vehicle rotating frame detection and recognition system, characterized in that: The system comprises: An acquisition module, used for acquiring an optical image, wherein the optical image includes a target vehicle to be detected; A feature extraction network, configured to output a plurality of first feature maps corresponding to the optical image according to the input optical image, wherein the plurality of first feature maps respectively have different numbers of channels and sizes; the feature extraction network is configured with a spatial state model; A feature pyramid network is used to output a plurality of second feature maps according to the plurality of first feature maps inputted, wherein the plurality of second feature maps respectively have the same number of channels and different sizes; The rotating target detection network is used to output the category of the target vehicle and a rectangular rotation box of the target vehicle on the optical image according to the input multiple second feature maps, and the rectangular rotation box is used to characterize the position and rotation angle of the target vehicle.

8. The system according to claim 7, characterized in that The feature extraction network includes an image processing module, a first feature extraction module and multiple second feature extraction modules connected in series in sequence, the image processing module is used to process the optical image to obtain multiple image blocks corresponding to the optical image, the first feature extraction module is used to determine a first feature map corresponding to the optical image according to the multiple image blocks corresponding to the optical image, the first second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the first feature extraction module, and the nth second feature extraction module is used to determine the first feature map corresponding to the optical image according to the first feature map determined by the n-1th second feature extraction module, wherein each second feature extraction module and the first feature map determined by the first feature extraction module have different channel numbers and sizes, respectively, and n is an integer greater than or equal to 2.

9. The system according to claim 8, characterized in that The first feature extraction module includes a visual state space unit and an attention fusion unit, and the visual state space unit is connected to the attention fusion unit; the second feature extraction module includes a downsampling processing unit, a visual state space unit and an attention fusion unit; the downsampling processing unit is connected to the visual state space unit, and the visual state space unit is connected to the attention fusion unit; the downsampling processing unit is used to: perform downsampling processing on the feature map; the visual state space unit is used to: perform cross-scan processing on the image block or the feature map to obtain feature information, map the feature information based on the hidden state variable to obtain mapping information corresponding to the feature information, and perform cross-merging processing on the mapping information to obtain the feature map; The formula for mapping feature information based on hidden state variables is: y t =Ch t-1 +Dx t ; Among them, h t is the hidden state variable of variable t; x t is the characteristic information of variable t; y t is the mapping information corresponding to the feature information of variable t; A∈R N×N ,B∈R N×1 ,C∈R 1×N and D∈R 1 is the weight matrix; The attention fusion unit is used to: perform weighted processing on the feature map based on the channel attention weight to obtain a channel attention weighted feature map, and perform weighted processing on the channel attention weighted feature map based on the spatial attention weight to obtain a first feature map; The formula for determining the channel attention weight is: Among them, M c F is the channel attention weight, σ() is the sigmoid activation function; W0 and W1 are the shared weights of the fully connected layer; is the feature map processed by average pooling, is the feature map processed by maximum pooling; The formula for determining the spatial attention weight is: Among them, M s F is the spatial attention weight; f 7×7 It is a convolution layer with a convolution kernel size of 7×7; is the channel attention weighted feature map after average pooling, It is the channel attention weighted feature map after maximum pooling.

10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the target vehicle rotation frame detection and recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Small target detection method based on Mama feature fusion

    CN118968019A

  • Detection system, method and equipment for automatically picking pears based on SRSMama

    CN119206315A

  • Remote sensing SAR image rotating vehicle target detection method based on YOLOX

    CN119206343A

  • Method and system for quickly positioning and identifying fine cracks of tunnel lining

    CN119360385A

  • Localization method and system based on deep learning

    WO2020173036A1