Method and system for identifying and tracking moving target on road, electronic equipment and storage medium

By adopting a mobile target tracking model with twin networks and attention mechanisms in road monitoring, combined with Taylor series approximate optical flow formula and attention weighting mechanism, the problem of mobile target recognition and tracking under low-light conditions is solved, and efficient and robust target recognition and tracking effects are achieved.

CN119992481AInactive Publication Date: 2025-05-13SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510346797.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Under low-light conditions, in road monitoring and traffic management, traditional computer vision algorithms are difficult to accurately identify and track mobile targets, and are affected by lighting changes, noise, complex backgrounds and real-time requirements.

Method used

A twin network and attention mechanism are used to build an improved mobile target tracking model, and the optical flow formula is initially screened through the Taylor series approximation calculation method. Multi-scale feature maps are extracted in combination with channel and spatial attention weighting mechanisms, and a RPN network is used to obtain the motion target candidate area map.

Benefits of technology

It improves the recognition accuracy and tracking stability of moving targets in low-light environments, enhances the robustness of the model to different lighting conditions, optimizes the spatial and temporal complexity of the algorithm, and meets the needs of real-time traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992481A_ABST
    Figure CN119992481A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for recognizing and tracking a moving target on a road, electronic equipment and a storage medium, and belongs to the technical field of machine vision, and the method comprises the steps: S1, collecting an image of the moving target on the road, carrying out the preprocessing of the image, and screening a moving target image in a weak light environment, and obtaining an initial image; s2, constructing an improved moving target tracking model by using a twin network and an attention mechanism; and S3, inputting the initial image into the improved moving target tracking model to obtain an identification and tracking result of the moving target on the road.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision technology, and in particular relates to a method, system, electronic device and storage medium for identifying and tracking a moving target on a road. Background Art

[0002] In road monitoring and traffic management, accurately identifying and continuously tracking moving targets is a critical but challenging task. Especially at night or in low light conditions, due to factors such as poor imaging quality, low contrast, and high noise, traditional computer vision algorithms are difficult to achieve ideal results. First, weak ambient lighting causes blurred object contours in the image, and there is a lack of obvious gradient differences between the background and the target, making the detection and segmentation of moving targets more difficult. Secondly, frequent changes in lighting can cause drastic fluctuations in the brightness and contrast of the image sequence, which brings additional interference to the extraction and expression of the target appearance features. Furthermore, low-light images generally have large noise, which can easily introduce a large number of false targets, resulting in subsequent tracking failures. In addition, there are some other special factors in road scenes. Vehicle headlights will produce strong spots and glare, resulting in overexposure in local areas, drowning out effective target information. Complex dynamic backgrounds such as road reflections, tree shades, and building projections will also cause serious interference to motion estimation and target positioning. At the same time, in practical applications, the real-time requirements of the algorithm also need to be considered. Low-light image enhancement, multi-target detection, and long-term tracking are all computationally intensive tasks. How to optimize the algorithm structure under limited hardware conditions and strike a balance between recognition accuracy and processing speed is an urgent problem to be solved. Therefore, it is urgent to study a robust and efficient method for road moving target recognition and tracking under low-light conditions. This method should be able to adapt to changes in illumination, overcome the impact of image degradation, and effectively distinguish complex background interference and extract stable and reliable target features. On this basis, the algorithm's temporal and spatial complexity is further optimized to meet the performance requirements of the actual system, and ultimately achieve all-weather and all-day road monitoring and traffic management. Summary of the invention

[0003] To solve the above problems, the present invention provides a method, system, electronic device and storage medium for identifying and tracking a moving target on a road, including:

[0004] Step S1, collecting images of moving targets on the road, pre-processing the images of moving targets on the road, and filtering the images of moving targets in a weak light environment to obtain an initial image;

[0005] Step S2: construct an improved mobile target tracking model using the twin network and attention mechanism;

[0006] Step S3: input the initial image into the improved moving target tracking model to obtain the moving target recognition and tracking results on the road.

[0007] Optionally, in step S1, the preprocessing process is:

[0008] Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula;

[0009] Using a channel and spatial attention weighting mechanism to extract global information features and local detail features from the road moving target image after preliminary screening, and fusing the global information features and the local detail features to obtain a multi-scale feature map;

[0010] The multi-scale feature map is cropped using the RPN network to obtain the candidate region map of the moving target and complete the image preprocessing.

[0011] Optionally, based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes:

[0012] In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is:

[0013] I(x,y,t)=I(δx+δx,y+δy+δt);

[0014] Using Taylor series approximation we get:

[0015]

[0016] After removing the common terms, the optical flow formula is obtained:

[0017]

[0018] Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

[0019] Optionally, in step S2, the contents of constructing an improved mobile target tracking model using the twin network and the attention mechanism specifically include:

[0020] The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

[0021] The present invention also discloses a system for identifying and tracking moving targets on a road, the system comprising:

[0022] An image preprocessing module is used to collect images of moving targets on the road, preprocess the images of moving targets on the road, and then filter the images of moving targets in a weak light environment to obtain an initial image;

[0023] Tracking target building module, which is used to build an improved mobile target tracking model using Siamese networks and attention mechanisms;

[0024] The recognition and tracking module is used to input the initial image into the improved mobile target tracking model to obtain the recognition and tracking results of the mobile target on the road.

[0025] Optionally, in the image preprocessing module, the preprocessing process is:

[0026] Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula;

[0027] Using a channel and spatial attention weighting mechanism to extract global information features and local detail features from the road moving target image after preliminary screening, and fusing the global information features and the local detail features to obtain a multi-scale feature map;

[0028] The multi-scale feature map is cropped using the RPN network to obtain the candidate region map of the moving target and complete the image preprocessing.

[0029] Optionally, based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes:

[0030] In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is:

[0031] I(x,y,t)=I(δx+δx,y+δy+δt);

[0032] Using Taylor series approximation we get:

[0033]

[0034] After removing the common terms, the optical flow formula is obtained:

[0035]

[0036] Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

[0037] Optionally, the tracking target construction module specifically includes:

[0038] The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

[0039] The present invention also discloses an electronic device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 4 when executing the program.

[0040] The present invention also discloses a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 4 is implemented.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] By using Taylor series to approximate the optical flow formula, the motion of the moving target in continuous frames can be captured more accurately, thereby improving the tracking accuracy. The image of the moving target is screened in a weak light environment, which enhances the robustness of the model under different lighting conditions. The global information features and local detail features are extracted using the channel and spatial attention weighted mechanisms, and these features are fused to improve the model's tracking ability for targets of different scales. The RPN network crops the multi-scale feature map to obtain the candidate region map of the moving target, which reduces unnecessary calculations and improves the running efficiency of the model. The combination of the twin network and the attention mechanism enables the model to adapt to changes in the appearance of the target and improves the tracking stability. The introduction of the attention mechanism helps to distinguish similar objects and reduce target confusion, especially when the target is similar to the background or there are multiple similar targets. By updating the template online and adaptively, the model can better cope with changes in the shape of the target and improve the generalization ability of the model in different scenarios. The introduction of the channel attention mechanism and the spatial attention mechanism enables the model to pay more attention to important features and optimize the allocation of computing resources. The attention mechanism helps to construct information features and improve the discrimination of the network, especially the model drift problem caused by the presence of similar backgrounds. Due to the lightweight design of the model, the complexity of model operations is reduced, thereby improving the tracking speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0044] Figure 1 A method step diagram of a method for identifying and tracking a moving target on a road designed for an embodiment of the present invention;

[0045] Figure 2 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Description of the drawings:

[0047] 1010, processor; 1020, memory; 1030, input / output interface; 1040, communication interface; 1050, bus. DETAILED DESCRIPTION

[0048] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. It should be clear that the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention. The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0049] Embodiment 1

[0050] A method for identifying and tracking a moving target on a road, such as Figure 1 As shown, including:

[0051] Step S1, collecting images of moving targets on the road, pre-processing the images of moving targets on the road, and filtering the images of moving targets in a weak light environment to obtain an initial image.

[0052] The preprocessing process is as follows: based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula; the channel and spatial attention weighted mechanism is used to extract the global information features and local detail features of the road moving target image after the preliminary screening, and the global information features and the local detail features are integrated to obtain a multi-scale feature map; the multi-scale feature map is cropped using the RPN network to obtain a moving target candidate area map, and the image preprocessing is completed.

[0053] Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result. The content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes:

[0054] In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is:

[0055] I(x,y,t)=I(δx+δx,y+δy+δt);

[0056] Using Taylor series approximation we get:

[0057]

[0058] After removing the common terms, the optical flow formula is obtained:

[0059]

[0060] Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

[0061] For a single-frame image input into the network, it first undergoes feature extraction through a deep convolutional neural network to generate six feature maps of different scales. The large-scale feature map has rich local detail information and can better learn the information of small targets during network training. The higher-level small-scale feature map is more friendly in terms of global semantic information, can better learn the association information between targets, and is more effective in distinguishing the background and foreground. These six feature maps of different scales first pass through the channel attention module to learn the channel weighted information at different scales. For feature maps of different scales, the network structure of the channel attention module is the same, but the weights learned by each are different, which gives them different focuses.

[0062] The input dimension of the feature map is F∈(C,H,W), where C is the channel dimension, H and W are the height and width of the feature map respectively. After entering the channel attention module, in order to better learn the information on the channel, the spatial dimension of the feature map is first compressed. The compression of the spatial dimension uses two parallel methods, namely global average pooling and global maximum pooling. After pooling, a channel dimension feature map with a dimension of (C,1,1) can be obtained.

[0063]

[0064] Among them, GAP S ,GMP Sis the average pooling and maximum pooling of the spatial dimension, and different Fs are feature maps after different pooling. During the training process, the maximum pooling records the position of the maximum value in the feature map, and returns the gradient value to the corresponding position during back propagation, so that the texture information of the feature can be better preserved; the average pooling calculates the average of each block and preserves the overall information of the feature map. The two work together to supplement the channel attention mechanism in SENet.

[0065] After parallel pooling operations, the multi-layer perceptron network is input to obtain the weighted channel attention feature results:

[0066]

[0067] Among them, σ is the sigmoid function, and is the feature after pooling, and X is the original input.

[0068] As calculated above, the weighted spatial attention feature is calculated based on the weighted channel attention feature:

[0069]

[0070] Among them, f 7×7 It is a convolution layer with a convolution kernel size of 7×7.

[0071] The RPN network further improves the feature robustness by fusing the spatial information in the feature map through 3x3 convolution processing, and divides the region (anchor) corresponding to the center point of the feature map string window, and obtains the candidate region map of the moving target by calculating the scaling processing of the candidate region.

[0072] Step S2: Use the twin network and attention mechanism to build an improved mobile target tracking model, specifically including:

[0073] The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

[0074] Introducing correlation filtering into the twin network AlexNet improves the performance of target tracking. The correlation filter template is set to n, which is the feature map of the d channel, and y is the expected response map, that is, the response map of the tracking process is obtained:

[0075]

[0076] The filters used are:

[0077]

[0078] in, The search area has the same features as the target area, and then the target area and the search area are convolved to obtain the corresponding map.

[0079] The channel attention mechanism consists of a pooling layer and two fully connected layers, a ReLu activation layer and a Sigmoid activation layer. Let the input feature be F, and the feature vector obtained after the average pooling layer be r. The obtained feature vector r is input into the next fully connected layer, where n represents the number of channels. After a fully connected layer and a nonlinear activation function, the output vector is obtained, and the output vector is multiplied by the input feature F to obtain the weighted channel attention feature map.

[0080] The input feature of spatial attention is B. After three 1×1 convolutional layers, feature maps B1, B2, and B3 are obtained. The feature maps B1 and B3 are converted to N and multiplied with the transposed matrix of B2 to obtain the spatial attention feature map.

[0081] Step S3: input the initial image into the improved moving target tracking model to obtain the moving target recognition and tracking results on the road.

[0082] Embodiment 2

[0083] A system for identifying and tracking moving targets on a road, the system comprising:

[0084] The image preprocessing module is used to collect images of moving targets on the road, preprocess the images of moving targets on the road, and then filter the images of moving targets in a weak light environment to obtain an initial image.

[0085] The preprocessing process is as follows: based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula; the channel and spatial attention weighted mechanism is used to extract the global information features and local detail features of the road moving target image after the preliminary screening, and the global information features and the local detail features are integrated to obtain a multi-scale feature map; the multi-scale feature map is cropped using the RPN network to obtain a moving target candidate area map, and the image preprocessing is completed.

[0086] Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result. The content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes:

[0087] In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is:

[0088] I(x,y,t)=I(δx+δx,y+δy+δt);

[0089] Using Taylor series approximation we get:

[0090]

[0091] After removing the common terms, the optical flow formula is obtained:

[0092]

[0093] Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

[0094] For a single-frame image input into the network, it first undergoes feature extraction through a deep convolutional neural network to generate six feature maps of different scales. The large-scale feature map has rich local detail information and can better learn the information of small targets during network training. The higher-level small-scale feature map is more friendly in terms of global semantic information, can better learn the association information between targets, and is more effective in distinguishing the background and foreground. These six feature maps of different scales first pass through the channel attention module to learn the channel weighted information at different scales. For feature maps of different scales, the network structure of the channel attention module is the same, but the weights learned by each are different, which gives them different focuses.

[0095] The input dimension of the feature map is F∈(C,H,W), where C is the channel dimension, H and W are the height and width of the feature map respectively. After entering the channel attention module, in order to better learn the information on the channel, the spatial dimension of the feature map is first compressed. The compression of the spatial dimension uses two parallel methods, namely global average pooling and global maximum pooling. After pooling, a channel dimension feature map with a dimension of (C,1,1) can be obtained.

[0096]

[0097] Among them, GAP S ,GMP S is the average pooling and maximum pooling of the spatial dimension, and different Fs are feature maps after different pooling. During the training process, the maximum pooling records the position of the maximum value in the feature map, and returns the gradient value to the corresponding position during back propagation, so that the texture information of the feature can be better preserved; the average pooling calculates the average of each block and preserves the overall information of the feature map. The two work together to supplement the channel attention mechanism in SENet.

[0098] After parallel pooling operations, the multi-layer perceptron network is input to obtain the weighted channel attention feature results:

[0099]

[0100] Among them, σ is the sigmoid function, and is the feature after pooling, and X is the original input.

[0101] As calculated above, the weighted spatial attention feature is calculated based on the weighted channel attention feature:

[0102]

[0103] Among them, f 7×7 It is a convolution layer with a convolution kernel size of 7×7.

[0104] The RPN network further improves the feature robustness by fusing the spatial information in the feature map through 3x3 convolution processing, and divides the region (anchor) corresponding to the center point of the feature map string window, and obtains the candidate region map of the moving target by calculating the scaling processing of the candidate region.

[0105] The tracking target building module is used to build an improved mobile target tracking model using the twin network and attention mechanism, including:

[0106] The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

[0107] Introducing correlation filtering into the twin network AlexNet improves the performance of target tracking. The correlation filter template is set to n, which is the feature map of the d channel, and y is the expected response map, that is, the response map of the tracking process is obtained:

[0108]

[0109] The filters used are:

[0110]

[0111] in, The search area has the same features as the target area, and then the target area and the search area are convolved to obtain the corresponding map.

[0112] The channel attention mechanism consists of a pooling layer and two fully connected layers, a ReLu activation layer and a Sigmoid activation layer. Let the input feature be F, and the feature vector obtained after the average pooling layer be r. The obtained feature vector r is input into the next fully connected layer, where n represents the number of channels. After a fully connected layer and a nonlinear activation function, the output vector is obtained, and the output vector is multiplied by the input feature F to obtain the weighted channel attention feature map.

[0113] The input feature of spatial attention is B. After three 1×1 convolutional layers, feature maps B1, B2, and B3 are obtained. The feature maps B1 and B3 are converted to N and multiplied with the transposed matrix of B2 to obtain the spatial attention feature map.

[0114] The recognition and tracking module is used to input the initial image into the improved mobile target tracking model to obtain the recognition and tracking results of the mobile target on the road.

[0115] Embodiment 3

[0116] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a method for identifying and tracking a mobile target on a road as described in any of the above embodiments is implemented.

[0117] Figure 2 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0118] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0119] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0120] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0121] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB (Universal Serial Bus), network cable, etc.), or through a wireless mode (such as mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).

[0122] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0123] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0124] The system of the above embodiment is used to implement a corresponding method for identifying and tracking a moving target on the road in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0125] Embodiment 4

[0126] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute a method for identifying and tracking mobile targets on the road as described in any of the above embodiments.

[0127] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0128] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute a method for identifying and tracking a moving target on a road as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0129] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0130] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0131] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0132] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.

[0133] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for identifying and tracking a moving target on a road, characterized in that: The method specifically comprises: Step S1, collecting images of moving targets on the road, pre-processing the images of moving targets on the road, and filtering the images of moving targets in a weak light environment to obtain an initial image; Step S2: construct an improved mobile target tracking model using the twin network and attention mechanism; Step S3: input the initial image into the improved moving target tracking model to obtain the moving target recognition and tracking results on the road.

2. The method for identifying and tracking a moving target on a road according to claim 1, characterized in that: In step S1, the pre-processing process is: Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula; Using a channel and spatial attention weighting mechanism to extract global information features and local detail features from the road moving target image after preliminary screening, and fusing the global information features and the local detail features to obtain a multi-scale feature map; The multi-scale feature map is cropped using the RPN network to obtain the candidate region map of the moving target and complete the image preprocessing.

3. The method for identifying and tracking a moving target on a road according to claim 2, characterized in that: Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result. The content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes: In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is: I(x,y,t)=I(δx+δx,y+δy+δt); Using Taylor series approximation we get: After removing the common terms, the optical flow formula is obtained: Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

4. The method for identifying and tracking a moving target on a road according to claim 1, characterized in that: In step S2, the contents of constructing an improved mobile target tracking model using the twin network and the attention mechanism specifically include: The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

5. A system for identifying and tracking mobile targets on a road, the system applying the method for identifying and tracking mobile targets according to any one of claims 1 to 4, characterized in that: The system includes: An image preprocessing module is used to collect images of moving targets on the road, preprocess the images of moving targets on the road, and then filter the images of moving targets in a weak light environment to obtain an initial image; Tracking target building module, which is used to build an improved mobile target tracking model using Siamese networks and attention mechanisms; The recognition and tracking module is used to input the initial image into the improved mobile target tracking model to obtain the recognition and tracking results of the mobile target on the road.

6. The system for identifying and tracking mobile targets on the road according to claim 5, characterized in that: In the image preprocessing module, the preprocessing process is: Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result, and the road moving target image is preliminarily screened based on the calculation result of the optical flow formula; Using a channel and spatial attention weighting mechanism to extract global information features and local detail features from the road moving target image after preliminary screening, and fusing the global information features and the local detail features to obtain a multi-scale feature map; The multi-scale feature map is cropped using the RPN network to obtain the candidate region map of the moving target and complete the image preprocessing.

7. The system for identifying and tracking moving targets on the road according to claim 6, characterized in that: Based on several consecutive frames of the road moving target image, the optical flow formula is calculated using the Taylor series approximation result. The content of preliminary screening of the road moving target image based on the calculation result of the optical flow formula specifically includes: In the continuous frames, let I be the image pixel intensity, which is a function of space (x, y) and time t. The distance moved by the moving target in δt is (δx, δy). Then the pixel intensity I is: I(x,y,t)=I(δx+δx,y+δy+δt); Using Taylor series approximation we get: After removing the common terms, the optical flow formula is obtained: Among them, u and v are the velocities of the moving target in the horizontal and vertical directions.

8. The system for identifying and tracking moving targets on the road according to claim 5, characterized in that: The tracking target building module specifically includes: The twin network AlexNet is used as the backbone network of the model, and the channel attention mechanism and spatial attention mechanism are introduced. The channel attention mechanism consists of a pooling layer and two fully connected layers connecting the ReLu activation layer and the Sigmoid activation layer; the spatial attention mechanism consists of three 1×1 convolutional layers and a Softmax activation function.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 4 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Target tracking method and device fusing optical flow information and Siamese framework

    CN110619655A

  • Target tracking algorithm based on optical flow and dynamic cascade RPN

    CN114359336A

  • Target tracking method based on gated attention mechanism and space-time memory network

    CN119131085A

  • Target tracking method and device fusing optical flow information and siamese framework

    WO2021035807A1