Pedestrian trajectory prediction method and system based on multi-target tracking

By combining visible and infrared light modal features in the multi-objective tracking model, the problem of inefficient pedestrian automatic identification and tracking in night monitoring is solved, and efficient safety monitoring is achieved.

CN115760921BActive Publication Date: 2025-05-13KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211501046.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-05-13
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The existing pedestrian trajectory prediction system based on multi-target tracking is inefficient in night monitoring, and cannot efficiently and automatically identify and track pedestrians, which poses safety risks.

Method used

A multi-objective tracking model is adopted, combining visible and infrared light modal features, weighted fusion and position prediction are performed through detectors and trackers to correct the trajectory to improve prediction accuracy.

Benefits of technology

It realizes efficient automatic identification and tracking of pedestrians during night monitoring, improves the efficiency of safety monitoring and reduces manual work burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760921B_ABST
    Figure CN115760921B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian trajectory prediction method and system based on multi-target tracking, which comprises the following steps: obtaining a video to be tracked, and marking multiple target pedestrians in the first frame of the video; inputting the video to be tracked into a trained multi-target tracking model, and the multi-target tracking model outputs target detection frames of multiple target pedestrians in each frame image, and tracking results of multiple target pedestrian trajectories; a detector of the multi-target tracking model performs target detection on non-first frames of the video to be tracked, and obtains visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and the infrared light modal features, and weighted fusion features are obtained; the weighted fusion features are positioned, and a coordinate frame of a target human body is obtained; a tracker predicts the position of the target pedestrian according to the center point position of the coordinate frame of the target human body; and the predicted position is associated with the current position of the target human body to obtain a predicted trajectory of the target human body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a pedestrian trajectory prediction method and system based on multi-target tracking. Background Art

[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous improvement of residents' living standards, the flow of leisure and entertainment passengers in large public places has gradually increased, and the density of pedestrian travel has also increased. A series of crowded accidents will occur, causing the entire passenger flow system to be paralyzed, thus affecting residents' living standards and physical and mental health. Predicting the movement trajectory of pedestrians and taking safe and effective preventive management measures in a timely manner are important measures to prevent such accidents. In order to effectively manage pedestrian traffic in the environment, discover potential risk areas for accidents, and ensure pedestrian traffic safety are the current research focuses of traffic management workers.

[0004] Intelligent video is derived from computer vision technology. It is to establish a relationship between images and image descriptions, so that computers can understand the content of video images through digital image processing and analysis, and achieve the purpose of automatically analyzing and extracting key information from video sources, that is, intelligent video analysis technology (IVS). As a major technology for strengthening the application of video surveillance systems, intelligent video analysis technology has been widely concerned by the industry in recent years. It separates the target of customer concern from the monitoring background through the analysis of video content, and associates the target's moving direction, speed, time and other parameters with certain behavioral characteristics, so as to achieve the purpose of active monitoring and defense.

[0005] In recent years, with the rapid development and interactive integration of big data, cloud computing, artificial intelligence and other fields, concepts such as smart e-commerce, smart transportation, and smart cities have attracted more and more attention. With people's yearning for a smarter, more convenient, and higher quality life, accompanied by great academic value and broad business prospects, many universities, scientific research institutions, and government departments have invested a lot of manpower, material resources, and financial resources in related industries. Artificial intelligence, known as the engine of the industrial revolution in the new era, is quietly infiltrating into all walks of life and changing our lifestyle. Computer vision is an important branch in the field of artificial intelligence, aiming to study how to make computers perceive, analyze, and process the real world intelligently like the human visual system. Various computer vision algorithms with images and videos as information carriers have long been infiltrated into the daily lives of the public, such as face recognition, human-computer interaction, commodity retrieval, intelligent monitoring, visual navigation, etc. Video target tracking technology, as one of the basic and important research directions in the field of computer vision, has always been a hot topic for researchers.

[0006] At the same time, pedestrian trajectory prediction also has important research significance for transportation. Pedestrian trajectory prediction for multi-target tracking can be widely used in unmanned driving technology, Internet travel, video surveillance, transportation hub information acquisition, human-computer interaction and other fields. It can effectively reduce the occurrence of traffic accidents and congestion accidents, and has an important role in promoting the construction and development of the transportation industry.

[0007] At present, the existing pedestrian trajectory prediction system based on multi-target tracking has the following problems:

[0008] 1. In view of the phenomenon that the target pedestrian is frequently blocked by obstacles and multiple people are tracked overlappingly, the existing technology is difficult to obtain an accurate prediction trajectory, and there are difficulties such as difficulty in eliminating the errors in the prediction results.

[0009] 2. Pedestrian trajectories have motion uncertainty and scale irregularity. The complexity of the urban road environment and the density of pedestrians put forward higher requirements for trajectory prediction technology. Traditional tracking and recognition modes are difficult to meet the requirements of accurate, real-time and intelligent pedestrian trajectory prediction.

[0010] 3. Due to the low population density, decreased vigilance of supervisors, poor ambient light and other factors, dusk to early morning has become a high-incidence period for traffic accidents, safety accidents and criminal incidents. Traditional cameras can no longer meet the monitoring needs at night, and there are great safety risks. They cannot automatically identify and track pedestrians efficiently. Summary of the invention

[0011] In order to address the deficiencies of the prior art, the present invention provides a pedestrian trajectory prediction method and system based on multi-target tracking; it solves the technical problem that traditional cameras can no longer meet nighttime monitoring needs, have great safety hazards, and cannot automatically identify and track pedestrians efficiently.

[0012] In a first aspect, the present invention provides a pedestrian trajectory prediction method based on multi-target tracking;

[0013] Pedestrian trajectory prediction method based on multi-target tracking, including:

[0014] Acquire a video to be tracked, and mark multiple target pedestrians in the first frame of the video; the video to be tracked is acquired by a digital camera and an infrared sensor;

[0015] The video to be tracked is input into the trained multi-target tracking model, which outputs the target detection frames of multiple target pedestrians in each frame image and the tracking results of multiple target pedestrian trajectories;

[0016] The trained multi-target tracking model includes a detector and a tracker connected to each other;

[0017] The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body;

[0018] The tracker predicts the position of the target pedestrian according to the center point position of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain the corrected trajectory.

[0019] In a second aspect, the present invention provides a pedestrian trajectory prediction system based on multi-target tracking;

[0020] Pedestrian trajectory prediction system based on multi-target tracking, including:

[0021] The acquisition module is configured to: acquire the video to be tracked and mark multiple target pedestrians in the first frame of the video; the video to be tracked is acquired by a digital camera and an infrared sensor;

[0022] A tracking module is configured to: input the video to be tracked into a trained multi-target tracking model, and the multi-target tracking model outputs target detection frames of multiple target pedestrians in each frame image, and tracking results of multiple target pedestrian trajectories;

[0023] The trained multi-target tracking model includes a detector and a tracker connected to each other;

[0024] The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body;

[0025] The tracker predicts the position of the target pedestrian according to the center point position of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain the corrected trajectory.

[0026] In a third aspect, the present invention further provides an electronic device, comprising:

[0027] a memory for non-transitory storage of computer-readable instructions; and

[0028] a processor for executing the computer readable instructions,

[0029] When the computer-readable instructions are executed by the processor, the method described in the first aspect is executed.

[0030] In a fourth aspect, the present invention further provides a storage medium that non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions of the method described in the first aspect are executed.

[0031] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, wherein the computer program is used to implement the method described in the first aspect when running on one or more processors.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] The present invention discloses a pedestrian trajectory prediction system based on multi-target tracking. Safety monitoring requires continuous detection and tracking of pedestrians in a specific area in order to promptly discover abnormal behaviors of pedestrians or safety hazards in the scene. It is widely used in all corners of daily life. Intelligent monitoring improves efficiency and greatly reduces people's workload by identifying and tracking suspicious pedestrians and automatically analyzing them. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0035] Figure 1 It is a detection and tracking flow chart of the present invention;

[0036] Figure 2 It is the overall flow chart of CBAM of the present invention. DETAILED DESCRIPTION

[0037] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0038] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0039] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0040] In this embodiment, all data is obtained in compliance with laws and regulations and based on the user's consent, and is used legally.

[0041] Embodiment 1

[0042] This embodiment provides a pedestrian trajectory prediction method based on multi-target tracking;

[0043] like Figure 1 As shown, the pedestrian trajectory prediction method based on multi-target tracking includes:

[0044] S101: Acquire a video to be tracked and mark multiple target pedestrians in the first frame of the video; the video to be tracked is acquired by a digital camera and an infrared sensor;

[0045] S102: Input the video to be tracked into the trained multi-target tracking model, and the multi-target tracking model outputs target detection frames of multiple target pedestrians in each frame image, as well as tracking results of multiple target pedestrian trajectories;

[0046] The trained multi-target tracking model includes a detector and a tracker connected to each other;

[0047] The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body;

[0048] The tracker predicts the position of the target pedestrian according to the center point position of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain the corrected trajectory.

[0049] Furthermore, the detector includes: a CAPSNet network for feature extraction, a concat function layer for feature fusion and an attention module; the tracker includes: a Kalman filter algorithm model, a Hungarian algorithm model and a trajectory correction model connected in sequence.

[0050] Furthermore, the training process of the trained multi-target tracking model includes:

[0051] Constructing a training set, wherein the training set is a video of each frame of which pedestrian labeling results are known;

[0052] The training set is input into the multi-target tracking model, and the model is trained. When the total loss function value of the model no longer decreases, the training is stopped to obtain the trained multi-target tracking model; the total loss function of the model refers to the loss function of the detector.

[0053] Furthermore, the detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features, specifically including:

[0054] The CAPsNet network is used to perform target detection on non-first frames of the video to be tracked, and visible light modal features and infrared light modal features are obtained.

[0055] Furthermore, the weighted fusion of the visible light modal features and the infrared light modal features to obtain the weighted fused features specifically includes:

[0056]

[0057] Among them, Γ i represents the eigenvalue after weighted fusion; f cat Represents the concat fusion function; A i Represents the visible light modal characteristics; B i Indicates the infrared light modal characteristics; W v represents the visible light modal features after being processed by the Sigmoid activation function; W i represents the infrared light modal feature after being processed by the Sigmoid activation function; W represents the sum of the visible light modal feature and the infrared light modal feature after being processed by the Sigmoid activation function; represents the weight of visible light modal features; represents the infrared light modal feature weight, f nin Indicates that the visible light modal features and infrared light modal features are compressed through the NIN network layer dimension.

[0058] f ninRepresents the dimensionality compression function under nin, which can be used to compress the dimensions of visible light features and infrared light features. Nin refers to Network in Network, i.e., NIN network, which can enhance the model discrimination. The principle is to use a micro neural network to slide the window to obtain a feature map, and then output the feature map to the next layer, thereby stacking small networks to achieve a deep network.

[0059] Furthermore, after obtaining the weighted fused features and before locating the weighted fused features, the method further includes: optimizing the fused features through a channel attention mechanism module and a spatial attention mechanism module of an attention module CBAM (Convolutional Block Attention Module). The specific optimization process includes:

[0060] (1) Use the channel attention mechanism module to process the fused features;

[0061] (2) Use the spatial attention mechanism module to process the output value of the channel attention mechanism module.

[0062] Furthermore, if Figure 2 As shown, the channel attention mechanism module is used to process the fused features, specifically including:

[0063] First, the weighted fused features are input into the channel attention mechanism module; then the entire feature is subjected to maximum pooling and average pooling to squeeze the features and obtain the global information between channels; then the weight of each channel is obtained through a multi-layer perceptron; finally, a weighted operation is performed to output the channel attention information;

[0064] Among them, the function of maximum pooling is to extract the most obvious feature in the filter and save it. It is used because the previous image contains more noise and information irrelevant to the target processing.

[0065] Average pooling takes the average value of the images in the pooling area. The feature information obtained in this way is more sensitive to background information and can better identify pedestrian trajectory targets and backgrounds.

[0066] The multi-layer perceptron is used for pedestrian targets and backgrounds. By giving different weights to different feature values ​​in the perceptron, the features can be better identified and the attention features can be strengthened.

[0067] Furthermore, the spatial attention mechanism module is used to process the output value of the channel attention mechanism module, specifically including:

[0068] The maximum pooling and average pooling are performed on the entire recognized image information features in sequence, and then the two pooled feature maps are stacked in the channel dimension, and the key information of the features is retained, so that the most important information points of the pedestrian trajectory can be obtained. Then, the convolution kernel is used to fuse the channel information. Finally, the convolution result is normalized to the spatial weight of the feature map through the Sigmoid activation function, and then the input feature map and the weight are multiplied. This can improve the accuracy of target tracking.

[0069] The overall flow chart for CBAM is as follows: Figure 2 As shown in the figure, the input feature map needs to pass through the channel attention mechanism first, the channel weight is multiplied by the input feature map, and then sent to the spatial attention mechanism, the normalized spatial weight is multiplied by the input feature map of the spatial attention mechanism, and the final weighted feature map is obtained.

[0070] It should be understood that in order to establish a more accurate network structure fusion mechanism, CBAM can establish a style cross-attention network, use the cross-cross method to learn and recognize the features of pedestrians in the image, and generate a sparse attention feature map. The image features after multiple cross-fusions will be more accurate.

[0071] Furthermore, the weighted fused features are positioned to obtain a coordinate frame of the target human body, specifically including:

[0072] Combined with pedestrian motion characteristics, the distance weight between any two nodes of the target human body is calculated.

[0073]

[0074] Among them, α and β represent pedestrian feature nodes α and β respectively, (x α ,y α ,z α )、(x β ,y β ,z β ) are the three-dimensional coordinates corresponding to the α and β nodes. The arbitrary two nodes are, for example, the knee node and the palm node of the human body.

[0075] Furthermore, the detector, its loss function Loss, includes:

[0076] Target box position loss function L box , target confidence loss function L obj And the category loss function L cls The summation result of .

[0077] Among them, the target box position loss function L box Represents the positioning error of the prediction box, which is used to represent the weight of the center coordinate error; the target box position loss function L box , specifically:

[0078]

[0079] Among them, the target confidence loss function L obj Represents the intersection-over-union error, which is used as a measurement indicator to describe the overlap between the actual box and the predicted box to ensure the precision and recall of the target detection network; target confidence loss function L obj , specifically:

[0080]

[0081] Among them, the category loss function L cls Represents the classification error, which ensures the accuracy of the network in predicting categories. The error uses the form of cross entropy loss function, and the category loss function L cls Specifically:

[0082]

[0083] Among them, S represents the number of grids into which the pedestrian feature image is divided; B represents the number of boxes predicted for each grid; Respectively represent whether the j-th prediction box of the i-th grid detects the target or not; x, y represent the horizontal and vertical coordinates of the center coordinates of the actual box of the pedestrian feature; w, h represent the length and width of the actual box of the pedestrian feature; Represents the horizontal and vertical coordinate values ​​of the center coordinates of the pedestrian feature prediction box; represents the length and width of the actual box of the pedestrian feature; c represents the actual confidence; represents the prediction confidence; coord , obj , noobj Represents the parameter value of the loss function; classes represents the predicted category parameter; p i (c) represents the actual probability that the detected target belongs to this category; Indicates the predicted probability that the detected target belongs to this category; the loss function Loss limits the length and width of the pedestrian prediction box and reduces the trajectory prediction coordinate offset value.

[0084] It should be understood that the target pedestrian is detected to address the problem of detection result errors caused by poor ambient light.

[0085] Furthermore, the tracker predicts the position of the target pedestrian according to the coordinate frame of the target human body, specifically including:

[0086] The Kalman filter algorithm is used to predict the position of the target pedestrian in the next frame of the image according to the coordinate frame of the target human body, and the motion characteristics are obtained, so that the coordinate frame of the human body moves along the predicted position direction.

[0087] Furthermore, associating the predicted position with the current position of the target body to obtain the predicted trajectory of the target body specifically includes:

[0088] The Hungarian algorithm is used to associate the predicted position with the current position of the target body to obtain the predicted trajectory of the target body.

[0089] For each frame, the tracker will give multiple tracks, each track is composed of several points. After the center point of a new frame is input, after multiple frames are fused, fast tracking is performed, and radar and infrared are combined for positioning and matching. The tracker gives a predicted value, and the predicted value and the actual distance are matched using the Hungarian algorithm. The Hungarian algorithm combines the color histogram features of the surface features. A distance threshold is set inside the tracker. When the threshold is exceeded, the original track will be saved in the system and a new track will be created.

[0090] Furthermore, after obtaining the predicted trajectory of the target human body and before correcting the predicted trajectory of the target human body, the method further includes:

[0091] Determine whether the pedestrian is blocked or overlapped. If so, make corrections; otherwise, do not make corrections.

[0092] It should be understood that the determination of whether a pedestrian is blocked or overlapped is implemented by using the SD-LSTM model. A pooling layer is added to the output end of the original LSTM to obtain the SD-LSTM.

[0093] Furthermore, the predicted trajectory of the target human body is corrected to obtain a corrected trajectory, specifically including:

[0094] The two-frame difference method is used to correct the predicted trajectory of the target human body and obtain the corrected trajectory.

[0095] It should be understood that the trajectory correction model is a two-frame difference method.

[0096] Furthermore, the two-frame difference method is used to correct the predicted trajectory of the target human body to obtain a corrected trajectory, which specifically includes:

[0097] Calculate the average height and width of the tracked target α human coordinate frame and record them as and Compare it with the human coordinate bounding box height h of the target β in the tth frame β and border width w β To compare the sizes, the threshold is set to σ∈{0.99,1.01} according to the dataset.

[0098] Then compare the center position of the β target in the t-1 frame and the t frame and The Euclidean distance between

[0099] like and

[0100]

[0101] This means that the detection targets of the previous and next two frames are the same human coordinate frame, and the pedestrian trajectory does not need to be corrected; otherwise, if the detection targets are different, correction is required, that is, switching the recognition target;

[0102] in, Indicates the center position of the β target in the t-1th frame The horizontal axis of Indicates the center position of the β target in the tth frame The horizontal axis of Indicates the center position of the β target in the t-1th frame The vertical coordinate of Indicates the center position of the β target in the tth frame The vertical coordinate of Indicates the center position and The discrimination threshold of the Euclidean distance between them.

[0103] It should be understood that the error trajectory in the occlusion phenomenon is corrected to prevent the forced interruption of the trajectory or the change of the recognition target. In summary, safety monitoring requires continuous detection and tracking of pedestrians in a specific area in order to promptly detect abnormal behavior of pedestrians or safety hazards in the scene. It is widely used in every corner of daily life. Intelligent monitoring can greatly reduce people's workload by identifying and tracking suspicious pedestrians and automatically analyzing them, improving efficiency; using video target tracking technology to monitor pedestrian trajectories in real time, further facilitating scene analysis, order maintenance and intelligent scheduling, and saving manpower and material resources; in addition, target tracking technology also plays an important role in video editing, three-dimensional reconstruction, robotics, mechanical automatic control and other fields.

[0104] Embodiment 2

[0105] This embodiment provides a pedestrian trajectory prediction system based on multi-target tracking;

[0106] Pedestrian trajectory prediction system based on multi-target tracking, including:

[0107] The acquisition module is configured to: acquire the video to be tracked and mark multiple target pedestrians in the first frame of the video; the video to be tracked is acquired by a digital camera and an infrared sensor;

[0108] A tracking module is configured to: input the video to be tracked into a trained multi-target tracking model, and the multi-target tracking model outputs target detection frames of multiple target pedestrians in each frame image, and tracking results of multiple target pedestrian trajectories;

[0109] The trained multi-target tracking model includes a detector and a tracker connected to each other;

[0110] The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body;

[0111] The tracker predicts the position of the target pedestrian according to the center point position of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain the corrected trajectory.

[0112] It should be noted that the acquisition module and the tracking module described above correspond to steps S101 to S102 in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules described above as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0113] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0114] The proposed system can be implemented in other ways. For example, the system embodiment described above is only illustrative, and the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0115] Embodiment 3

[0116] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the method described in the above embodiment one.

[0117] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0118] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0119] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software.

[0120] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0121] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0122] Embodiment 4

[0123] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is completed.

[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A pedestrian trajectory prediction method based on multi-target tracking, characterized by: include: Get the video to be tracked and mark multiple target pedestrians in the first frame of the video; The video to be tracked is obtained by a digital camera and an infrared sensor; The video to be tracked is input into the trained multi-target tracking model, which outputs the target detection frames of multiple target pedestrians in each frame image and the tracking results of multiple target pedestrian trajectories; The trained multi-target tracking model includes a detector and a tracker connected to each other; The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body; The tracker predicts the position of the target pedestrian according to the center point of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain a corrected trajectory; The step of correcting the predicted trajectory of the target human body to obtain a corrected trajectory specifically includes: The predicted trajectory of the target human body is corrected by using the two-frame difference method to obtain the corrected trajectory; The two-frame difference method is used to correct the predicted trajectory of the target human body to obtain the corrected trajectory, which specifically includes: Counting Tracked Targets The average height and width of the human coordinate frame are recorded as and , and compare it with Frame target The height of the human body coordinate frame and border width Compare the sizes and set the threshold according to the data set. ; Compare the Frame and In frame The center position of the target and The Euclidean distance between ; like , ,and This means that the detection targets of the previous and next frames are the same human coordinate frame, and the pedestrian trajectory does not need to be corrected; otherwise, if the detection targets are different, correction is required, that is, switching the recognition target; in, Indicates In the frame, The center position of the target The horizontal axis of Indicates In the frame, The center position of the target The horizontal axis of Indicates In the frame, The center position of the target The vertical coordinate of Indicates In the frame, The center position of the target The vertical coordinate of Indicates the center position and The discrimination threshold of the Euclidean distance between them.

2. The pedestrian trajectory prediction method based on multi-target tracking as claimed in claim 1, characterized in that: The detector includes: a CAPsNet network for feature extraction, a concat function layer for feature fusion and an attention module; the tracker includes: a Kalman filter algorithm model, a Hungarian algorithm model and a trajectory correction model connected in sequence.

3. The pedestrian trajectory prediction method based on multi-target tracking as claimed in claim 1, characterized in that: The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features, specifically including: The CAPsNet network is used to detect targets in non-first frames of the video to be tracked, and visible light modal features and infrared light modal features are obtained; The weighted fusion of the visible light modal features and the infrared light modal features to obtain the weighted fused features specifically includes: in, Represents the eigenvalue after weighted fusion; Represents the concat fusion function; Represents the visible light modal characteristics; Indicates the infrared light modal characteristics; Represents the visible light modal features after being processed by the Sigmoid activation function; Represents the infrared light modal features after being processed by the Sigmoid activation function; Represents the sum of visible light modal features and infrared light modal features after being processed by the Sigmoid activation function; represents the weight of visible light modal features; represents the infrared light modal feature weight, Indicates that the visible light modal features and infrared light modal features are compressed through the NIN network layer dimension.

4. The pedestrian trajectory prediction method based on multi-target tracking as claimed in claim 1, characterized in that: After obtaining the weighted fused features and before locating the weighted fused features, the method further includes: optimizing the fused features through the channel attention mechanism module and the spatial attention mechanism module of the attention module CBAM. The specific optimization process includes: Use the channel attention mechanism module to process the fused features; Use the spatial attention mechanism module to process the output value of the channel attention mechanism module; The processing of the fused features by using the channel attention mechanism module specifically includes: First, the weighted fused features are input into the channel attention mechanism module; then the entire feature is subjected to maximum pooling and average pooling to squeeze the features and obtain the global information between channels; then the weight of each channel is obtained through a multi-layer perceptron; finally, a weighted operation is performed to output the channel attention information; The processing of the output value of the channel attention mechanism module by using the spatial attention mechanism module specifically includes: The entire recognized image information features are subjected to maximum pooling and average pooling in turn, and then the two pooled feature maps are stacked in the channel dimension and the key information of the features is retained to obtain the most important information points of the pedestrian trajectory. Then, the convolution kernel is used to fuse the channel information. Finally, the convolution result is normalized through the Sigmoid activation function to normalize the spatial weight of the feature map, and then the input feature map and the weight are multiplied.

5. The pedestrian trajectory prediction method based on multi-target tracking as claimed in claim 1, characterized in that: The detector, its loss function ,include: Target box position loss function , target confidence loss function And the category loss function The summation result of Among them, the target box position loss function Represents the positioning error of the prediction box, which is used to represent the weight of the center coordinate error; target box position loss function , specifically: Among them, the target confidence loss function Represents the intersection-over-union error, which is used as a measurement indicator to describe the overlap between the actual box and the predicted box to ensure the precision and recall of the target detection network; target confidence loss function , specifically: Among them, the category loss function Represents the classification error, ensuring the accuracy of the network in predicting categories. The error uses the form of cross entropy loss function, and the category loss function Specifically: in, Indicates the number of grids into which the pedestrian feature image is divided; Indicates the number of boxes predicted for each grid; Respectively indicate whether the j-th prediction box of the i-th grid detects the target or not; Indicates the horizontal and vertical coordinate values ​​of the center coordinates of the actual frame of the pedestrian feature; Indicates the length and width of the actual box of the pedestrian feature; Represents the horizontal and vertical coordinate values ​​of the center coordinates of the pedestrian feature prediction box; Indicates the length and width of the actual box of the pedestrian feature; Indicates the actual confidence level; Indicates the prediction confidence; Represents the parameter value of the loss function; represents the predicted category parameter; Indicates the actual probability that the detected target belongs to this category; Indicates the predicted probability that the detected target belongs to this category; loss function , limit the length and width of the pedestrian prediction box and reduce the trajectory prediction coordinate offset value.

6. The pedestrian trajectory prediction method based on multi-target tracking as claimed in claim 1, characterized in that: The tracker predicts the position of the target pedestrian according to the center point position of the coordinate frame of the target human body, specifically including: The Kalman filter algorithm is used to predict the position of the target pedestrian in the next frame of the image according to the coordinate frame of the target human body, and the motion characteristics are obtained, so that the coordinate frame of the human body moves along with the predicted position direction; The step of associating the predicted position with the current position of the target body to obtain the predicted trajectory of the target body specifically includes: The Hungarian algorithm is used to associate the predicted position with the current position of the target body to obtain the predicted trajectory of the target body; After obtaining the predicted trajectory of the target human body and before correcting the predicted trajectory of the target human body, the method further includes: Determine whether the pedestrian is blocked or overlapped. If so, make corrections; otherwise, do not make corrections.

7. A pedestrian trajectory prediction system based on multi-target tracking, characterized by: include: An acquisition module is configured to: acquire a video to be tracked and mark multiple target pedestrians in the first frame of the video; The video to be tracked is obtained by a digital camera and an infrared sensor; A tracking module is configured to: input the video to be tracked into a trained multi-target tracking model, and the multi-target tracking model outputs target detection frames of multiple target pedestrians in each frame image, and tracking results of multiple target pedestrian trajectories; The trained multi-target tracking model includes a detector and a tracker connected to each other; The detector performs target detection on non-first frames of the video to be tracked to obtain visible light modal features and infrared light modal features; weighted fusion is performed on the visible light modal features and infrared light modal features to obtain weighted fused features; the weighted fused features are located to obtain the coordinate frame of the target human body; The tracker predicts the position of the target pedestrian according to the center point of the target person's coordinate frame; associates the predicted position with the current position of the target person to obtain the predicted trajectory of the target person; and corrects the predicted trajectory of the target person to obtain a corrected trajectory; The step of correcting the predicted trajectory of the target human body to obtain a corrected trajectory specifically includes: The predicted trajectory of the target human body is corrected by using the two-frame difference method to obtain the corrected trajectory; The two-frame difference method is used to correct the predicted trajectory of the target human body to obtain the corrected trajectory, which specifically includes: Counting Tracked Targets The average height and width of the human coordinate frame are recorded as and , and compare it with Frame target The height of the human body coordinate frame and border width Compare the sizes and set the threshold according to the data set. ; Compare the Frame and In frame The center position of the target and The Euclidean distance between ; like , ,and This means that the detection targets of the previous and next frames are the same human coordinate frame, and the pedestrian trajectory does not need to be corrected; otherwise, if the detection targets are different, correction is required, that is, switching the recognition target; in, Indicates In the frame, The center position of the target The horizontal axis of Indicates In the frame, The center position of the target The horizontal axis of Indicates In the frame, The center position of the target The vertical coordinate of Indicates In the frame, The center position of the target The vertical coordinate of Indicates the center position and The discrimination threshold of the Euclidean distance between them.

8. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; as well as a processor for executing the computer readable instructions, Wherein, when the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 6 is executed.

9. A storage medium, characterized in that: The computer-readable instructions are stored, wherein when the computer-readable instructions are executed by a computer, the instructions of the method according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Multi-target tracking method for yellow-feather chickens in flat-feeding chicken house

    CN113470076A

  • Intersection anti-overflow signal control system and method

    CN113643533A