Trajectory determination method and device

By obtaining the pedestrian frame confidence and adjacent occlusion rate of the video frame image, using the Kalman filter and pedestrian feature matching to screen the target trajectory frame, the confusion and error problems in pedestrian trajectory recognition are solved, and the recognition accuracy and speed are improved.

CN115661201BActive Publication Date: 2025-09-12SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211314218.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-09-12
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

Existing pedestrian trajectory recognition technology is prone to confusion and errors when multiple user trajectories intersect in the video image, and has slow recognition speed and low efficiency.

Method used

By obtaining the pedestrian frame of the video frame image and its confidence and adjacent occlusion rate, the Kalman filter is used to predict the trajectory frame. Combined with pedestrian feature matching, the target trajectory frame is screened out to distinguish each trajectory and avoid confusion and error.

Benefits of technology

It improves the accuracy and speed of pedestrian trajectory recognition, solves the recognition confusion problem under the intersection of multiple user trajectories, and improves processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661201B_ABST
    Figure CN115661201B_ABST
Patent Text Reader

Abstract

The present disclosure provides a trajectory determination method and device. This method uses the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in each frame to filter out the target trajectory frame. Therefore, when multiple user trajectories intersect in a video, the target trajectory frame can be used to distinguish the individual trajectories, avoiding the problem of trajectory confusion and errors in the pedestrian trajectory recognition results, thereby improving the accuracy of the trajectory recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a trajectory determination method and device. Background Art

[0002] Pedestrian trajectory recognition technology has high application value in public security, criminal investigation, and image retrieval. However, current pedestrian trajectory recognition technology is prone to confusion and errors when multiple users' trajectories intersect in a video. Furthermore, pedestrian trajectory recognition is slow and inefficient. Therefore, a new pedestrian trajectory recognition method is urgently needed. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure provide a trajectory determination method, apparatus, computer device, and computer-readable storage medium to address the problems of slow trajectory determination speed, low efficiency, and low accuracy of trajectory determination results in the prior art.

[0004] A first aspect of the present disclosure provides a trajectory determination method, the method comprising:

[0005] Obtaining a video to be processed, and determining a pedestrian frame corresponding to each frame image of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame;

[0006] Determine, according to the confidence level and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed, the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame;

[0007] For the i-th frame of the video to be processed, determine the target trajectory frame corresponding to the i-th frame and the next predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame, the confidence score and adjacent occlusion rate corresponding to the pedestrian frame, and the next predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame; i is a positive integer greater than 1;

[0008] The pedestrian trajectory in the video to be processed is determined according to the target trajectory frame of each frame image of the video to be processed.

[0009] According to a second aspect of the present disclosure, a trajectory determination device is provided, the device comprising:

[0010] A first determining unit is configured to obtain a video to be processed, and determine a pedestrian frame corresponding to each frame of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame;

[0011] A second determining unit is configured to determine a target trajectory frame of the first frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame according to a confidence level and an adjacent occlusion rate corresponding to a pedestrian frame of the first frame image of the video to be processed;

[0012] a third determining unit, configured to determine, for an i-th frame image of the video to be processed, a target trajectory frame corresponding to the i-th frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame based on a pedestrian frame corresponding to the i-th frame image, a confidence score and an adjacent occlusion rate corresponding to the pedestrian frame, and a subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; i is a positive integer greater than 1;

[0013] The fourth determining unit is configured to determine a pedestrian trajectory in the video to be processed according to a target trajectory frame of each frame of the video to be processed.

[0014] According to a third aspect of an embodiment of the present disclosure, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0015] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0016] Compared with the prior art, the embodiments of the present disclosure have the following beneficial effects: the embodiments of the present disclosure have a trajectory determination method. After obtaining a video to be processed, the method first determines the pedestrian frame corresponding to each frame image of the video to be processed and the confidence and adjacent occlusion rate corresponding to the pedestrian frame; then, based on the confidence and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed, the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame can be determined; then, for the i-th frame image of the video to be processed, based on the pedestrian frame corresponding to the i-th frame image, the confidence and adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image, the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame are determined; finally, the pedestrian trajectory in the video to be processed can be determined based on the target trajectory frame of each frame image of the video to be processed. In the present application, since the adjacent occlusion rate corresponding to the pedestrian frame can reflect the probability of occurrence of overlapping and intersecting trajectories, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in each frame image can be used to screen out the target trajectory frame. Therefore, when there are multiple users' trajectories intersecting in the video screen, the target trajectory frame can be used to distinguish each trajectory, thereby avoiding the problem of trajectory confusion and errors in the recognition results of pedestrian trajectories, thereby improving the accuracy of the trajectory recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 is a schematic diagram of an application scenario of an embodiment of the present disclosure;

[0019] Figure 2 is a flow chart of a trajectory determination method provided by an embodiment of the present disclosure;

[0020] Figure 3 is a block diagram of a trajectory determination device provided by an embodiment of the present disclosure;

[0021] Figure 4 Schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary detail.

[0023] A trajectory determination method and apparatus according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0024] In the prior art, when the trajectories of multiple users intersect in a video image, pedestrian trajectory recognition results are prone to trajectory confusion and errors, and the speed of pedestrian trajectory recognition is slow and inefficient.

[0025] In order to solve the above problems, the present invention provides a trajectory determination method. After obtaining a video to be processed, the method first determines the pedestrian box corresponding to each frame image of the video to be processed and the confidence and adjacent occlusion rate corresponding to the pedestrian box; then, based on the confidence and adjacent occlusion rate corresponding to the pedestrian box of the first frame image of the video to be processed, the target trajectory box of the first frame image and the subsequent predicted trajectory box corresponding to the target trajectory box can be determined; then, for the i-th frame image of the video to be processed, based on the pedestrian box corresponding to the i-th frame image, the confidence and adjacent occlusion rate corresponding to the pedestrian box, and the subsequent predicted trajectory box corresponding to the target trajectory box of the i-1-th frame image, the target trajectory box corresponding to the i-th frame image and the subsequent predicted trajectory box corresponding to the target trajectory box are determined; finally, the pedestrian trajectory in the video to be processed can be determined based on the target trajectory box of each frame image of the video to be processed. In the present application, since the adjacent occlusion rate corresponding to the pedestrian frame can reflect the probability of occurrence of overlapping and intersecting trajectories, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in each frame image can be used to screen out the target trajectory frame. Therefore, when there are multiple users' trajectories intersecting in the video screen, the target trajectory frame can be used to distinguish each trajectory, thereby avoiding the problem of trajectory confusion and errors in the recognition results of pedestrian trajectories, thereby improving the accuracy of the trajectory recognition results.

[0026] For example, the embodiment of the present invention can be applied to Figure 1 The application scenario shown in FIG. In this scenario, a terminal device 1 and a server 2 may be included.

[0027] The terminal device 1 can be hardware or software. When the terminal device 1 is hardware, it can be various electronic devices with image acquisition functions and supporting communication with the server 2, including but not limited to smart phones, tablet computers, laptop computers and desktop computers; when the terminal device 1 is software, it can be installed in the electronic devices mentioned above. The terminal device 1 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present disclosure do not limit this. The server 2 can be a server that provides various services, for example, a background server that receives requests sent by terminal devices that establish communication connections with it, and the background server can receive and analyze the requests sent by the terminal devices, and generate processing results. The server 2 can be a single server, or a server cluster composed of several servers, or a cloud computing service center, and the embodiments of the present disclosure do not limit this.

[0028] It should be noted that the server 2 can be either hardware or software. When the server 2 is hardware, it can be various electronic devices that provide various services to the terminal device 1. When the server 2 is software, it can be multiple software programs or software modules that provide various services to the terminal device 1, or it can be a single software program or software module that provides various services to the terminal device 1, and this is not limited in the present embodiment.

[0029] The terminal device 1 and the server 2 can be connected to each other through a network. The network can be a wired network connected by coaxial cable, twisted pair, or optical fiber, or a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), infrared, etc., which is not limited in the embodiments of the present disclosure.

[0030] Specifically, a user can input a video to be processed through terminal device 1, which then sends the video to server 2. Server 2 first determines the pedestrian frame corresponding to each frame of the video to be processed, as well as the confidence and adjacent occlusion ratio of the pedestrian frame. Then, based on the confidence and adjacent occlusion ratio of the pedestrian frame in the first frame of the video to be processed, server 2 determines the target trajectory frame for the first frame and the subsequent predicted trajectory frame corresponding to the target trajectory frame. Next, for the i-th frame of the video to be processed, server 2 can determine the target trajectory frame for the i-th frame and the subsequent predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame, the confidence and adjacent occlusion ratio of the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame for the i-1-th frame. Subsequently, server 2 can determine the pedestrian trajectory in the video to be processed based on the target trajectory frame for each frame of the video to be processed. Server 2 returns the pedestrian trajectory in the video to terminal device 1 so that terminal device 1 can display the pedestrian trajectory recognition result corresponding to the video to be processed to the user. In this way, not only the processing speed and efficiency of trajectory determination are improved, but also the accuracy of the trajectory determination results is improved.

[0031] It should be noted that the specific types, quantities and combinations of the terminal devices 1, the server 2 and the network can be adjusted according to the actual needs of the application scenario, and the embodiments of the present disclosure do not limit this.

[0032] It should be noted that the above application scenarios are only shown to facilitate understanding of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0033] Figure 2 This is a flowchart of a trajectory determination method provided by an embodiment of the present disclosure. Figure 2A trajectory determination method can be obtained by Figure 1 The terminal device or server executes. Figure 2 As shown, the trajectory determination method includes:

[0034] S201: Obtain a video to be processed, and determine a pedestrian frame corresponding to each frame image of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame.

[0035] In this embodiment, the video to be processed can be understood as the video for which trajectory determination is required. As an example, the video to be processed can be captured by a surveillance camera installed in a fixed location, captured by a mobile terminal device, or read from a storage device that pre-stores images. It should be noted that in one implementation, the video to be processed can be a segment captured from a video.

[0036] After obtaining the video to be processed, pedestrian frame detection can be performed on each frame image or partial frame image (such as key frame image) of the video to be processed. Pedestrian frame detection can be understood as detecting whether a pedestrian appears in the image. If a pedestrian is detected in the image, a frame (such as a rectangular frame) is used to mark the image area where the pedestrian appears. In this embodiment, this frame can be referred to as a pedestrian frame. It should be noted that each pedestrian frame has a confidence score.

[0037] The confidence score corresponding to a pedestrian frame can reflect the probability of the pedestrian frame's existence. The lower the confidence score corresponding to the pedestrian frame, the lower the probability that the image area corresponding to the pedestrian frame is a pedestrian image. Conversely, the higher the confidence score corresponding to the pedestrian frame, the higher the probability that the image area corresponding to the pedestrian frame is a pedestrian image. In one implementation, if the confidence score corresponding to the pedestrian frame is lower than a preset confidence threshold, the pedestrian frame will be removed.

[0038] The adjacent occlusion rate corresponding to a pedestrian frame can reflect whether the pedestrian frame is occluded by other pedestrian frames and the degree of occlusion. A higher adjacent occlusion rate for a pedestrian frame indicates a larger area of ​​the pedestrian frame occluded by other pedestrian frames and a higher probability of trajectory intersection at that moment. Conversely, a lower adjacent occlusion rate for a pedestrian frame indicates a smaller area of ​​the pedestrian frame occluded by other pedestrian frames and a lower probability of trajectory intersection at that moment.

[0039] S202: Determine a target trajectory frame of the first frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame according to the confidence level and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed.

[0040] Since the first frame can be understood as the first frame in the playback time dimension of the video to be processed, the trajectory needs to be initialized. Specifically, based on the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in the first frame of the video to be processed, the target trajectory frame that can serve as the starting point of the pedestrian trajectory can be filtered out from the pedestrian frame in the first frame of the video to be processed. It should be noted that the starting point of each pedestrian trajectory is only one pedestrian frame.

[0041] In order to facilitate subsequent calculations and improve the accuracy of distinguishing between various trajectories when trajectories intersect, after determining the target trajectory frame of the first frame image, it is necessary to predict the corresponding prediction frame of the target trajectory frame of the first frame image in the second frame image, that is, the position where the target trajectory frame of the first frame image appears in the next frame image (that is, the second frame image). In this embodiment, the position where the target trajectory frame of the previous frame image appears in the adjacent subsequent frame image is called the subsequent predicted trajectory frame. It can be understood that the predicted pedestrian trajectory corresponding to the target trajectory frame of the first frame image can include the subsequent predicted trajectory frame corresponding to the target trajectory frame.

[0042] S203: For the i-th frame image of the video to be processed, determine the target trajectory frame corresponding to the i-th frame image and the next predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame image, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame, and the next predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; i is a positive integer greater than 1.

[0043] In this embodiment, for images other than the first frame, the target trajectory frame corresponding to the frame can be screened based on the subsequent predicted trajectory frame corresponding to the target trajectory frame of the adjacent previous frame, as well as the pedestrian frame corresponding to the frame, the confidence level of the pedestrian frame, and the adjacent occlusion rate. Furthermore, the subsequent predicted trajectory frame corresponding to the target trajectory frame can be determined based on the target trajectory frame corresponding to the frame. In this way, the target trajectory frames corresponding to the second to the last frames of the video to be processed can be determined.

[0044] It can be understood that the target trajectory box corresponding to the i-th frame image can be understood as the position where the target trajectory box corresponding to the i-1-th frame image appears in the i-th frame image, that is, the target trajectory box corresponding to the i-1-th frame image and the target trajectory box corresponding to the i-th frame image can be connected to generate the pedestrian trajectory of the pedestrian.

[0045] S204: Determine the pedestrian trajectory in the video to be processed according to the target trajectory frame of each frame image of the video to be processed.

[0046] Since the target trajectory box of each frame image of the video to be processed can include the upper and lower association relationship between the target trajectory box of the frame image and the target trajectory box of the adjacent frame image of the frame image, the target trajectory box of each frame image of the video to be processed can be connected according to the upper and lower association relationship between the target trajectory boxes of each frame image, so that the pedestrian trajectory of the pedestrian in the video to be processed can be obtained. Therefore, the upper and lower association relationship between the target trajectory boxes of each frame image can be used to distinguish each trajectory when there are multiple users' trajectories intersecting in the video picture, so as to avoid the problem of trajectory confusion and error in the recognition result of the pedestrian trajectory. It should be noted that in this embodiment, the pedestrian trajectories of all pedestrians or some pedestrians in the video to be processed can be determined, and this is not limited here.

[0047] It can be seen that in this embodiment, after obtaining the video to be processed, the pedestrian frame corresponding to each frame image of the video to be processed and the confidence and adjacent occlusion rate corresponding to the pedestrian frame are first determined; then, the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame can be determined based on the confidence and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed; then, for the i-th frame image of the video to be processed, the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame are determined based on the pedestrian frame corresponding to the i-th frame image, the confidence and adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; finally, the pedestrian trajectory in the video to be processed can be determined based on the target trajectory frame of each frame image of the video to be processed. In the present application, since the adjacent occlusion rate corresponding to the pedestrian frame can reflect the probability of occurrence of overlapping and intersecting trajectories, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in each frame image can be used to screen out the target trajectory frame. Therefore, when there are multiple users' trajectories intersecting in the video screen, the target trajectory frame can be used to distinguish each trajectory, thereby avoiding the problem of trajectory confusion and errors in the recognition results of pedestrian trajectories, thereby improving the accuracy of the trajectory recognition results.

[0048] Next, an implementation of "determining a pedestrian frame corresponding to each frame of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame" in S201 will be described. In this embodiment, the step of determining a pedestrian frame corresponding to each frame of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame, may include the following steps:

[0049] S201a: For each frame image of the video to be processed, input the frame image into a trained pedestrian detection model to obtain a number of pedestrian frames corresponding to the frame image and the confidence level corresponding to each pedestrian frame.

[0050] The video to be processed is a video stream, that is, a series of images along the time axis. A trained pedestrian detection model can be used to detect pedestrians in each image. The pedestrian detection model will output several pedestrian frames corresponding to each frame and the confidence score corresponding to each pedestrian frame. For example, the first frame is detected to obtain k0 pedestrian frames and corresponding confidence scores; the second frame is detected to obtain k1 pedestrian frames and corresponding confidence scores; the third frame is detected to obtain k2 pedestrian frames and corresponding confidence scores, and so on. The first and second frames are adjacent video frames, and the playback time of the first frame is earlier than the playback time of the second frame. The second and third frames are adjacent video frames, and the playback time of the second frame is earlier than the playback time of the third frame.

[0051] S201b: Determine, based on the plurality of pedestrian frames, the adjacent occlusion rate corresponding to each pedestrian frame.

[0052] In this embodiment, for each pedestrian frame, the intersection-and-union ratios between the pedestrian frame and each pedestrian frame may be determined first. Then, the adjacent occlusion rate corresponding to the pedestrian frame may be determined based on the intersection-and-union ratios between the pedestrian frame and each pedestrian frame.

[0053] It can be understood that in a frame of image, for each pedestrian frame, the intersection-over-union ratio (i.e., IoU value) between the pedestrian frame and each pedestrian frame other than the pedestrian frame in the frame image can be calculated. It should be noted that the intersection-over-union (IoU) value is calculated by calculating the ratio of the intersection area size and the union area size between the two pedestrian frames. The union of the two pedestrian frames is the entire area containing the two pedestrian frames, and the intersection is the size of the overlap of the intersection of the two pedestrian frames. Then, the ratio of the intersection area size divided by the union area size can be used as the intersection-over-union ratio (i.e., IoU value) between the two pedestrian frames. In addition, the sum of the intersection-over-union ratios between the pedestrian frame and each pedestrian frame can be used as the adjacent occlusion rate corresponding to the pedestrian frame. For example, it can be calculated by the following formula:

[0054] Among them, u dj Refers to the sum of the intersection-and-union ratios of the jth pedestrian frame in the dth video frame and all other pedestrian frames in the video frame; k d is the number of pedestrian boxes in the dth video frame, b dj Refers to the jth pedestrian box in the dth video frame, b di It refers to the i-th pedestrian box in the d-th video frame, and IoU() represents the intersection-over-union ratio calculation function.

[0055] Next, an implementation of "determining the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame based on the confidence level and the adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed" in S202 will be introduced. In this embodiment, the step of determining the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame based on the confidence level and the adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed may include the following steps:

[0056] The pedestrian frame whose confidence and adjacent occlusion rate meet the initialization conditions is used as the target trajectory frame;

[0057] The Kalman filter is used to determine the next predicted trajectory frame corresponding to the target trajectory frame.

[0058] In this embodiment, the initialization condition may be that the confidence is greater than the first confidence threshold, and the adjacent occlusion rate is less than the first adjacent occlusion rate threshold. As an example, the first confidence threshold may be 0.5, and the first adjacent occlusion rate threshold may be 0.2. If the confidence and adjacent occlusion rate of the pedestrian frame in the first frame image meet the above initialization conditions, the pedestrian frame can be used as the target trajectory frame, that is, the starting trajectory frame of the initial trajectory. Then, the Kalman filter can be used to calculate the next position of the target trajectory frame, that is, the subsequent predicted trajectory frame corresponding to the target trajectory frame. For example, assuming that 10 pedestrian frames are detected in the first frame image, and 6 pedestrian frames meet the initialization conditions, they can be initialized to 6 trajectories (that is, 6 target pedestrian frames) and 6 subsequent predicted trajectory frames can be calculated.

[0059] Next, an implementation method of "for the i-th frame image of the video to be processed, determining the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame image, the confidence level and the adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image" in S203 will be introduced. In this embodiment, for the i-th frame image of the video to be processed, determining the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame image, the confidence level and the adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image may include the following steps:

[0060] S203a: Determine the trajectory interlaced candidate frame corresponding to the i-th frame image according to the pedestrian frame corresponding to the i-th frame image, the confidence and adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image.

[0061] If a pedestrian trajectory does not intersect and has no intersection or overlap with other pedestrian trajectories, then this pedestrian trajectory only needs to match one pedestrian frame in each video frame, and the accuracy of the pedestrian trajectory recognition result corresponding to the pedestrian frame is relatively high. If a pedestrian trajectory intersects and overlaps with other pedestrian trajectories, that is, only one pedestrian frame can be matched at the moment when multiple pedestrian trajectories intersect, then trajectory confusion is likely to occur, resulting in trajectory errors. Therefore, in this embodiment, when pedestrian trajectories intersect, a candidate queue is generated to store the pedestrian frames corresponding to the intersection of pedestrian trajectories. It can be understood that in this embodiment, the pedestrian frames in the candidate queue can be referred to as trajectory intersection candidate frames, and how to determine the trajectory intersection candidate frames will be described later. In other words, the trajectory intersection candidate frame is a pedestrian frame that has overlapping and intersecting conditions in the image corresponding to the moment when the pedestrian trajectory intersects. For example, in an image at a certain moment, a pedestrian trajectory may correspond to multiple pedestrian frames that may belong to the pedestrian trajectory. These pedestrian frames that may belong to the pedestrian trajectory are caused by the intersection of trajectories. For example, if there are 10 pedestrian trajectories in an image at a certain moment, among them, 5 pedestrian trajectories correspond to 1 pedestrian frame, 2 pedestrian trajectories correspond to 2 pedestrian frames, 2 pedestrian trajectories correspond to 3 pedestrian frames, and 1 pedestrian trajectory corresponds to 5 pedestrian frames.

[0062] After determining the track intersection candidate frames, pedestrian frame screening can be performed only on the track intersection candidate frames, so that the overlapping pedestrian tracks can be separated by screening the track intersection candidate frames. For the non-track intersection candidate frames, screening is not required, thereby improving the speed and efficiency of pedestrian track recognition. It should be emphasized that in one implementation, in order to avoid an excessive number of track intersection candidate frames in the candidate queue, which would lead to cumbersome data processing, the size of the candidate queue can be set to a preset number, for example 5.

[0063] Next, we will introduce how to determine the candidate frame for trajectory intersection. In this embodiment, the preset first filtering condition of the matching frame can be used to determine the first pedestrian frame in the pedestrian frame corresponding to the i-th frame image. It can be understood that the first pedestrian frame can be a pedestrian frame whose confidence and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the first filtering condition of the matching frame. In one implementation, the first filtering condition of the matching frame can be that the confidence of the pedestrian frame is greater than the second confidence threshold and the adjacent occlusion rate is less than the second adjacent occlusion rate threshold. As an example, the second confidence threshold can be 0.45, and the second adjacent occlusion rate threshold can be 0.3. It should be noted that, in one implementation, the method may further include: if the confidence and adjacent occlusion rate of the pedestrian frame corresponding to the i-th frame image do not meet the first filtering condition of the matching frame but meet the initialization condition, then the pedestrian frame is used as the target trajectory frame corresponding to the i-th frame image; that is, although the confidence and adjacent occlusion rate of the pedestrian frame corresponding to the i-th frame image do not meet the first filtering condition of the matching frame, but meet the initialization condition, it means that the pedestrian frame is suitable as the starting pedestrian frame of a pedestrian trajectory.

[0064] Then, a preset matching box occlusion screening condition can be used to determine a second pedestrian frame within the pedestrian frame corresponding to the i-th image frame; the second pedestrian frame can be understood as a pedestrian frame within the pedestrian frame corresponding to the i-th image frame whose confidence and adjacent occlusion ratio satisfy the matching box occlusion screening condition. In one implementation, the matching box occlusion screening condition can be that the pedestrian frame's confidence is greater than a third confidence threshold and the adjacent occlusion ratio is greater than a third adjacent occlusion ratio threshold. As an example, the third confidence threshold can be 0.4, and the third adjacent occlusion ratio threshold can be 0.3.

[0065] Then, the trajectory interlaced candidate frame corresponding to the i-th frame image can be determined based on the subsequent predicted trajectory frame corresponding to the first pedestrian frame, the second pedestrian frame, and the target trajectory frame of the i-1-th frame image. As an example, the intersection-and-union ratios of the first pedestrian frame and the second pedestrian frame and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image can be first determined, and the first pedestrian frame and the second pedestrian frame whose intersection-and-union ratios meet a first preset intersection-and-union ratio threshold (such as 0.5) are used as the trajectory interlaced candidate frames corresponding to the i-th frame image.

[0066] Specifically, the intersection-and-union ratios of the first pedestrian frame and the subsequent predicted trajectory frames corresponding to each target trajectory frame of the i-1th frame image can be calculated, and the first pedestrian frame and the corresponding subsequent predicted trajectory frame whose intersection-and-union ratios meet the first preset intersection-and-union ratio threshold are selected, and the first pedestrian frame is used as the trajectory interlaced candidate frame corresponding to the i-th frame image; and the intersection-and-union ratios of the second pedestrian frame and the subsequent predicted trajectory frames corresponding to each target trajectory frame of the i-1th frame image can be calculated, and the second pedestrian frame and the corresponding subsequent predicted trajectory frame whose intersection-and-union ratios meet the first preset intersection-and-union ratio threshold are selected, and the second pedestrian frame is used as the trajectory interlaced candidate frame corresponding to the i-th frame image. It should be emphasized that if the intersection-and-union ratios of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image are higher, it means that the similarity and matching degree of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image are higher. Conversely, if the intersection-and-union ratios of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image are lower, it means that the similarity and matching degree of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image are lower.

[0067] Next, a third pedestrian frame can be determined within the pedestrian frame corresponding to the i-th image frame using the preset second matching frame filtering condition. The third pedestrian frame is a pedestrian frame within the pedestrian frame corresponding to the i-th image frame whose confidence and adjacent occlusion ratio satisfy the third matching frame filtering condition. In one implementation, the second matching frame filtering condition can be that the pedestrian frame's confidence is less than a second confidence threshold and greater than a fourth confidence threshold, and that the adjacent occlusion ratio is greater than a second adjacent occlusion ratio threshold. As an example, the second confidence threshold can be 0.45, the fourth confidence threshold can be 0.25, and the second adjacent occlusion ratio threshold can be 0.3.

[0068] Finally, a trajectory interleaving candidate frame corresponding to the i-th frame can be determined based on the third pedestrian frame and the pedestrian frames in the pedestrian frame corresponding to the i-th frame whose confidence and adjacent occlusion ratio do not satisfy the first screening condition of the matching frame or the occlusion screening condition of the matching frame. Specifically, an intersection-over-union (IoU) value can be determined between the third pedestrian frame and the pedestrian frames in the pedestrian frame corresponding to the i-th frame whose confidence and adjacent occlusion ratio do not satisfy the first screening condition of the matching frame or the occlusion screening condition of the matching frame, and the third pedestrian frame whose IoU value satisfies a second preset IoU threshold value can be selected as the trajectory interleaving candidate frame corresponding to the i-th frame. That is to say, it is possible to determine the pedestrian frame in the pedestrian frame corresponding to the i-th frame image that does not match the corresponding subsequent predicted trajectory frame, calculate the intersection-and-union ratio of the pedestrian frame and each third pedestrian frame, and if there is an intersection-and-union ratio in the calculated intersection-and-union ratio that satisfies the second preset intersection-and-union ratio threshold, then the third pedestrian frame whose corresponding intersection-and-union ratio satisfies the second preset intersection-and-union ratio threshold can be used as the trajectory interlaced candidate frame corresponding to the i-th frame image.

[0069] S203b: Determine the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame according to the trajectory interlaced candidate frame and the target trajectory frame of the i-1-th frame image.

[0070] In this embodiment, if the adjacent occlusion rate of the trajectory interlaced candidate frame is less than or equal to the preset occlusion rate threshold (for example, 0.3), it means that the probability that the trajectory interlaced candidate frame is a pedestrian frame in the image where there is pedestrian trajectory interlacing is relatively small, so there is no need to further judge which pedestrian trajectory the trajectory interlaced candidate frame belongs to. The trajectory interlaced candidate frame can be used as the target trajectory frame corresponding to the i-th frame image.

[0071] If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, it means that the probability that the trajectory interlaced candidate frame is a pedestrian frame in the image with pedestrian trajectory interlacing is relatively high, and the trajectory interlaced candidate frame needs to be further judged to which pedestrian trajectory it belongs to. The target trajectory frame corresponding to the i-th frame image can be determined based on the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1-th frame image.

[0072] As an example, if the adjacent occlusion rate of the trajectory intersection candidate frame is greater than the preset occlusion rate threshold, the trajectory intersection candidate frame and the target trajectory frame of the i-1th frame image can be respectively input into the pedestrian re-identification network to obtain the pedestrian features corresponding to the trajectory intersection candidate frame and the pedestrian features of the target trajectory frame of the i-1th frame image. Among them, the pedestrian features can be high-dimensional identity discrimination information that can reflect the pedestrian. It should be noted that in this embodiment, the pedestrian feature extraction operation is performed before and after the pedestrian trajectory is intertwined. It can be understood that the previous moment is one trajectory and the current moment is another trajectory. If one of the trajectories is intertwined, that is, the adjacent occlusion rate of the pedestrian frame of a certain trajectory at the current moment is greater than the preset occlusion rate threshold, that is, the pedestrian features are extracted only when the pedestrian trajectory encounters an intersection, and after the pedestrian features are extracted, the pedestrian features are saved. That is to say, if a pedestrian trajectory corresponds to multiple trajectory interlaced candidate frames, then it is necessary to extract the pedestrian features corresponding to the trajectory interlaced candidate frames, that is, to extract the pedestrian features of the pedestrian in the trajectory interlaced candidate frames of the pedestrian trajectory at the previous moment.

[0073] Then, the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image can be determined based on the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1 frame image. As an example, the inner product (i.e., dot product) of the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1 frame image can be calculated, and the inner product can be used as the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image. It should be noted that the trajectory interlaced candidate frame and the target trajectory frame used to extract pedestrian features are both pedestrian frames corresponding to the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image when no pedestrian trajectory intersection occurs.

[0074] If the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image is greater than a preset similarity threshold (such as 0.5), the trajectory interlaced candidate frame can be used as the target trajectory frame corresponding to the i-1 frame image. It can be understood that if the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image is greater than the preset similarity threshold, all other frames in the candidate queue can be deleted, and only this trajectory interlaced candidate frame can be retained; if the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1 frame image is less than the preset similarity threshold, then this trajectory interlaced candidate frame is deleted from the candidate queue. Finally, based on the target trajectory frame corresponding to the i-1 frame image, the subsequent predicted trajectory frame corresponding to the target trajectory frame can be determined, for example, the subsequent predicted trajectory frame corresponding to the target trajectory frame can be determined using a Kalman filter.

[0075] It can be seen that this implementation performs pedestrian feature extraction conditionally, that is, pedestrian features are extracted from the image area of ​​the trajectory intersection candidate frame only when the pedestrian trajectory encounters an intersection, rather than performing pedestrian feature extraction on all pedestrian frames of all pedestrian trajectories. This can greatly save the computing power of feature extraction, thereby improving the speed and efficiency of pedestrian trajectory determination. In addition, in this embodiment, the timing of pedestrian feature extraction is when the pedestrian trajectory does not intersect. In other words, the pedestrian features collected at this time for the trajectory intersection candidate frame are pedestrian features collected at a time when the adjacent occlusion rate is low. Since there is no occlusion problem, the extracted pedestrian features are relatively accurate.

[0076] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0077] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0078] Figure 3 Schematic diagram of the trajectory determination device provided by the embodiment of the present disclosure. Figure 3 As shown, the trajectory determination device includes:

[0079] The first determining unit 301 is configured to obtain a video to be processed, and determine a pedestrian frame corresponding to each frame of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame;

[0080] The second determining unit 302 is configured to determine a target trajectory frame of the first frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame according to the confidence level and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed;

[0081] The third determining unit 303 is configured to determine, for the i-th frame image of the video to be processed, a target trajectory frame corresponding to the i-th frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame image, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; i is a positive integer greater than 1;

[0082] The fourth determining unit 304 is configured to determine a pedestrian trajectory in the video to be processed according to a target trajectory frame of each frame of the video to be processed.

[0083] Optionally, the first determining unit 301 is configured to:

[0084] For each frame image of the video to be processed, the frame image is input into the trained pedestrian detection model to obtain several corresponding pedestrian frames in the frame image and the confidence level corresponding to each pedestrian frame; based on the several pedestrian frames, the adjacent occlusion rate corresponding to each pedestrian frame is determined.

[0085] Optionally, the first determining unit 301 is configured to:

[0086] For each pedestrian frame, determine the intersection-over-union ratios between the pedestrian frame and each pedestrian frame; and determine the adjacent occlusion rate corresponding to the pedestrian frame based on the intersection-over-union ratios between the pedestrian frame and each pedestrian frame.

[0087] Optionally, the second determining unit 302 is configured to:

[0088] The pedestrian frame whose confidence and adjacent occlusion rate meet the initialization conditions is used as the target trajectory frame;

[0089] The Kalman filter is used to determine the next predicted trajectory frame corresponding to the target trajectory frame.

[0090] Optionally, the third determining unit 303 is configured to:

[0091] Determine the trajectory interlaced candidate frame corresponding to the i-th frame image according to the pedestrian frame corresponding to the i-th frame image, the confidence score and adjacent occlusion rate corresponding to the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image;

[0092] According to the trajectory interlaced candidate frame and the target trajectory frame of the (i-1)th frame image, the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame are determined.

[0093] Optionally, the third determining unit 303 is configured to:

[0094] Determining a first pedestrian frame in the pedestrian frame corresponding to the i-th frame image using a preset first matching frame screening condition; the first pedestrian frame is a pedestrian frame in the pedestrian frame corresponding to the i-th frame image whose confidence level and adjacent occlusion rate meet the first matching frame screening condition;

[0095] Using a preset matching box occlusion screening condition, determine a second pedestrian frame in the pedestrian frame corresponding to the i-th frame image; the second pedestrian frame is a pedestrian frame in the pedestrian frame corresponding to the i-th frame image whose confidence level and adjacent occlusion rate meet the matching box occlusion screening condition;

[0096] Determine a trajectory interlaced candidate frame corresponding to the i-th frame image according to the first pedestrian frame, the second pedestrian frame and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image;

[0097] Determining a third pedestrian frame in the pedestrian frame corresponding to the i-th frame image using the preset second matching frame screening condition; the third pedestrian frame is a pedestrian frame in the pedestrian frame corresponding to the i-th frame image whose confidence level and adjacent occlusion rate meet the third matching frame screening condition;

[0098] According to the pedestrian frame whose confidence and adjacent occlusion rate between the third pedestrian frame and the pedestrian frame corresponding to the i-th frame image do not meet the first screening condition of the matching frame or the occlusion screening condition of the matching frame, the trajectory interlaced candidate frame corresponding to the i-th frame image is determined.

[0099] Optionally, the third determining unit 303 is configured to:

[0100] Determine the intersection-and-union ratios of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image, and use the first pedestrian frame and the second pedestrian frame whose intersection-and-union ratios meet a first preset intersection-and-union ratio threshold as trajectory interlaced candidate frames corresponding to the i-th frame image.

[0101] Optionally, the third determining unit 303 is configured to:

[0102] Determine the intersection-and-union ratio of the third pedestrian frame and the pedestrian frame corresponding to the i-th frame image, whose confidence and adjacent occlusion rate do not meet the first filtering condition of the matching frame or the occlusion filtering condition of the matching frame, and use the third pedestrian frame whose intersection-and-union ratio meets the second preset intersection-and-union ratio threshold as the trajectory interlaced candidate frame corresponding to the i-th frame image.

[0103] Optionally, the third determining unit 303 is configured to:

[0104] If the adjacent occlusion rate of the trajectory interlaced candidate frame is less than or equal to the preset occlusion rate threshold, the trajectory interlaced candidate frame is used as the target trajectory frame corresponding to the i-th frame image;

[0105] If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, the target trajectory frame corresponding to the i-th frame image is determined based on the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1-th frame image;

[0106] According to the target trajectory frame corresponding to the i-th frame image, a subsequent predicted trajectory frame corresponding to the target trajectory frame is determined.

[0107] Optionally, the third determining unit 303 is configured to:

[0108] If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, the trajectory interlaced candidate frame and the target trajectory frame of the i-1th frame image are respectively input into the pedestrian re-identification network to obtain the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1th frame image;

[0109] Determine the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the (i-1)th frame image based on the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the (i-1)th frame image;

[0110] If the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1th frame image is greater than a preset similarity threshold, the trajectory interlaced candidate frame is used as the target trajectory frame corresponding to the i-th frame image.

[0111] Optionally, the device further includes: a fifth determining unit, configured to

[0112] If the confidence and adjacent occlusion rate of the pedestrian frame corresponding to the i-th frame image do not meet the first screening condition of the matching frame and meet the initialization condition, the pedestrian frame is used as the target trajectory frame corresponding to the i-th frame image.

[0113] As can be seen, the trajectory determination device of this embodiment includes: a first determination unit for acquiring a video to be processed and determining a pedestrian frame corresponding to each frame of the video to be processed, as well as the confidence and adjacent occlusion rate of the pedestrian frame; a second determination unit for determining a target trajectory frame for the first frame and a subsequent predicted trajectory frame corresponding to the target trajectory frame based on the confidence and adjacent occlusion rate of the pedestrian frame in the first frame of the video to be processed; a third determination unit for determining, for the i-th frame of the video to be processed, a target trajectory frame for the i-th frame and a subsequent predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame, the confidence and adjacent occlusion rate of the pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame; i is a positive integer greater than 1; and a fourth determination unit for determining the pedestrian trajectory in the video to be processed based on the target trajectory frame of each frame of the video to be processed. In the present application, since the adjacent occlusion rate corresponding to the pedestrian frame can reflect the probability of occurrence of overlapping and intersecting trajectories, the confidence level and adjacent occlusion rate corresponding to the pedestrian frame in each frame image can be used to screen out the target trajectory frame. Therefore, when there are multiple users' trajectories intersecting in the video screen, the target trajectory frame can be used to distinguish each trajectory, thereby avoiding the problem of trajectory confusion and errors in the recognition results of pedestrian trajectories, thereby improving the accuracy of the trajectory recognition results.

[0114] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0115] Figure 4 Schematic diagram of the computer device 4 provided in the embodiment of the present disclosure. Figure 4 As shown, the computer device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned apparatus embodiments are implemented.

[0116] For example, computer program 403 may be divided into one or more modules / units, which are stored in memory 402 and executed by processor 401 to implement the present disclosure. One or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of computer program 403 in computer device 4.

[0117] The computer device 4 can be a desktop computer, a notebook computer, a palmtop computer, a cloud server, etc. The computer device 4 can include but is not limited to a processor 401 and a memory 402. It will be understood by those skilled in the art that Figure 4 This is merely an example of the computer device 4 and does not constitute a limitation on the computer device 4. The computer device 4 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.

[0118] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0119] Memory 402 can be an internal storage unit of computer device 4, such as a hard disk or memory of computer device 4. Memory 402 can also be an external storage device of computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on computer device 4. Furthermore, memory 402 can include both an internal storage unit of computer device 4 and an external storage device. Memory 402 is used to store computer programs and other programs and data required by the computer device. Memory 402 can also be used to temporarily store data that has been output or is about to be output.

[0120] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this disclosure. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0121] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0122] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0123] In the embodiments provided in the present disclosure, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.

[0124] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0125] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0126] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0127] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the scope of protection of the present disclosure.

Claims

1. A trajectory determination method, characterized in that: The method comprises: Acquire a video to be processed, and determine a pedestrian frame corresponding to each frame image of the video to be processed and a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame; Determining a target trajectory frame of the first frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame according to a confidence level and an adjacent occlusion rate corresponding to a pedestrian frame of the first frame image of the video to be processed; For the i-th frame image of the video to be processed, determine the target trajectory frame corresponding to the i-th frame image and the next predicted trajectory frame corresponding to the target trajectory frame based on the pedestrian frame corresponding to the i-th frame image, the confidence score and adjacent occlusion rate corresponding to the pedestrian frame, and the next predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; i is a positive integer greater than 1; Determining a pedestrian trajectory in the video to be processed according to a target trajectory frame of each frame image of the video to be processed; The method of determining the target trajectory frame corresponding to the i-th frame image and the next predicted trajectory frame corresponding to the target trajectory frame according to the pedestrian frame corresponding to the i-th frame image, the confidence level and the adjacent occlusion rate corresponding to the pedestrian frame, and the next predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image includes: Determining a first pedestrian frame in the pedestrian frame corresponding to the i-th frame image using a preset first matching frame screening condition; the first pedestrian frame is a pedestrian frame whose confidence level and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the first matching frame screening condition; Determining a second pedestrian frame in the pedestrian frame corresponding to the i-th frame image using a preset matching frame occlusion screening condition; the second pedestrian frame being a pedestrian frame whose confidence level and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the matching frame occlusion screening condition; Determine a trajectory interlaced candidate frame corresponding to the i-th frame image according to the first pedestrian frame, the second pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; Determining a third pedestrian frame in the pedestrian frame corresponding to the i-th frame image using a preset second matching frame screening condition; the third pedestrian frame being a pedestrian frame whose confidence level and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the third matching frame screening condition; Determining a trajectory interleaving candidate frame corresponding to the i-th frame image based on pedestrian frames whose confidences and adjacent occlusion rates between the third pedestrian frame and the pedestrian frame corresponding to the i-th frame image do not satisfy the first matching frame screening condition or the matching frame occlusion screening condition; According to the trajectory interlaced candidate frame and the target trajectory frame of the (i-1)th frame image, the target trajectory frame corresponding to the (i)th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame are determined.

2. The method according to claim 1, characterized in that The determining of the pedestrian frame corresponding to each frame image of the video to be processed and the confidence level and adjacent occlusion rate corresponding to the pedestrian frame includes: For each frame image of the video to be processed, the frame image is input into the trained pedestrian detection model to obtain several corresponding pedestrian frames in the frame image and the confidence corresponding to each pedestrian frame; based on the several pedestrian frames, the adjacent occlusion rate corresponding to each pedestrian frame is determined.

3. The method according to claim 2, characterized in that The determining, based on the plurality of pedestrian frames, the adjacent occlusion rate corresponding to each pedestrian frame, includes: For each pedestrian frame, determine the intersection-over-union ratios between the pedestrian frame and each pedestrian frame; and determine the adjacent occlusion rate corresponding to the pedestrian frame based on the intersection-over-union ratios between the pedestrian frame and each pedestrian frame.

4. The method according to claim 1, wherein The determining, based on the confidence level and adjacent occlusion rate corresponding to the pedestrian frame of the first frame image of the video to be processed, the target trajectory frame of the first frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame includes: The pedestrian frame whose confidence and adjacent occlusion rate meet the initialization conditions is used as the target trajectory frame; A Kalman filter is used to determine a subsequent predicted trajectory frame corresponding to the target trajectory frame.

5. The method according to claim 1, wherein The determining, based on the first pedestrian frame, the second pedestrian frame, and the subsequent predicted trajectory frame corresponding to the target trajectory frame of the (i-1)th frame image, a trajectory interlaced candidate frame corresponding to the i-th frame image includes: Determine the intersection-and-union ratios of the first pedestrian frame and the second pedestrian frame with the subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1th frame image, and use the first pedestrian frame and the second pedestrian frame whose intersection-and-union ratios meet a first preset intersection-and-union ratio threshold as the trajectory interlaced candidate frames corresponding to the i-th frame image.

6. The method according to claim 1, characterized in that The determining, based on the pedestrian frame whose confidence and adjacent occlusion rate between the third pedestrian frame and the pedestrian frame corresponding to the i-th frame image do not satisfy the first matching frame screening condition or the matching frame occlusion screening condition, a trajectory interlaced candidate frame corresponding to the i-th frame image, includes: Determine the intersection-and-union ratio of the third pedestrian frame and the pedestrian frame corresponding to the i-th frame image, whose confidence and adjacent occlusion rate do not meet the first filtering condition of the matching frame or the matching frame occlusion filtering condition, and use the third pedestrian frame whose intersection-and-union ratio meets the second preset intersection-and-union ratio threshold as the trajectory interlaced candidate frame corresponding to the i-th frame image.

7. The method according to claim 1, characterized in that The determining, based on the trajectory interleaving candidate frame and the target trajectory frame of the (i-1)th frame image, the target trajectory frame corresponding to the (i)th frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame, includes: If the adjacent occlusion rate of the trajectory interlaced candidate frame is less than or equal to a preset occlusion rate threshold, the trajectory interlaced candidate frame is used as the target trajectory frame corresponding to the i-th frame image; If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, determining the target trajectory frame corresponding to the i-th frame image according to the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1-th frame image; According to the target trajectory frame corresponding to the i-th frame image, a subsequent predicted trajectory frame corresponding to the target trajectory frame is determined.

8. The method according to claim 7, characterized in that If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, determining the target trajectory frame corresponding to the i-th frame image according to the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1-th frame image, including: If the adjacent occlusion rate of the trajectory interlaced candidate frame is greater than the preset occlusion rate threshold, the trajectory interlaced candidate frame and the target trajectory frame of the i-1th frame image are respectively input into the pedestrian re-identification network to obtain the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the i-1th frame image; Determining the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the (i-1)th frame image according to the pedestrian features corresponding to the trajectory interlaced candidate frame and the pedestrian features of the target trajectory frame of the (i-1)th frame image; If the similarity between the trajectory interlaced candidate frame and the target trajectory frame of the i-1th frame image is greater than a preset similarity threshold, the trajectory interlaced candidate frame is used as the target trajectory frame corresponding to the i-th frame image.

9. The method according to claim 1, characterized in that The method further comprises: If the confidence and adjacent occlusion rate of the pedestrian frame corresponding to the i-th frame image do not meet the first screening condition of the matching frame and meet the initialization condition, the pedestrian frame is used as the target trajectory frame corresponding to the i-th frame image.

10. A trajectory determination device, characterized in that: The device comprises: A first determining unit is configured to obtain a video to be processed, and determine a pedestrian frame corresponding to each frame of the video to be processed, as well as a confidence level and an adjacent occlusion rate corresponding to the pedestrian frame; A second determining unit is configured to determine a target trajectory frame of the first frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame according to a confidence level and an adjacent occlusion rate corresponding to a pedestrian frame of the first frame image of the video to be processed; a third determining unit, configured to determine, for an i-th frame image of the video to be processed, a target trajectory frame corresponding to the i-th frame image and a subsequent predicted trajectory frame corresponding to the target trajectory frame based on a pedestrian frame corresponding to the i-th frame image, a confidence score and an adjacent occlusion rate corresponding to the pedestrian frame, and a subsequent predicted trajectory frame corresponding to the target trajectory frame of the i-1-th frame image; i is a positive integer greater than 1; A fourth determining unit, configured to determine a pedestrian trajectory in the video to be processed according to a target trajectory frame of each frame of the video to be processed; The third determining unit is specifically configured to: determine a first pedestrian frame in the pedestrian frame corresponding to the i-th frame image by using a preset matching frame first screening condition; the first pedestrian frame is a pedestrian frame whose confidence and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the matching frame first screening condition; determine a second pedestrian frame in the pedestrian frame corresponding to the i-th frame image by using a preset matching frame occlusion screening condition; the second pedestrian frame is a pedestrian frame whose confidence and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the matching frame occlusion screening condition; determine the target trajectory frame corresponding to the i-th frame image according to the subsequent predicted trajectory frame corresponding to the first pedestrian frame, the second pedestrian frame and the target trajectory frame of the i-1-th frame image. trajectory interlaced candidate frame; using the preset matching frame second filtering condition, determine a third pedestrian frame in the pedestrian frame corresponding to the i-th frame image; the third pedestrian frame is a pedestrian frame whose confidence and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image meet the third filtering condition of the matching frame; according to the third pedestrian frame and the pedestrian frame whose confidence and adjacent occlusion rate in the pedestrian frame corresponding to the i-th frame image do not meet the first filtering condition of the matching frame or the occlusion filtering condition of the matching frame, determine the trajectory interlaced candidate frame corresponding to the i-th frame image; according to the trajectory interlaced candidate frame and the target trajectory frame of the i-1-th frame image, determine the target trajectory frame corresponding to the i-th frame image and the subsequent predicted trajectory frame corresponding to the target trajectory frame.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Pedestrian target tracking method based on convolution association network in automatic driving scene

    CN111652903A

  • Behavior recognition device, behavior recognition method, and program

    JP2020087312A