A method and device for target detection and tracking

By using a method of combining deep convolutional networks and significant images in the object detection tracking system, the problems of unstable and inaccurate object detection in the prior art are solved, and more efficient and accurate object detection tracking are achieved.

CN113744304BActive Publication Date: 2025-05-02YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010480879.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-30
Publication Date
2025-05-02
Estimated Expiration
2040-05-30

AI Technical Summary

Technical Problem

The existing target detection and tracking methods have problems such as missing target detection and excessive interference with targets in complex scenarios and environment changes, resulting in unstable and inaccurate detection.

Method used

A target detection and tracking system is adopted, which includes a data acquisition unit, a target detection unit, a preprocessing unit, and a target tracking unit. The image is detected and significant by a deep convolutional network, and the target object is determined and tracked by combining the significance image.

Benefits of technology

It improves the efficiency and accuracy of target detection, enhances the tracking stability and accuracy in the case of missed target detection, and reduces the error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113744304B_ABST
    Figure CN113744304B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for target detection and tracking, which is used to provide a target detection and tracking method with better stability and robustness. The method includes: acquiring at least one image; performing target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, wherein N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predicting the next movement trajectory of the target object in the recommended object, wherein the target object is a recommended object whose intersection ratio is not less than a ratio threshold, and the intersection ratio is the ratio of the intersection and union between the position of the recommended object in the image and the saliency area. Through this method, when determining the target object, combined with the saliency image, the target detection efficiency is effectively improved, and the accuracy and stability of tracking in the case of missed target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic detection, and in particular to a method and device for target detection and tracking. Background Art

[0002] With the continuous development of artificial intelligence, visual technology and other technologies, autonomous driving has gradually become a new trend in smart cars. In the process of autonomous driving, driving safety is particularly important. It is necessary to detect and track pedestrians and vehicles in a timely manner, obtain the category, location information, expected movement trajectory, etc. of pedestrians and vehicles, so as to effectively avoid and drive safely. Therefore, multi-target detection and tracking is one of the important tasks in the autonomous driving system. Figure 1 Figure 2 shows the process of pedestrian detection and matching with tracked targets.

[0003] Among them, in the process of detecting and tracking multiple targets, due to factors such as complex scenes, changes in the surrounding environment, and occlusion of targets, there are often situations where targets are missed or there are too many interfering targets, making it impossible to accurately detect and track targets.

[0004] In summary, the stability, accuracy and robustness of current target detection and tracking methods are poor. Summary of the invention

[0005] The present application provides a method and device for target detection and tracking, so as to provide a target detection and tracking method with better efficiency, accuracy, stability and robustness.

[0006] It should be understood that the method for performing target detection and tracking provided in the embodiments of the present application can be executed by a target detection system.

[0007] In a possible implementation, the target detection and tracking system includes a data acquisition unit, a target detection unit, a preprocessing unit, and a target tracking unit.

[0008] Wherein, the data acquisition unit is used to continuously collect data and send continuous pictures to the target detection unit and the preprocessing unit.

[0009] The target detection unit is used to obtain the current frame image.

[0010] The data acquisition unit may be a camera installed on the smart vehicle and / or at least one sensor with data acquisition and transmission functions.

[0011] The target detection unit is used to perform target detection and saliency detection on the input current frame image to obtain an initialization image and a saliency image of the input image.

[0012] The preprocessing unit is used to receive the initialization image and the saliency image of the current frame image output by the target detection unit 901.

[0013] At the same time, the preprocessing unit is also used to extract the apparent features of the target from the image data input by the data acquisition unit.

[0014] The target tracking unit is used to track the detected pedestrian and vehicle targets according to the category, position, salient area and apparent feature information of the input target.

[0015] The output unit is used to output the location information and tracking ID of the tracking target.

[0016] It should be noted that the target detection and tracking system described in the present application can be a single device with target detection and tracking functions. It can also be a combination of at least two devices, that is, at least two devices are combined into an integral system with target detection and tracking functions. When the target detection and tracking system is a combination of at least two devices, the two devices in the target detection and tracking system can communicate with each other through a communication method selected from Bluetooth, wired connection, or wireless transmission.

[0017] Among them, the target detection and tracking system described in the embodiments of the present application can be installed on a mobile device, such as a vehicle, and is used by the intelligent vehicle to detect and track pedestrians and surrounding vehicles, and is particularly suitable for the detection and tracking of surrounding targets by the intelligent vehicle in an autonomous driving scenario. In addition, in addition to being installed on a mobile device, the target detection and tracking system can also be installed on a fixed device, for example, on a road side unit (RSU) and other equipment, to detect and track surrounding pedestrians and vehicles, and notify the corresponding device of the detection and tracking results.

[0018] In a first aspect, an embodiment of the present application provides a method for target detection and tracking, including:

[0019] Acquire at least one image; perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predict the next movement trajectory of the target object in the recommended objects, where the target object is a recommended object whose intersection ratio is not less than a ratio threshold.

[0020] Among them, optionally, when the above method is executed by the target detection and tracking system, at least one image is acquired through the data acquisition unit; the target detection unit performs target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; the target object among the recommended objects is determined through the preprocessing unit; and the next movement trajectory of the target object among the recommended objects is predicted through the target tracking unit.

[0021] Based on the above method, when performing target detection and tracking, the input image is simultaneously subjected to target detection and saliency detection through a deep convolutional network, so that when determining the target object, the saliency image is combined to effectively improve the target detection efficiency. In addition, the combination of the saliency image can improve the accuracy and stability of tracking in the case of target miss detection.

[0022] Among them, for ease of understanding, those skilled in the art can know based on common knowledge in the art that the position of the recommended object described in the embodiment of the present application in the current frame image can be represented by a recommended rectangular box.

[0023] In a possible implementation manner, the intersection ratio is a ratio of the intersection and union between the position of the proposed object in the image and the salient region.

[0024] In a possible implementation manner, the target object is determined from the suggested objects according to a salient region in the salient image.

[0025] Based on the above method, when determining the target object, the selected suggested object is constrained by the saliency image. On the one hand, it can effectively reduce the search interval and improve the efficiency of determining the target object. On the other hand, it can also be corrected according to the saliency image, effectively avoiding missed detection and reducing the error rate.

[0026] In a possible implementation, the at least N pieces of suggested object information further include appearance features of the suggested objects; wherein the appearance features include one or more of the color, brightness, direction, shape, size, and model of the object.

[0027] Optionally, when the above method is executed by the target detection and tracking system, the apparent features of the proposed object are determined by a preprocessing unit.

[0028] Based on the above method, in order to better determine the target object and improve the target tracking success rate in the embodiment of the present application, the apparent characteristics of the target object are also determined.

[0029] In a possible implementation manner, before predicting the next movement trajectory of the target object in the suggested objects, when there is a new target object in the target objects, a new tracking identification ID is allocated to the new target object.

[0030] Optionally, when the above method is executed by the target detection and tracking system, a new tracking ID is also allocated to the new object through a target tracking unit.

[0031] Based on the above method, the embodiment of the present application effectively distinguishes the target object by adding a tracking ID to the target object. Therefore, after determining that there is a new target object, a new tracking ID is assigned to the new target object.

[0032] In a possible implementation, after allocating a new tracking ID to the new target object, the new target object obtained by the current image detection is added to the target objects obtained by the previous image detection.

[0033] Among them, in the embodiment of the present application, there is a corresponding relationship between the target object and the tracking ID, and the corresponding relationship is stored in a relationship library for subsequent tracking of the motion trajectory of the target object.

[0034] Optionally, when the above method is executed by the target detection and tracking system, the new target object is also stored in the relationship library through the target tracking unit as an object for which the next motion trajectory needs to be predicted. For ease of understanding, in the embodiment of the present application, the object for which the next motion trajectory needs to be predicted is referred to as a tracking object.

[0035] Based on the above method, in an embodiment of the present application, if there is a new target object in the captured image, it means that the new target object did not exist in the relationship library before. Therefore, in order to better perform detection and tracking, the new target object is stored in the relationship library as an object whose next motion trajectory needs to be predicted, that is, a tracking object.

[0036] In a possible implementation manner, the target object obtained by the previous image detection is updated according to the target object obtained by the current image detection.

[0037] It should be noted that, in order to effectively reduce the amount of information stored in the relationship library in the embodiment of the present application, the relevant information of the target object can also be stored in a memory. The memory can be located in the target detection and tracking system; or the memory has a wired or wireless communication connection with the target detection and tracking system.

[0038] In one possible implementation, the first target object obtained by the previous image detection is updated to the first target object obtained by the current image detection, and the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

[0039] Optionally, when the above method is executed by the target detection and tracking system, the first target object obtained by the previous image detection is updated to the first target object obtained by the current image detection through the target tracking unit, and the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

[0040] Based on the above method, an embodiment of the present application provides a detailed process for tracking the motion trajectory of a target object.

[0041] In a possible implementation, the method further includes: deleting a second target object obtained by the last image detection, where the second target object is a target object that has not been successfully matched in two consecutive frames among the target objects obtained by the last image detection.

[0042] Optionally, when the above method is executed by the target detection and tracking system, the target tracking unit determines whether there is a second target object that has not been successfully matched with the target object for two consecutive frames. If so, the target tracking unit ends the detection and tracking of the second target object.

[0043] Based on the above method, in the embodiment of the present application, if there is a second target object that has not been successfully matched with the target object for two consecutive frames, the detection and tracking of the second target object is terminated, thereby effectively reducing unnecessary predictions and saving system overhead.

[0044] In one possible implementation, when the degree of overlap between the position of the target object detected in this image and the position of the target object detected in the previous image is not less than a first threshold, it is determined that the target object detected in this image matches the target object detected in the previous image successfully; or when the similarity between the apparent features of the target object detected in this image and the apparent features of the target object detected in the previous image is not less than a second threshold, it is determined that the target object detected in this image matches the target object detected in the previous image successfully; or when the weight of the first threshold and the second threshold is not less than a third threshold, it is determined that the target object detected in this image matches the target object detected in the previous image successfully.

[0045] Optionally, when the above method is executed by the target detection and tracking system, the target tracking unit determines whether the position of the target object obtained by the current image detection is successful compared with the target object obtained by the previous image detection.

[0046] Based on the above method, the embodiment of the present application provides multiple ways to determine whether the position of the target object obtained by the current image detection is successful compared with the target object obtained by the previous image detection.

[0047] In a possible implementation, the next movement trajectory is predicted based on the third target object obtained by the previous image detection, and the third target object is a target object that is not successfully matched with the target object obtained by the current image detection.

[0048] In one possible implementation, the third target object is subjected to particle filtering prediction to obtain a first predicted position of the third target object in the image for the next step; the degree of overlap between the first predicted position of the third target object in the image and a salient area of ​​the third target object is determined; and when the degree of overlap is not less than a fourth threshold, the first predicted position of the third target object in the image is determined as the position of the third object in the image for the next step.

[0049] In one possible implementation, the third target object is subjected to particle filtering prediction to obtain a predicted position of the third target object in the image for the next step; the degree of overlap between the predicted position of the third target object in the image and a salient area of ​​the third target object is determined; when the degree of overlap is less than a fourth threshold, the third target object is subjected to linear prediction to obtain a second predicted position of the third target object in the image; and the second predicted position is determined as the position of the third target object in the image for the next step.

[0050] Based on the above method, an embodiment of the present application provides a method for predicting the next position of the third target object that has not been successfully matched.

[0051] In a possible implementation, saliency detection is performed on the current frame image through a deep convolutional network to obtain a saliency image; wherein, the deep convolutional network uses an adversarial network to predict the saliency map salGANd to supervise the generated saliency image.

[0052] Optionally, when the above method is executed by the target detection and tracking system, the target detection unit performs saliency detection on the current frame image through a deep convolutional network to obtain a saliency image.

[0053] Based on the above method, the embodiment of the present application processes the current frame image through a deep convolutional network. At the same time, salGANd is used to supervise the generation of a new salient image, which effectively improves the region proposal search efficiency of target detection and the accuracy of the salient image.

[0054] In a possible implementation, after predicting the next movement trajectory of the target object in the suggested objects, the next movement trajectory of the target object is announced through a voice announcement device in the vehicle.

[0055] In one possible implementation, after predicting the next movement trajectory of the target object in the recommended object, when it is determined that there is a dangerous object based on the next movement trajectory of the target object, the vehicle driver or the vehicle automatic driving system is notified to perform emergency avoidance, and the dangerous object is a target object whose next running position is less than a safe distance from the vehicle.

[0056] Based on the above method, the target tracking unit in the embodiment of the present application can promptly confirm whether the target object is a dangerous object by predicting the next movement trajectory of the target object. Therefore, when the target object is a dangerous object, the user can be promptly notified of the fact that the target object is a dangerous object, so that the user who receives the notification can effectively avoid the safety hazards caused by the dangerous object.

[0057] It should be noted that the method described in the embodiments of the present application can be executed locally or in the cloud, and the specific embodiments of the present application are not limited.

[0058] In a second aspect, an embodiment of the present application further provides a target detection and tracking device, which can be used to perform the operations in the above-mentioned first aspect and any possible implementation of the first aspect. For example, the device may include a module or unit for performing each operation in the above-mentioned first aspect or any possible implementation of the first aspect. For example, it includes a transceiver module and a processing module.

[0059] In a third aspect, an embodiment of the present application provides a chip system, comprising a processor and optionally a memory; wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a target detection and tracking device equipped with the chip system executes any method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0060] In a fourth aspect, an embodiment of the present application provides a vehicle, at least one camera, at least one memory, at least one transceiver and at least one processor;

[0061] The camera is used to acquire at least one image;

[0062] The memory is used to store one or more programs and data information; wherein the one or more programs include instructions;

[0063] The transceiver is used for data transmission with the communication device in the vehicle and for data transmission with the cloud;

[0064] The processor is used to perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predict the next movement trajectory of the target object among the recommended objects, where the target object is a recommended object whose intersection ratio is not less than a ratio threshold.

[0065] In a possible implementation, the vehicle further includes a display screen, a voice broadcast device, and at least one sensor;

[0066] The display screen is used to display the motion trajectory of the target object;

[0067] The voice broadcasting device is used to broadcast the movement trajectory of the target object;

[0068] The sensor is used to detect the position and distance of the target object.

[0069] Among them, the camera described in the embodiments of the present application can be a camera of a driver monitoring system, a cockpit camera, an infrared camera, a driving recorder (i.e., a recording terminal), a reversing image camera, etc., and is not limited to the specific embodiments of the present application.

[0070] The shooting area of ​​the camera may be the external environment of the vehicle. For example, when the vehicle is moving forward, the shooting area is the area in front of the front of the vehicle; when the vehicle is reversing, the shooting area is the area behind the rear of the vehicle; when the camera is a 360-degree multi-angle camera, the shooting area may be the 360-degree area around the vehicle, etc.

[0071] The sensor described in the embodiments of the present application can be one or more of a photoelectric sensor, a photosensitive sensor, an ultrasonic sensor, a sound sensor, a ranging sensor, a visual sensor, and an image sensor.

[0072] In the fifth aspect, an embodiment of the present application provides a computer program product, which includes: a computer program code. When the computer program code is executed by a communication module, a processing module or a transceiver, or a processor of a target detection and tracking device, the target detection and tracking device executes any method in the above-mentioned first aspect or any possible implementation of the first aspect.

[0073] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a program, and the program enables a target detection and tracking device to execute any method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 A schematic diagram of a target detection and tracking image provided in an embodiment of the present application;

[0075] Figure 2 It is a schematic diagram of the first existing target detection and tracking method;

[0076] Figure 3 It is a schematic diagram of the second existing target detection and tracking method;

[0077] Figure 4 It is a schematic diagram of the third existing target detection and tracking method;

[0078] Figure 5 A schematic diagram of a target detection and tracking method provided in an embodiment of the present application;

[0079] Figure 6 A schematic diagram of an installation scenario of a target detection and tracking system provided in an embodiment of the present application;

[0080] Figure 7 A schematic diagram of an autonomous driving scenario provided in an embodiment of the present application;

[0081] Figure 8 A schematic diagram of another target detection and tracking system installation scenario provided in an embodiment of the present application;

[0082] Fig. 9 An architecture diagram of a target detection and tracking system provided in an embodiment of the present application;

[0083] Fig.10 A schematic diagram of a neural network model provided in an embodiment of the present application;

[0084] Fig.11 A schematic diagram of generating a saliency map using salGANd supervision provided in an embodiment of the present application;

[0085] Fig.12 A flowchart of a target detection and tracking method provided in an embodiment of the present application;

[0086] Fig.13 A schematic diagram of a collection scenario provided in an embodiment of the present application;

[0087] Fig.14 A schematic diagram of a current frame image collected according to an embodiment of the present application;

[0088] Fig.15 A schematic diagram of a current frame image after target detection provided by an embodiment of the present application;

[0089] Fig.16 A schematic diagram of the intersection ratio of a suggested rectangular frame and a significant area provided in an embodiment of the present application;

[0090] Fig.17 A schematic diagram of a target tracking process provided in an embodiment of the present application;

[0091] Fig.18 A schematic diagram of a tracking object prediction process provided in an embodiment of the present application;

[0092] Fig.19 A particle filter prediction process diagram provided in an embodiment of the present application;

[0093] Fig. 20 A schematic diagram of a first target detection and tracking device provided in this application;

[0094] Fig.21 A schematic diagram of a second target detection and tracking device provided in this application;

[0095] Fig. 22 A schematic diagram of a vehicle provided for this application. DETAILED DESCRIPTION

[0096] With the continuous development of artificial intelligence, visual technology and other technologies, autonomous driving has gradually become a new trend in smart cars. In the process of autonomous driving, driving safety is particularly important. In order to effectively improve driving safety, Figure 1 As shown in the figure, it is necessary to detect and track pedestrians and vehicles in a timely manner, and obtain the categories, location information, expected movement trajectories, etc. of pedestrians and vehicles, so as to effectively avoid and drive safely. Therefore, multi-target detection and tracking is one of the important tasks in the autonomous driving system.

[0097] At present, there are mainly the following methods for target detection and tracking:

[0098] Method 1: Figure 2 As shown in the figure, during the target detection and tracking process, the saliency model is first used to calculate the saliency map of the color, brightness, and direction of the input video, and the saliency map is used to define simple scenes and complex scenes. If it is a simple scene, a rectangular box is established with the saliency area as the tracking target. If it is a complex scene, the tracking box needs to be established manually. When tracking fails, the saliency model is used to detect each frame of the image, and each area in the saliency map is matched with the previous tracking results, and the area with the highest similarity is tracked.

[0099] However, the target detection and tracking solution provided by the method 1 requires the establishment of a saliency model separately, and also requires the determination of the scene complexity based on the saliency map, so as to establish a tracking frame for tracking. In actual application, this method is complicated to operate, and often requires the manual establishment of a tracking frame, which has a high labor cost and a large error problem.

[0100] In addition, when tracking fails, this method needs to repeatedly use the saliency model to calculate the saliency map of each frame. When the features of multiple regions are similar or the tracking target is occluded, it often leads to target tracking failure and lacks tracking prediction function.

[0101] Method 2: Figure 3 As shown in the figure, during the target detection and tracking process, the saliency detection technology is used to distinguish the foreground and background in the image, and the image is contrast enhanced and segmented. Image segmentation destroys the background information in the original rectangular frame and retains the target information, making the characteristics of the target information stronger than the background information. The rectangular frame containing only the foreground information is used as the tracking target.

[0102] However, in the target detection and tracking solution provided by the second method, the modules for target detection and saliency map generation are independent of each other. Only after the object is detected can the saliency map be used to segment the target frame and then used to initialize the tracking target. Once a target is missed, the saliency map cannot be effectively used.

[0103] Method 3: If Figure 4 As shown in the figure, during the target detection and tracking process, the visual attention method is first used to extract the salient area in the first frame of the video, and the background is removed to obtain the moving target. Then, the improved mean shift method is used to track the detected moving target. For example, the original image is decomposed into 8 sub-images of different resolutions, and the features of the sub-images are extracted respectively. Then, each feature map is processed by the central peripheral difference operator and nonlinear normalization operation to synthesize a total salient feature map.

[0104] However, in the target detection and tracking solution provided by the method 3, the process of generating the feature map is complicated. In addition, since the salient region also includes non-moving targets or non-focused targets, in the process of extracting effective salient regions from the feature map and determining whether they include moving targets, and taking the moving targets as tracking targets, invalid tracking targets will be generated.

[0105] In summary, the existing methods for target detection and tracking are relatively complex, unable to perform more accurate target detection and tracking, and the prediction error is relatively large.

[0106] To solve this problem, the embodiments of the present application provide a method and device for target detection and tracking, so as to perform target detection and tracking efficiently and accurately.

[0107] The technical solutions of the embodiments of the present application can be applied to various communication systems, for example: long term evolution (LTE) system, worldwide interoperability for microwave access (WiMAX) communication system, future fifth generation (5G) system, such as new radio access technology (NR), and future communication systems, such as 6G system.

[0108] Taking the 5G system (also known as the New Radio system) as an example, specifically, in the embodiment of the present application, during the detection process, the tasks of target detection and generating a saliency image are simultaneously completed for the input image, so that the initialization image obtained according to the target detection and the saliency image complement and constrain each other to obtain the target rectangular frame of the input image, thereby better improving the accuracy of target detection and tracking. At the same time, it can effectively avoid target tracking failures caused by missed detections during the target detection and tracking process, as well as possible security issues.

[0109] like Figure 5 As shown, the target detection and tracking method provided in the embodiment of the present application includes the following steps:

[0110] Step 500: Acquire at least one image.

[0111] Step 501: Perform object detection and saliency detection on the image through a deep convolutional network to obtain at least N pieces of suggested object information and a saliency image.

[0112] Wherein, N is a positive integer, and the suggested object information includes the position of the suggested object in the image and the category of the suggested object.

[0113] Step 502 , predicting the next movement trajectory of a target object among the suggested objects, wherein the target object is a suggested object whose intersection ratio is not less than a ratio threshold.

[0114] It should be noted that, in order to facilitate subsequent understanding and explanation, in the embodiment of the present application, the position of the suggested object in the current frame image is represented by a suggested rectangular box; the position of the target object in the current frame image is represented by a target rectangular box.

[0115] Among them, there are many ways to determine the intersection ratio in the embodiments of the present application, which are not limited to the following:

[0116] Mode 1: The intersection ratio is the ratio of the intersection and union between the position of the proposed object in the image and the salient region.

[0117] That is, the intersection ratio is the intersection ratio between the suggestion rectangular box corresponding to the suggestion object and the salient area.

[0118] Method 2: Divide the salient area into at least 4 small rectangular frames, and determine a small rectangular frame A from the small rectangular frames, where the small rectangular frame A refers to a small rectangular frame that overlaps with the suggested rectangular frame corresponding to the suggested object.

[0119] Furthermore, the ratio of the number of the small rectangular frames A to the total number of the small rectangular frames is determined as the intersection ratio.

[0120] Method 3: Determine the center position of the suggested rectangular box corresponding to the suggested object, draw a circle with a preset radius using the center position as the point to obtain a circular area, and determine the ratio of the area of ​​the circular area to the area of ​​the salient area as the intersection ratio.

[0121] In addition, in the embodiment of the present application, when the intersection ratio determined by any one of the above methods 1 to 3 is not an integer, it can be rounded off.

[0122] Among them, those skilled in the art can understand that in the process of target detection and tracking, when using the target detection network for target detection, it can be divided into two stages. Among them, in the first stage, multiple prediction rectangular boxes are generated in the collected current frame image (wherein, the method of generating multiple rectangular boxes is not limited in the embodiment of the present application), and then, the top N rectangular boxes with high credibility are selected from the predicted rectangular boxes as the recommended rectangular boxes, and the recommended rectangular boxes refer to the rectangular boxes in which the possibility of the target object is high. Finally, the recommended rectangular box is sent to the second stage for further tracking processing.

[0123] In the embodiment of the present application, a saliency image is also generated in the network of the first stage, so that when the rectangular frame is sent to the second stage, the saliency image is further used to constrain the N suggestion frames. Among them, those that have an intersection with the saliency area and whose intersection ratio is greater than the ratio threshold are sent to the second stage, and those that do not meet the requirement are filtered out. Therefore, the target detection efficiency is effectively improved. In addition, the combination of saliency images can improve the accuracy and stability of tracking in the case of missed target detection.

[0124] To facilitate understanding of the embodiments of the present application, the embodiments of the present application provide a system architecture for target detection and tracking. The target detection and tracking system may be a single device having a target detection and tracking function. It may also be a combination of at least two devices, that is, at least two devices are combined into a system having a target detection and tracking function as a whole. When the target detection and tracking system is a combination of at least two devices, the two devices in the target detection and tracking system may communicate via Bluetooth, wired connection, or a communication method selected from wireless transmission.

[0125] In addition, an optional method of the embodiment of the present application is as follows: Figure 6 As shown, the target detection and tracking system described in the present application can be installed in an intelligent vehicle and used by the intelligent vehicle to detect and track pedestrians and surrounding vehicles, and is particularly suitable for the intelligent vehicle to detect and track surrounding targets in an autonomous driving scenario.

[0126] Among them, Figure 7 As shown, it is a schematic diagram of a possible application scenario of an embodiment of the present application. The above application scenarios can be unmanned driving, automatic driving, intelligent driving, networked driving, etc. The target detection and tracking system can be installed in motor vehicles (such as unmanned vehicles, intelligent vehicles, electric vehicles, digital vehicles, etc.), drones, rail vehicles, bicycles, signal lights, speed measuring devices or network equipment (such as base stations and terminal devices in various systems), etc.

[0127] In addition, in addition to being installed on a mobile device (e.g., installed on a vehicle) to detect and track pedestrians and vehicles around the vehicle, the target detection and tracking system can also be installed on a fixed device, such as Figure 8 As shown, it is installed on a road side unit (RSU) and other equipment to detect and track pedestrians and vehicles in the surrounding area, and notify the corresponding equipment of the detection and tracking results. That is, the embodiment of the present application does not limit the installation location and function of the target detection and tracking system.

[0128] The target detection and tracking system architecture is not limited to the following structure.

[0129] like Fig. 9 As shown, the target detection and tracking system includes a data acquisition unit 900, a target detection unit 901, a preprocessing unit 902, a target tracking unit 903 and an output unit 904. The data acquisition unit 900 is used to continuously collect data and send continuous pictures to the target detection unit 901 and the preprocessing unit 902.

[0130] The target detection unit 900 is used to obtain a current frame image.

[0131] Among them, when the target detection and tracking system is a system architecture composed of multiple devices, the data acquisition unit 900 can be a camera installed on an intelligent vehicle and / or at least one sensor with data collection and transmission functions, etc.

[0132] Among them, the camera described in the embodiments of the present application can be a camera of a driver monitoring system, a cockpit camera, an infrared camera, a driving recorder (i.e., a recording terminal), a reversing image camera, etc., and is not limited to the specific embodiments of the present application.

[0133] The shooting area of ​​the camera may be the external environment of the vehicle. For example, when the vehicle is moving forward, the shooting area is the area in front of the front of the vehicle; when the vehicle is reversing, the shooting area is the area behind the rear of the vehicle; when the camera is a 360-degree multi-angle camera, the shooting area may be the 360-degree area around the vehicle, etc.

[0134] The sensor described in the embodiments of the present application can be one or more of a photoelectric sensor, a photosensitive sensor, an ultrasonic sensor, a sound sensor, a ranging sensor, a visual sensor, and an image sensor.

[0135] The target detection unit 901 is used to perform target detection and saliency detection on the input current frame image at the same time to obtain at least N pieces of suggested object information and a saliency image.

[0136] Among them, in an optional manner in an embodiment of the present application, when the target detection unit 901 processes the input current frame image, in order to improve the efficiency of the candidate region (region proposal) search for target detection, a deep convolutional neural network is applied to perform saliency detection and target detection on the current frame image.

[0137] Exemplarily, the deep convolutional neural network used in the embodiments of the present application is as follows: Fig.10 The neural network model shown.

[0138] Furthermore, in order to improve the accuracy of the acquired saliency image, Fig.11 As shown, during the training process of the neural network model, the output of salGANd is used as supervision for generating a saliency map.

[0139] Among them, the neural network model described in the embodiment of the present application can simultaneously complete the tasks of target detection and saliency detection. The acquired current frame image is input into the deep convolutional neural network, VGG16 is used as the backbone network to extract image features, and the hierarchical relative priority (Relevant Point Nought, RPN) and recommended target region (region of interest, ROI) pooling are used to obtain semantic information of different levels in the Conv4-3 layer and Conv5-3 of VGG16.

[0140] Among them, the Conv4-3 layer features in the neural network model can reduce the loss of target space information and are used to obtain information about small-sized targets. The categories and position boxes output by the two-layer RPN are integrated and processed by the Network Management System (NMS) to generate ROIs. After two-layer ROI Pooling, the ROIs are connected into a final feature for the final target classification and position prediction. Furthermore, the neural network model generates the salient area (obj Mask in the figure) of the corresponding target through the feature map output by the Conv4-3 and Conv5-3 layers and the RPN of each layer. Among them, the salient area serves as the anchor constraint of the target.

[0141] The preprocessing unit 902 is used to receive the initialization image and the saliency image of the current frame image output by the target detection unit 901. At the same time, the preprocessing unit 902 is also used to extract the apparent features of the target from the picture data input by the data acquisition unit 900.

[0142] Among them, the apparent features in the embodiments of the present application include one or more of the color, brightness, direction, shape, size, and model of the target.

[0143] For example, the appearance features include the model, color, size, brand, etc. of the vehicle displayed in the current frame image; and may also include the gender, height, clothing color, etc. of the person displayed in the current frame image.

[0144] The target tracking unit 903 is used to track the detected pedestrian and vehicle targets according to the category, position, salient area and apparent feature information of the input target.

[0145] The output unit 904 is used to output the location information and tracking ID of the tracking target.

[0146] The system architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Furthermore, it is known to those skilled in the art that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. It should be understood that Figure 6 as well as Fig.11 This is a simplified schematic diagram for ease of understanding. The system architecture may also include other devices or other acquisition devices. Figure 6 as well as Fig.11 Not drawn in.

[0147] Some of the terms used in the embodiments of the present application are explained below for easier understanding.

[0148] 1) Autonomous driving refers to a technology that relies mainly on the collaborative efforts of artificial intelligence, visual computing, radar, monitoring devices, and global positioning systems to enable the vehicle system to automatically and safely operate a motor vehicle without the need for driver control.

[0149] 2) Intelligent vehicle refers to a comprehensive vehicle that integrates environmental perception, planning and decision-making, multi-level assisted driving and other functions. The intelligent vehicle shown is equipped with TV cameras, electronic computers and automatic control systems, and it uses computers, modern sensors, information fusion, communications, artificial intelligence and automatic control technologies, and is a typical high-tech complex.

[0150] 3) Computational theory of vision was proposed by American artificial intelligence expert Marr in 1977. The theory holds that vision is a multi-level, bottom-up analysis process, in which a series of different representations of objects are generated, and these representations provide more and more detailed information about the relevant visual environment. In the process of object visual perception, three different levels of representation are generated.

[0151] 4) Target detection refers to locating multiple target objects from an image, including the category and position of the target. The position is generally marked with a rectangular box, which is also called a bounding box.

[0152] 5) Object classification refers to determining the category of the object in the image.

[0153] 6) Multi-target tracking refers to calculating the complete motion trajectory of an object in a continuous image sequence based on the target position in the previous frame. Multi-target tracking means calculating the exact position of the target in the next frame based on a given image sequence, and matching the moving objects in different frames one by one to give the motion trajectories of different objects.

[0154] 7) Data association is a typical processing method often used in multi-target tracking tasks. It determines whether the targets in each frame belong to the same target based on the similarity of their features.

[0155] 8) Visual saliency detection refers to simulating human visual characteristics through algorithms to extract salient areas (i.e. areas of human interest) in images. Humans automatically process areas of interest and selectively ignore areas of no interest. These areas of human interest are called salient areas. Object saliency detection in images is generally achieved by simulating human attention mechanisms.

[0156] 9) Salient area refers to the area of ​​interest or the area containing the target of attention.

[0157] 10) Saliency map refers to the collection of salient areas in an image.

[0158] 11) Region proposal refers to the process of obtaining the region that may contain the features of the target object through RPN (region proposal Net, a network) in the first stage of the two-stage target detection process.

[0159] 12) Feature map refers to the high-order feature map extracted by the neural network during the target detection process.

[0160] Among them, the term "at least one" in the embodiments of the present application refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. The following at least one item (items) or similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one item (items) of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0161] Unless otherwise stated, ordinal numbers such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects. In addition, the terms "including" and "having" in the embodiments of the present application, the claims and the drawings are not exclusive. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, and may also include steps or modules that are not listed.

[0162] Based on the above Fig. 9 The target detection and tracking system, the embodiment of the present application provides a target detection and tracking method, such as Fig.12 As shown, an object detection method provided in an embodiment of the present application includes the following specific steps:

[0163] Step 1200: A data acquisition unit acquires at least one image.

[0164] like Fig.13 As shown, assuming that A is the threshold range for the data acquisition unit in vehicle 1 to acquire images, the current frame image acquired by the data acquisition unit is as follows: Fig.14 shown.

[0165] Step 1201: the data acquisition unit sends the acquired image to a target detection unit and a preprocessing unit.

[0166] Step 1202: The object detection unit performs object detection and saliency detection on the received image to obtain at least N pieces of suggested object information and a saliency image.

[0167] Wherein, N is a positive integer not less than 1, and the suggested object information includes the position of the suggested object in the current frame image and the category of the suggested object. In the embodiment of the present application, determining the category of the object refers to determining whether the object is a pedestrian or a vehicle, or other object.

[0168] Exemplarily, the target detection unit performs target detection on the current frame image 14 to obtain the position and category of the target object in the current frame image 14, that is, Fig.15 , where the dotted box indicates that the target category is a vehicle, and the solid box indicates that the target category is a pedestrian. In addition, after the target detection unit performs saliency detection on the current frame image 14, it also obtains a saliency image corresponding to the current frame image, that is, Fig.16 .

[0169] Step 1203: The target detection unit sends the N suggested object information and the saliency image to a preprocessing unit.

[0170] Step 1204: the preprocessing unit receives the N suggested object information and the saliency image sent by the object detection unit.

[0171] Step 1205: The preprocessing unit determines a recommended object whose position in the image intersects with a salient area in the salient image and whose intersection ratio is not less than a ratio threshold.

[0172] Optionally, the intersection ratio is a ratio of the intersection and union between the position of the proposed object in the image and the salient region.

[0173] Step 1206: The preprocessing unit determines the suggested objects whose intersection ratio is not less than a ratio threshold as target objects.

[0174] In addition, in order to save system storage space and overhead, the pre-processing unit may also delete the suggested objects whose ratio is smaller than a threshold.

[0175] For example, assuming that the ratio threshold is 60%, from the above Fig.15 It can be seen from the content that Fig.15 By detecting the current frame image, 6 suggested rectangular frames are obtained, namely suggested rectangular frames 1 to 6, with 6 significant regions. Fig.16As shown, by determining the intersection ratios of the six suggested rectangular boxes and the six significant regions, it can be seen that only the intersection ratios of the suggested rectangular boxes 1 to 4 and the corresponding significant regions are not less than 60%. Therefore, the suggested boxes 5 and 6 are filtered out. In the embodiment of the present application, the step 1206, when determining the target object, combines the significant image to effectively improve the target detection efficiency. In addition, combining the significant image can improve the accuracy and stability of tracking in the case of missed target detection.

[0176] Step 1207: The preprocessing unit extracts the appearance features corresponding to the target object from the image.

[0177] Among them, in an optional manner in the embodiment of the present application, the embodiment of the present application extracts the apparent features of the target object in the current frame image in the following manner:

[0178] Specifically, the current frame image is converted from an RGB color mode (RGB color mode) color space to a Hue-Saturation-Value (HSV) color space.

[0179] Step 1208: The preprocessing unit sends the target object information to a target tracking unit.

[0180] The target object information includes information such as the position of the target object in the image, the category of the target object, and the appearance characteristics.

[0181] Step 1209: the target tracking unit receives the target object information.

[0182] Step 1210: The target tracking unit predicts the next movement trajectory of the target object.

[0183] Among them, in an embodiment of the present application, the target tracking unit can promptly confirm whether the target object is a dangerous object by predicting the next movement trajectory of the target object, so that when the target object is a dangerous object, the user can be notified of the information that the target object is a dangerous object in a timely manner, so that the user who receives the notification can effectively avoid the safety hazards caused by the dangerous object.

[0184] Exemplarily, taking a driving scene as an example, when the target detection and tracking system in the driving vehicle detects the collected image, it determines four valid target objects, for example, target objects 1 to 4. After predicting the next movement trajectory of the four target objects respectively, the target detection and tracking system learns that the next movement trajectory of target object 1 may enter the safe driving distance of the vehicle, that is, the next position of target object 1 is less than the safe distance of the vehicle, and the remaining three target objects are safe objects. Therefore, the target detection and tracking system outputs the next movement trajectory of the four target objects and prompts that target object 1 is determined to be a dangerous object.

[0185] When the vehicle is autonomously driving, the vehicle's autonomous driving system can make timely adjustments based on the actual situation based on the next movement trajectory of the target object output by the target detection and tracking, and the notification that the target object 1 is a dangerous object, such as emergency avoidance, deceleration, pulling over, and re-planning the driving route based on the movement trajectory of the target object.

[0186] When the vehicle is driven manually, the driver makes timely adjustments based on the actual situation according to the next movement trajectory of the target object announced by the voice broadcasting device in the vehicle and the notification that the target object 1 is a dangerous object, such as emergency avoidance, deceleration, pulling over, and re-planning the driving route according to the movement trajectory of the target object.

[0187] It should be noted that in the embodiment of the present application, the target detection and tracking system can only be used to output the next movement trajectory of the target object, without determining which are dangerous objects and which are safe objects. The execution subject of determining the dangerous object from the target object can be further confirmed by the driving system that receives the movement trajectory, etc., and the embodiment of the present application is not limited here, and any method applicable to the embodiment of the present application belongs to the protection scope of the present application.

[0188] Further, in the embodiment of the present application, based on the step 1210, when the target tracking unit tracks the target object in the target rectangular frame, the specific process can be seen in Fig.17 shown.

[0189] Step 1700, input target object information.

[0190] The information input by the target tracking unit data includes one or more of the following:

[0191] HSV color space map, target object category, target object location and target salient area.

[0192] The target object information is sent by the preprocessing unit to the target tracking unit.

[0193] Exemplarily, the data input may be represented by a vector:

[0194] For example, the detected target can be represented by a vector x = [x left ,y top ,x right ,y bottom ,object_class].

[0195] Step 1701: Determine whether there is a new target object in the input target objects.

[0196] Step 1702: assign a new tracking ID to the determined new target, and store the new target as a new tracking object.

[0197] In the embodiment of the present application, the tracking object can be understood as an object for which the next movement trajectory needs to be predicted. For example, the tracking object is a target object determined by the previous image detection and tracking.

[0198] Among them, in the embodiment of the present application, there is a corresponding relationship between the target object and the tracking ID, and the corresponding relationship is stored in a relationship library for subsequent tracking of the motion trajectory of the target object.

[0199] Specifically, in the embodiment of the present application, the target object detected last time is stored in the relationship library for matching with the target object determined in the next image acquisition. The next time an image is acquired and the corresponding target object is obtained, it is matched with the target object previously stored in the relationship library.

[0200] It is understandable that if there is a new target object in the image collected this time, it means that the new target object did not exist in the relationship library before. Therefore, in order to better detect and track, the new target object is stored in the relationship library as the next tracking object.

[0201] It should be noted that, in the embodiment of the present application, in order to ensure the validity and practicality of the information stored in the relationship library, the relationship library regularly updates the stored data.

[0202] For example, assuming that the smart vehicle has been parked for more than 30 minutes, the surrounding pedestrians and vehicles have been updated several times because of the long parking time. Therefore, after the smart vehicle is started again, the relationship library is formatted and updated.

[0203] Step 1703, according to the target object determined this time, match the target object stored in the relationship library last time, if there is a successfully matched object, execute 1704, if there is an unsuccessfully matched object, execute 1705.

[0204] Further, in an optional manner of the embodiment of the present application, it can be determined that the target object detected by the current image successfully matches the target object detected by the previous image in a variety of ways, which are not limited to the following:

[0205] Determination method 1: When the degree of overlap between the position of the target object detected in this image and the position of the target object detected in the previous image is not less than a first threshold, it is determined that the target object detected in this image matches the target object detected in the previous image successfully.

[0206] Determination method 2: When the similarity between the apparent features of the target object detected in this image and the apparent features of the target object detected in the previous image is not less than a second threshold, it is determined that the target object detected in this image and the target object detected in the previous image are successfully matched.

[0207] Determination method 3: when the weight of the first threshold and the second threshold is not less than the third threshold, it is determined that the target object obtained by the current image detection successfully matches the target object obtained by the previous image detection.

[0208] Step 1704: Update the first target object obtained by the previous image detection to the first target object obtained by the current image detection.

[0209] Among them, the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

[0210] An optional method in the embodiment of the present application is to predict the next position of the target object through Kalman filtering.

[0211] Exemplarily, a Kalman filter is used to predict the displacement of a target between two frames based on a constant velocity model.

[0212] The predictor variable can be represented as X:

[0213] X=[x left ,δx left ,y top ,δy top ,x right ,δx right ,y bottom,δy bottom ]

[0214] Where δ represents the change in the sub-image coordinates of the bounding box. When the target object and the tracked object are associated, the bounding box of the target object updates the tracked object state.

[0215] Step 1705 , predicting the next movement trajectory according to the third target object obtained by the previous image detection, wherein the third target object is a target object that is not successfully matched with the target object obtained by the current image detection.

[0216] Furthermore, the second target object obtained by the last image detection is deleted, where the second target object is a target object that has not been successfully matched in two consecutive frames among the target objects obtained by the last image detection.

[0217] This can effectively reduce unnecessary predictions of the target position.

[0218] In addition, in the embodiment of the present application, in step 1705, the next movement trajectory is predicted based on the third target object obtained by the previous image detection. For details, see Fig.18 The steps shown are used to make predictions.

[0219] Step 1800: Perform particle filtering prediction on the third target object to obtain a first predicted position of the third target object in the image in the next step.

[0220] For example, when performing particle filter prediction in the embodiment of the present application, the specific process can be found in Fig.19 shown.

[0221] S1900, particle initialization.

[0222] Exemplarily, assuming time T=0, the following matrix is ​​obtained:

[0223]

[0224]

[0225] Where F is the transfer matrix, dT = 1, represents the sampling time;

[0226] Q is the covariance matrix of the treatments;

[0227] H is the observation model;

[0228] The covariance matrix of R measurements, where Gr = 1 / 16, pedestrian detection L1 = 1, and vehicle L = 0.5;

[0229] P pre-covariance matrix;

[0230] S1901, particle prediction.

[0231] Exemplarily, the current position is calculated based on the observation state of the previous frame:

[0232] time T = 1, 2, 3…;

[0233] The filter time update equation is:

[0234]

[0235]

[0236] The filter state equation is:

[0237]

[0238]

[0239]

[0240]

[0241]

[0242] Where y represents the change between the measured value (detection result) and the predicted value, which is used to calculate the target position change parameter δ in the x and y directions x ,δ y , and these two parameters are used to calculate the standard deviation of particle distribution σ x ,σ y The target position state is represented by the upper left vertex and lower right vertex of its bounding box rectangle, which can not only accurately represent the target position, but also adjust the size of the rectangle box according to the change of the target size.

[0243] δ x =(x left_d -x left_p )+(x right_d -x right_p )

[0244] δ y =(y top_d -y top_p )+(y bottom_d -y bottom_p )

[0245]

[0246]

[0247] After the target position prediction is completed, the particle weight is calculated according to the predicted position and the detected position, wherein the ratio of the area of ​​the intersection and union of the two rectangular boxes of the bounding box (Intersection over Union, IOU) and the HSV histogram of the target feature are used as the weight.

[0248] The value ranges of the hue channel and saturation channel in the HSV image are [0,180] and [0,255] respectively. The number of histograms is set to 50 and 60 respectively, the histograms of the two channels are calculated, the histograms are normalized respectively, and then concatenated into a total histogram.

[0249] Finally, the particle with the largest weight is taken as the final prediction value.

[0250] When the particle weight is updated, steps S1902, particle weight update, and S1903, resampling, etc. will continue to be executed.

[0251] Step 1801, determine whether the overlap between the first predicted position of the third target object in the image and the salient area of ​​the third target object is not less than a fourth threshold, if so, execute step 1802, if not, execute step 1803.

[0252] Step 1802: determine the first predicted position of the third target object in the image as the next position of the third object in the image.

[0253] Step 1803: linearly predict the third target object to obtain a second predicted position of the third target object in the image.

[0254] Step 1804: determine the second predicted position as the next position of the third target object in the image.

[0255] It should be noted that the method described in the embodiment of the present application can also be applied to the field of monitoring. For example, after determining the next movement trajectory of the target object, the embodiment of the present application can display the next movement trajectory of the target object on a display screen.

[0256] Through the above introduction to the scheme of the present application, it can be understood that, in order to realize the above functions, the above-mentioned implementation devices include hardware structures and / or software units corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0257] like Fig. 20 As shown, an embodiment of the present invention provides a device for target detection and tracking, and the device for target detection and tracking includes a processor 2000, a memory 2001, and a transceiver 2002;

[0258] The processor 2000 is responsible for managing the bus architecture and general processing, and the memory 2001 can store data used by the processor 2000 when performing operations. The transceiver 2002 is used to receive and send data under the control of the processor 2000 to perform data communication with the memory 2001.

[0259] The bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 2000 and memory represented by memory 2001. The bus architecture may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. Processor 2000 is responsible for managing the bus architecture and general processing, and memory 2001 may store data used by processor 2000 when performing operations.

[0260] The process disclosed in the embodiment of the present invention can be applied to the processor 2000, or implemented by the processor 2000. In the implementation process, each step of the process of safe driving monitoring can be completed by the hardware integrated logic circuit or software instructions in the processor 2000. The processor 2000 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware processor to be executed, or a combination of hardware and software modules in the processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 2001, and the processor 2000 reads the information in the memory 2001 and completes the steps of the signal processing process in combination with its hardware.

[0261] In an optional manner of the present application, the processor 2000 is used to read the program in the memory 2001 and execute the following Figure 5 The method flow in S500-S503 shown in FIG. 10 ; or executing the method in S500-S503 shown in FIG. 10 ; Fig.12 The method flow in S1200-S1208 shown in FIG. 1 is as follows: Fig.17 The method flow in S1700-S1705 shown in FIG. 17 is as follows: Fig.18 The method flow in S1800-S1805 shown in FIG. 1800-1805; or executing the method in S1800-1805 shown in FIG. Fig.19 The method flow in S1900-S1903 is shown.

[0262] like Fig.21 As shown, the present invention provides a device for target detection and tracking, which includes a transceiver module 2100 and a processing module 2101.

[0263] The communication module 2100 is used to obtain at least one image;

[0264] The processing module 2101 is used to perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predict the next movement trajectory of the target object among the recommended objects, where the target object is a recommended object whose intersection ratio is not less than a ratio threshold, and the intersection ratio is the ratio of the intersection and union between the position of the recommended object in the image and the saliency area.

[0265] In one implementation, the at least N pieces of suggested object information further include appearance features of the suggested objects;

[0266] The appearance characteristics include one or more of the color, brightness, direction, shape, size, and model of the object.

[0267] In one implementation, the processing module 2101 is further configured to:

[0268] When it is determined that a new target object exists among the target objects, a new tracking identification ID is allocated to the new target object.

[0269] In one implementation, the processing module 2101 is further configured to:

[0270] The target object obtained by the previous image detection is updated according to the target object obtained by the current image detection.

[0271] In one implementation, the processing module 2101 is specifically used for:

[0272] The first target object obtained by the previous image detection is updated to the first target object obtained by the current image detection, wherein the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

[0273] In one implementation, the processing module 2101 is further configured to:

[0274] The new target object obtained by this image detection is added to the target object obtained by the previous image detection.

[0275] In one implementation, the processing module 2101 is further configured to:

[0276] The second target object obtained by the last image detection is deleted, where the second target object is a target object that has not been successfully matched in two consecutive frames among the target objects obtained by the last image detection.

[0277] In one implementation, when the degree of coincidence between the position of the target object detected in the current image and the position of the target object detected in the previous image is not less than a first threshold, it is determined that the target object detected in the current image is successfully matched with the target object detected in the previous image; or

[0278] When the similarity between the apparent features of the target object detected in the current image and the apparent features of the target object detected in the previous image is not less than a second threshold, determining that the target object detected in the current image and the target object detected in the previous image are successfully matched; or

[0279] When the weight of the first threshold and the second threshold is not less than a third threshold, it is determined that the target object obtained by the current image detection successfully matches the target object obtained by the previous image detection.

[0280] In one implementation, the processing module 2101 is specifically used for:

[0281] The next movement trajectory is predicted based on the third target object obtained by the previous image detection, wherein the third target object is a target object that is not successfully matched with the target object obtained by the current image detection.

[0282] In one implementation, the processing module 2101 is specifically used for:

[0283] Perform particle filtering prediction on the third target object to obtain a first predicted position of the third target object in the image in the next step;

[0284] Determining a degree of overlap between a first predicted position of the third target object in the image and a salient area of ​​the third target object;

[0285] When the degree of overlap is not less than a fourth threshold, the first predicted position of the third target object in the image is determined as the next position of the third object in the image.

[0286] In one implementation, the processing module 2101 is specifically used for:

[0287] Perform particle filtering prediction on the third target object to obtain a predicted position of the third target object in the image in the next step;

[0288] Determining a degree of overlap between a predicted position of the third target object in the image and a salient area of ​​the third target object;

[0289] When the overlap degree is less than a fourth threshold, linearly predicting the third target object to obtain a second predicted position of the third target object in the image;

[0290] The second predicted position is determined as the position of the third target object in the image in the next step.

[0291] In one implementation, the processing module 2101 is specifically used for:

[0292] Performing saliency detection on the current frame image through a deep convolutional network to obtain a saliency image;

[0293] Among them, the adversarial network is used in the deep convolutional network to predict the saliency map salGANd to supervise the generation of new salient maps.

[0294] In one implementation, after predicting the next movement trajectory of the target object in the suggested objects, the processing module 2101 is further configured to:

[0295] The next movement trajectory of the target object is announced through a voice announcement device in the vehicle.

[0296] In one implementation, after predicting the next movement trajectory of the target object in the suggested objects, the processing module 2101 is further configured to:

[0297] When a dangerous object is determined to exist according to the next movement trajectory of the target object, the vehicle driver or the vehicle automatic driving system is notified to perform emergency avoidance. The dangerous object is a target object whose next position is less than the safe distance from the vehicle.

[0298] Above Fig.21 The functions of the communication module 2100 and the processing module 2101 shown may be performed by the processor 2000 running the program in the memory 2001 , or may be performed by the processor 2000 alone.

[0299] like Fig. 22 As shown, the present invention provides a vehicle, wherein the device includes at least one camera 2201, at least one memory 2202, at least one transceiver 2203 and at least one processor 2204;

[0300] The camera 2201 is used to obtain at least one image;

[0301] The memory 2202 is used to store one or more programs and data information; wherein the one or more programs include instructions;

[0302] The transceiver 2203 is used for data transmission with the communication device in the vehicle and for data transmission with the cloud;

[0303] The processor 2204 is used to perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N pieces of suggested object information and a saliency image, where N is a positive integer, and the suggested object information includes the position of the suggested object in the image and the category of the suggested object; predict the next movement trajectory of the target object among the suggested objects, where the target object is a suggested object whose intersection ratio is not less than a ratio threshold, and the intersection ratio is the ratio of the intersection and union between the position of the suggested object in the image and the saliency area.

[0304] In one implementation, the vehicle further includes a display screen 2205 , a voice broadcast device 2206 , and at least one sensor 2207 ;

[0305] The display screen 2205 is used to display the motion trajectory of the target object;

[0306] The voice broadcasting device 2206 is used to broadcast the movement trajectory of the target object;

[0307] The sensor 2207 is used to detect the position and distance of the target object.

[0308] In some possible implementations, various aspects of the target detection and tracking method provided by an embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program code is run on a computer device, the program code is used to enable the computer device to execute the steps of the target detection and tracking method according to various exemplary embodiments of the present invention described in this specification.

[0309] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0310] The program product for target detection and tracking according to an embodiment of the present invention may be a portable compact disk read-only memory (CD-ROM) and include program code, and may be run on a server device. However, the program product of the present invention is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by a communication transmission, an apparatus or a device, or used in combination therewith.

[0311] The readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than a readable storage medium, which may transmit, propagate, or transfer a program for use by or in conjunction with a periodic network action system, apparatus, or device.

[0312] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0313] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device.

[0314] The method for target detection and tracking in the embodiment of the present application also provides a computing device readable storage medium, that is, the content is not lost after power failure. The storage medium stores a software program, including program code, and when the program code is run on a computing device, the software program can implement any of the above target detection and tracking solutions in the embodiment of the present application when read and executed by one or more processors.

[0315] The present application is described above with reference to the block diagrams and / or flow charts showing the methods, devices (systems) and / or computer program products according to the embodiments of the present application. It should be understood that a block of the block diagram and / or flow chart diagram and a combination of blocks of the block diagram and / or flow chart diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer and / or other programmable data processing device to produce a machine, so that the instructions executed by the computer processor and / or other programmable data processing device create a method for implementing the functions / actions specified in the block diagram and / or flow chart block.

[0316] Accordingly, the present application may also be implemented with hardware and / or software (including firmware, resident software, microcode, etc.). Furthermore, the present application may take the form of a computer program product on a computer-usable or computer-readable storage medium, which has a computer-usable or computer-readable program code implemented in the medium, for use by an instruction execution system or in conjunction with an instruction execution system. In the context of the present application, a computer-usable or computer-readable medium may be any medium that may contain, store, communicate, transmit, or convey a program for use by an instruction execution system, device, or apparatus, or in conjunction with an instruction execution system, device, or apparatus.

[0317] This application describes multiple embodiments in detail in conjunction with multiple flow charts, but it should be understood that these flow charts and the related descriptions of their corresponding embodiments are only examples for ease of understanding and should not constitute any limitation to this application. Each step in each flow chart is not necessarily required to be executed, for example, some steps can be skipped. In addition, the execution order of each step is not fixed and is not limited to that shown in the figure. The execution order of each step should be determined by its function and internal logic.

[0318] The multiple embodiments described in this application can be arbitrarily combined or the steps can be executed alternately. The execution order of each embodiment and the execution order between the steps of each embodiment are not fixed and are not limited to those shown in the figures. The execution order of each embodiment and the cross-execution order of each step of each embodiment should be determined by their functions and internal logic.

[0319] Although the present application has been described in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are deemed to have covered any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for target detection and tracking, characterized in that: The method comprises: Obtain at least one image; Performing target detection and saliency detection on the image through a deep convolutional network to obtain at least N pieces of suggested object information and a saliency image, where N is a positive integer, and the suggested object information includes a position of the suggested object in the image and a category of the suggested object; Predicting the next movement trajectory of a target object among the suggested objects, wherein the target object is a suggested object whose intersection ratio is not less than a ratio threshold.

2. The method according to claim 1, characterized in that The at least N pieces of suggested object information also include appearance features of the suggested objects; The appearance characteristics include one or more of the color, brightness, direction, shape, size, and model of the object.

3. The method according to claim 1 or 2, characterized in that: Before predicting the next movement trajectory of the target object in the suggested objects, the method further includes: When a new target object exists among the target objects, a new tracking identification ID is allocated to the new target object.

4. The method according to claim 1 or 2, characterized in that: The method further comprises: The target object obtained by the previous image detection is updated according to the target object obtained by the current image detection.

5. The method according to claim 4, characterized in that The updating of the target object obtained by the previous image detection according to the target object obtained by the current image detection includes: The first target object obtained by the previous image detection is updated to the first target object obtained by the current image detection, wherein the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

6. The method according to claim 5, characterized in that The updating of the target object obtained by the previous image detection according to the target object obtained by the current image detection includes: The new target object obtained by this image detection is added to the target object obtained by the previous image detection.

7. The method according to claim 5 or 6, characterized in that: The method further comprises: The second target object obtained by the last image detection is deleted, where the second target object is a target object that has not been successfully matched in two consecutive frames among the target objects obtained by the last image detection.

8. The method according to claim 5 or 6, characterized in that: When the degree of coincidence between the position of the target object detected in the current image and the position of the target object detected in the previous image is not less than a first threshold, determining that the target object detected in the current image and the target object detected in the previous image are successfully matched; or When the similarity between the apparent features of the target object detected in the current image and the apparent features of the target object detected in the previous image is not less than a second threshold, determining that the target object detected in the current image and the target object detected in the previous image are successfully matched; or When the weight of the first threshold and the second threshold is not less than a third threshold, it is determined that the target object obtained by the current image detection successfully matches the target object obtained by the previous image detection.

9. The method according to claim 5 or 6, characterized in that: The method further comprises: The next movement trajectory is predicted based on the third target object obtained by the previous image detection, wherein the third target object is a target object that is not successfully matched with the target object obtained by the current image detection.

10. The method according to claim 9, characterized in that The predicting of the next movement trajectory of the third target object obtained by the previous image detection includes: Perform particle filtering prediction on the third target object to obtain a first predicted position of the third target object in the image in the next step; Determining a degree of overlap between a first predicted position of the third target object in the image and a salient area of ​​the third target object; When the degree of overlap is not less than a fourth threshold, the first predicted position of the third target object in the image is determined as the next position of the third target object in the image.

11. The method according to claim 9, characterized in that The predicting of the next movement trajectory of the third target object obtained by the previous image detection includes: Perform particle filtering prediction on the third target object to obtain a predicted position of the third target object in the image in the next step; Determining a degree of overlap between a predicted position of the third target object in the image and a salient area of ​​the third target object; When the overlap degree is less than a fourth threshold, linearly predicting the third target object to obtain a second predicted position of the third target object in the image; The second predicted position is determined as the position of the third target object in the image in the next step.

12. The method according to claim 1, 2, 5, 6, 10 or 11, characterized in that The performing target detection and saliency detection on the image by using a deep convolutional network includes: Performing saliency detection on the image through a deep convolutional network to obtain a saliency image; Among them, the adversarial network is used in the deep convolutional network to predict the saliency map salGANd to supervise the generation of new salient maps.

13. The method according to claim 1, 2, 5, 6, 10 or 11, characterized in that After predicting the next movement trajectory of the target object in the suggested objects, the method further includes: The next movement trajectory of the target object is announced through a voice announcement device in the vehicle.

14. The method according to claim 1, 2, 5, 6, 10 or 11, characterized in that After predicting the next movement trajectory of the target object in the suggested objects, the method further includes: When a dangerous object is determined to exist according to the next movement trajectory of the target object, the vehicle driver or the vehicle automatic driving system is notified to perform emergency avoidance. The dangerous object is a target object whose next position is less than the safe distance from the vehicle.

15. A device for target detection and tracking, characterized in that: include: A transceiver module, used for acquiring at least one image; A processing module is used to perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predict the next movement trajectory of the target object in the recommended objects, where the target object is a recommended object whose intersection ratio is not less than a ratio threshold.

16. The device according to claim 15, characterized in that The at least N pieces of suggested object information also include appearance features of the suggested objects; The appearance characteristics include one or more of the color, brightness, direction, shape, size, and model of the object.

17. The device according to claim 15 or 16, characterized in that The processing module is also used for: When it is determined that a new target object exists among the target objects, a new tracking identification ID is allocated to the new target object.

18. The device according to claim 15 or 16, characterized in that The processing module is also used for: The target object obtained by the previous image detection is updated according to the target object obtained by the current image detection.

19. The device according to claim 18, characterized in that The processing module is specifically used for: The first target object obtained by the previous image detection is updated to the first target object obtained by the current image detection, wherein the first target object obtained by the previous image detection is the target object that successfully matches the first target object obtained by the current image detection, and the first target object obtained by the current image detection is the target object that successfully matches the first target object obtained by the previous image detection.

20. The device according to claim 19, characterized in that The processing module is specifically used for: The new target object obtained by this image detection is added to the target object obtained by the previous image detection.

21. The device according to claim 19 or 20, characterized in that The processing module is also used for: The second target object obtained by the last image detection is deleted, where the second target object is a target object that has not been successfully matched in two consecutive frames among the target objects obtained by the last image detection.

22. The device according to claim 19 or 20, characterized in that When the degree of coincidence between the position of the target object detected in the current image and the position of the target object detected in the previous image is not less than a first threshold, determining that the target object detected in the current image and the target object detected in the previous image are successfully matched; or When the similarity between the apparent features of the target object detected in the current image and the apparent features of the target object detected in the previous image is not less than a second threshold, determining that the target object detected in the current image and the target object detected in the previous image are successfully matched; or When the weight of the first threshold and the second threshold is not less than a third threshold, it is determined that the target object obtained by the current image detection successfully matches the target object obtained by the previous image detection.

23. The device according to claim 19 or 20, characterized in that The processing module is specifically used for: The next movement trajectory is predicted based on the third target object obtained by the previous image detection, wherein the third target object is a target object that is not successfully matched with the target object obtained by the current image detection.

24. The device according to claim 23, characterized in that The processing module is specifically used for: Perform particle filtering prediction on the third target object to obtain a first predicted position of the third target object in the image in the next step; Determining a degree of overlap between a first predicted position of the third target object in the image and a salient area of ​​the third target object; When the degree of overlap is not less than a fourth threshold, the first predicted position of the third target object in the image is determined as the next position of the third target object in the image.

25. The device according to claim 23, characterized in that The processing module is specifically used for: Perform particle filtering prediction on the third target object to obtain a predicted position of the third target object in the image in the next step; Determining a degree of overlap between a predicted position of the third target object in the image and a salient area of ​​the third target object; When the overlap degree is less than a fourth threshold, linearly predicting the third target object to obtain a second predicted position of the third target object in the image; The second predicted position is determined as the position of the third target object in the image in the next step.

26. The device according to claim 15, 16, 19, 20, 24 or 25, characterized in that The processing module is specifically used for: Performing saliency detection on the image through a deep convolutional network to obtain a saliency image; Among them, the adversarial network is used in the deep convolutional network to predict the saliency map salGANd to supervise the generation of new salient maps.

27. The device according to claim 15, 16, 19, 20, 24 or 25, characterized in that The processing module is also used for: The next movement trajectory of the target object is announced through a voice announcement device in the vehicle.

28. The device of claim 15, 16, 19, 20, 24 or 25, characterized in that The processing module is specifically used for: When a dangerous object is determined to exist according to the next movement trajectory of the target object, the vehicle driver or the vehicle automatic driving system is notified to perform emergency avoidance. The dangerous object is a target object whose next position is less than the safe distance from the vehicle.

29. A device for target detection and tracking, characterized in that: include: one or more processors; Memory; Transceiver; The memory is used to store one or more programs and data information; wherein the one or more programs include instructions; The processor is configured to execute the method according to any one of claims 1 to 14 according to at least one or more programs in the memory.

30. A vehicle, characterized in that: include: at least one camera, at least one memory, at least one transceiver and at least one processor; The camera is used to acquire at least one image; The memory is used to store one or more programs and data information; wherein the one or more programs include instructions; The transceiver is used for data transmission with the communication device in the vehicle and for data transmission with the cloud; The processor is used to perform target detection and saliency detection on the image through a deep convolutional network to obtain at least N recommended object information and a saliency image, where N is a positive integer, and the recommended object information includes the position of the recommended object in the image and the category of the recommended object; predict the next movement trajectory of the target object among the recommended objects, where the target object is a recommended object whose intersection ratio is not less than a ratio threshold.

31. The vehicle according to claim 30, characterized in that The vehicle also includes a display screen, a voice broadcast device and at least one sensor; The display screen is used to display the motion trajectory of the target object; The voice broadcasting device is used to broadcast the movement trajectory of the target object; The sensor is used to detect the position and distance of the target object.

32. A computer-readable storage medium, characterized in that: It includes computer instructions. When the computer instructions are executed on a target detection and tracking device, the target detection and tracking device executes the method steps as claimed in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Target tracking method, device and system

    CN110706247A

  • Layered multi-target tracking method based on significance detection

    CN111027505A