Dynamic target detection method, alarm control method and related equipment

By combining the detection results of mobile objects and object detection results in the video, and using deep learning technology to determine dynamic targets, the problem of low recognition accuracy of existing automobile sentry mode in dense targets and light changing environments is solved, improving the accuracy of dynamic target recognition and the safety performance of the vehicle.

CN120032344APending Publication Date: 2025-05-23BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093074.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing automobile sentry mode has low accuracy in dynamic target recognition in environments with dense targets and strong lighting changes, resulting in poor vehicle safety performance.

Method used

By obtaining the detection results of moving objects and object detection results in the video, combined with deep learning technology, dynamic targets are determined and the accuracy of dynamic target recognition is improved.

Benefits of technology

It improves the accuracy of dynamic target recognition, reduces interference from complex driving environments on dynamic target recognition, ensures that the vehicle promptly calls the alarm when it detects a dynamic target, and ensures the safety of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032344A_ABST
    Figure CN120032344A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic target detection method and device, an alarm control method and device, electronic equipment, a vehicle, a storage medium and a program product. The method comprises the steps of obtaining a moving object detection result in a video; obtaining a target detection result in the video; and determining a dynamic target based on the moving object detection result and the target detection result. The dynamic target detection accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automatic control technology, and in particular to a dynamic target detection method, an alarm control method, a device, an electronic device, a vehicle, a storage medium and a product. Background Art

[0002] With the development of automobile intelligence and interconnection, the safety performance of automobiles has also been greatly improved. Among them, the car sentry mode is regarded as an important part of automobile safety. It monitors the environment around the vehicle, activates the car's security anti-theft mechanism, and provides real-time alarm function to prevent the vehicle from being stolen or damaged.

[0003] However, the existing sentinel mode is generally easily affected by the environment in which the vehicle is located, such as an environment with dense targets and an environment with drastic changes in lighting, resulting in low recognition accuracy of dynamic targets and poor safety performance of the vehicle. Summary of the invention

[0004] The embodiments of the present application provide a dynamic target detection method, an alarm control method, an apparatus, an electronic device, a vehicle, a storage medium and a program product to improve the accuracy of dynamic target recognition so as to at least partially solve the above-mentioned technical problems.

[0005] In order to achieve the above object, according to a first aspect of the present application, a dynamic target detection method is provided, comprising:

[0006] Get the moving object detection results in the video;

[0007] Obtaining target detection results in the video;

[0008] Based on the moving object detection result and the target detection result, a dynamic target is determined.

[0009] Optionally, the moving object detection result includes first moving area information of the moving object in the first frame image in the video;

[0010] The step of obtaining a moving object detection result in a video includes:

[0011] The first moving area information is obtained through a first recognition model based on a first frame image in a video and adjacent frame images of the first frame image.

[0012] Optionally, obtaining the first moving area information based on a first frame image and adjacent frame images of the first frame image in a video by using a first recognition model includes:

[0013] Extracting feature maps of the first frame image and the adjacent frame images respectively through the first recognition model;

[0014] Based on the feature map, obtaining a target thermal map of the first frame image;

[0015] According to the target thermal map, first moving area information of the moving object in the first frame image is determined.

[0016] Optionally, extracting feature maps of the first frame image and the adjacent frame images respectively by using the first recognition model includes:

[0017] Through the convolutional neural network with shared weights in the first recognition model, feature extraction is performed on the first frame image and the adjacent frame images respectively to obtain feature maps corresponding to the first frame image and the adjacent frame images.

[0018] Optionally, obtaining a target heat map of the first frame image based on the feature map includes:

[0019] The feature map of the first frame image is subtracted from the feature map of the adjacent frame image to obtain a target thermal map.

[0020] Optionally, determining first moving area information of the moving object in the first frame image according to the target heat map includes:

[0021] Determine a target pixel whose pixel value is a local maximum value among the pixels according to the pixel value of the pixel in the target heat map and the neighborhood of the pixel;

[0022] According to the target pixel point, first moving area information of the moving object in the first frame image is determined.

[0023] Optionally, the target detection result includes second moving area information of the target;

[0024] The obtaining of the target detection result in the video includes:

[0025] Target detection is performed based on the first frame image in the video using a second recognition model, and second moving area information of the target in the first frame image is output.

[0026] Optionally, the determining the dynamic target based on the moving object detection result and the target detection result includes:

[0027] Determining distance information between the moving object and the target according to the first moving area information and the second moving area information;

[0028] A dynamic target in the first frame of image is determined according to the distance information.

[0029] Optionally, the moving object detection result further includes a target heat map corresponding to the first frame image;

[0030] The step of determining the dynamic target in the first frame image according to the distance information includes:

[0031] If the distance information is less than a preset distance threshold, the dynamic target in the first frame image is determined according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information.

[0032] Optionally, determining the dynamic target in the first frame image according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information includes:

[0033] Determine a first number of pixels corresponding to the second moving area information in the target heat map, and a second number of non-zero pixels in the pixels;

[0034] If the ratio between the first number and the second number is smaller than a preset threshold, it is determined that the target corresponding to the second moving area information is a dynamic target.

[0035] Optionally, the method further comprises:

[0036] If the ratio between the first number and the second number is greater than or equal to the preset threshold, it is determined that the target corresponding to the second moving area information is not a dynamic target.

[0037] Optionally, the first moving area information includes a center point position of a moving object in the first frame image, and the second moving area information includes detection frame information.

[0038] Optionally, the method further comprises:

[0039] Performing moving object detection based on the sample video using the first recognition model to be trained, obtaining a feature map of the image in the sample video;

[0040] Obtaining a predicted heat map of the sample video based on a feature map of adjacent frame images using a first recognition model to be trained;

[0041] The first recognition model to be trained is trained according to the predicted heat map to obtain a first recognition model.

[0042] Optionally, the training of the first recognition model to be trained according to the predicted heat map to obtain the first recognition model includes:

[0043] Obtaining the loss of the first recognition model according to the predicted heat map and the preset real heat map;

[0044] The first recognition model is trained based on the loss to obtain a trained first recognition model.

[0045] Optionally, the step of generating the real heat map includes:

[0046] For a video frame in the video, obtain a circumscribed rectangular frame of a moving object in the video frame, and obtain a center point position of the circumscribed rectangular frame and a pixel point position in the video frame;

[0047] The thermal value of the pixel point is determined according to the pixel point position and the center point position to obtain a real thermal map of the video frame.

[0048] Optionally, determining the thermal value of the pixel point according to the pixel point position and the center point position to obtain a real thermal map of the video frame includes:

[0049] Determine the target circumscribed rectangular frame corresponding to the pixel point;

[0050] According to the position of the pixel point and the center point position of the target circumscribed rectangular frame, the thermal value of the pixel point is determined to obtain a real thermal map of the video frame.

[0051] Optionally, determining a target circumscribed rectangular frame corresponding to the pixel point includes:

[0052] Determine the circumscribed rectangular frame to which the pixel point belongs;

[0053] If there is one circumscribed rectangular frame, the circumscribed rectangular frame is the target circumscribed rectangular frame;

[0054] If there are multiple circumscribed rectangular frames, a target circumscribed rectangular frame is selected from the circumscribed rectangular frames according to the priorities of the circumscribed rectangular frames.

[0055] Optionally, determining the thermal value of the pixel point according to the pixel point position and the center point position of the target circumscribed rectangular frame to obtain a true thermal map of the video frame includes:

[0056] Obtaining distance information between the pixel point position and the center point position of the target circumscribed rectangular frame;

[0057] The thermal value of the pixel is determined according to the distance information to obtain a real thermal map of the video frame.

[0058] According to a second aspect of the present application, there is provided an alarm control method, the method comprising:

[0059] Determine whether to alarm based on the relative position between the dynamic target and the alarm subject determined by the dynamic target detection method.

[0060] According to a third aspect of the present application, a dynamic target detection device is provided, comprising:

[0061] The first acquisition module is used to obtain the moving object detection result in the video;

[0062] A second acquisition module is used to obtain the target detection result in the video;

[0063] A determination module is used to determine a dynamic target based on the moving object detection result and the target detection result.

[0064] According to a fourth aspect of the present application, there is provided an alarm control device, comprising:

[0065] The alarm module is used to determine whether to alarm according to the relative position between the dynamic target determined by the dynamic target detection device and the vehicle or the alarm control method.

[0066] In a fifth aspect, this embodiment further provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.

[0067] In a sixth aspect, this embodiment also provides a vehicle, which includes at least one of the above-mentioned dynamic target detection device, alarm control device and electronic device.

[0068] In a seventh aspect, this embodiment further provides a computer-readable storage medium, which includes a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the above method.

[0069] In the eighth aspect, this embodiment also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the steps of the above method.

[0070] In summary, in the embodiment of the present application, through the above technical solution, the vehicle can obtain the moving object detection result in the first frame image of the video, and then determine the dynamic target in the first frame image based on the moving object detection result and the target detection result of the first frame image. It can be seen that the present application combines the moving object detection result based on the video and the target detection result based on the single image to determine the dynamic target in the first frame image, and can accurately detect the dynamic target, avoiding the interference of the complex driving environment on the dynamic target recognition.

[0071] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative work.

[0073] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same figure numbers represent the same parts in the following description.

[0074] Figure 1 is a schematic diagram of an existing sentinel mode provided in an exemplary embodiment of the present application;

[0075] Figure 2 is a schematic diagram of a target detection process provided in an exemplary embodiment of the present application;

[0076] Figure 3 is a first schematic diagram of a motion detection model provided in an exemplary embodiment of the present application;

[0077] Figure 4 is a second schematic diagram of a motion detection model provided in an exemplary embodiment of the present application;

[0078] Figure 5 is a schematic diagram of a dynamic target detection device provided in an exemplary embodiment of the present application;

[0079] Figure 6 is a schematic diagram of an alarm device provided in an exemplary embodiment of the present application;

[0080] Figure 7 It is a schematic diagram of the architecture of an electronic device provided in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0081] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0082] Combined with the above background technology description of this application, with the development of automobile intelligence and interconnection, the safety performance of automobiles has also been greatly improved. Among them, the car sentry mode is regarded as an important part of automobile safety. It monitors the surrounding environment of the vehicle, activates the safety anti-theft mechanism of the vehicle, and provides a real-time alarm function to prevent the vehicle from being stolen or damaged.

[0083] How to realize the precise sentinel function and help the car owners avoid the losses caused by accidents during the parking of the vehicle is a topic worthy of study.

[0084] like Figure 1 As shown, the current mainstream sentinel mode includes three modules: sensor system, data processing, and early warning system:

[0085] Sensor system: Sentry mode relies on a series of sensors to collect information around the vehicle. These sensors include radar, cameras, ultrasonic sensors, etc. With these sensors, objects around the vehicle, road conditions, and the behavior of other vehicles or pedestrians can be collected in real time.

[0086] Data processing: The information collected by the sensor will be further transmitted to the data processing module for analysis and processing. This module uses machine learning algorithms to determine whether there are potential dangers around the vehicle by learning and training on a large amount of data.

[0087] Early warning system: Once the data processing module detects danger, it will send a signal to the early warning system, which will then transmit it to the car owners through sound, light, vibration, etc., to remind them of potential dangers and avoid damage to the vehicle.

[0088] The core of Sentinel Mode lies in whether the intelligent algorithm in the data processing module is implemented accurately. The current data processing methods include the following:

[0089] 1) Use a vibration sensor to send a warning message of the corresponding level according to the vibration level when the vehicle is detected to be vibrating;

[0090] 2) Using ultrasonic radar, when the distance between the specified object and the vehicle is detected to be less than a pre-defined threshold, a warning message of the corresponding level is sent according to the distance level;

[0091] 3) Use the on-board camera and pure visual algorithm to detect whether there are any suspicious targets around the vehicle, and send corresponding level of warning information based on the results returned by the algorithm.

[0092] Among them, the camera-based visual algorithm has become the main means to implement the sentinel mode due to its low cost and low power consumption. The main algorithms include two categories:

[0093] 1) Algorithm based on target detection + multi-target tracking.

[0094] This type of algorithm first trains a target detection model to detect the predefined position of the target object in the video; then uses a multi-target tracking algorithm to match the detection frames of the previous and next frames in the video to obtain the complete motion trajectory of the target; finally, based on the trajectory, it determines whether the target has entered the target area and is a suspicious object.

[0095] The difficulty of this type of algorithm lies in the real-time and accuracy of the multi-target tracking algorithm. In places where targets are densely packed, the multi-target tracking algorithm is very likely to have mismatches due to occlusion, thus affecting the final judgment.

[0096] 2) Algorithms based on dynamic target detection.

[0097] The core of this type of algorithm is to detect whether there is any object movement in the previous and next frames of the video, that is, to separate the changed area or dynamic target from the background image. It mainly uses traditional image processing algorithms, such as frame difference method, background elimination method, optical flow method, etc.

[0098] This type of algorithm is generally simple to implement and does not require a lot of computing power. At the same time, the difficulty lies in the fact that the robustness of the algorithm is difficult to ensure and it is easily affected by external factors such as changes in lighting, which can lead to false alarms.

[0099] In order to solve the above problems and improve the recognition accuracy of dynamic targets, such as pedestrians and other vehicles, this application provides a dynamic target detection method based on deep learning based on the vehicle-mounted camera, which avoids the high computing power and mismatching of the multi-target tracking algorithm, and is different from the low precision based on traditional image processing.

[0100] It is worth noting that the present application adopts a new two-stage vehicle sentinel mode implementation method and a new dynamic target detection method, which integrates dynamic target detection with target detection, improves the accuracy of the dynamic target detection algorithm, and then promptly issues an alarm when suspicious targets appear around the vehicle.

[0101] Considering that the current vehicle sentry mode based on pure vision solutions is mainly implemented by combining target detection with multi-target tracking, the real-time and accuracy of the multi-target tracking algorithm are difficult to guarantee, especially in places with dense targets. The multi-target tracking algorithm may have mismatches due to occlusion, and may easily cause string ID or trajectory loss problems, which may lead to false alarms of the sentinel.

[0102] Therefore, the two-stage sentinel method in this application no longer uses multi-target tracking, but instead uses a framework that combines dynamic target detection with target detection. First, the moving objects in each frame of the video are detected through a moving object detection model, and then the moving objects are fused with the target detection results to obtain the final dynamic target. Finally, based on the position of the dynamic target, the sentinel is judged whether to sound an alarm.

[0103] Specifically, the dynamic target detection method in this application can be applied to vehicles, such as Figure 2 As shown, the dynamic target detection method in the present application may include the following steps:

[0104] S10, obtaining a moving object detection result in the video;

[0105] In this embodiment, after the sentry mode is turned on, the vehicle can collect a video of the environment in which the vehicle is currently parked, and then determine the moving object detection result in the first frame image based on at least two adjacent frames of images in the video. The first frame image can be any first frame image in the video.

[0106] In addition, the video may also be a historical parking video, and there is no limitation on this.

[0107] It is understandable that this embodiment can detect dynamic targets based on images in the video, such as pedestrians, vehicles, animals, etc. Compared with the multi-target tracking algorithm, which is very easy to have mismatching due to occlusion in places with dense targets, this embodiment can use an algorithm based on dynamic target detection to detect dynamic targets.

[0108] S20, obtaining the target detection result in the video;

[0109] In this embodiment, the vehicle can also obtain target detection results in the video.

[0110] It should be noted that, in this embodiment, the target type to be detected can be set according to the preset suspicious dynamic target, such as pedestrians, vehicles, animals, etc., and the target detection model can be trained using yolov5 (target detection algorithm), and the target detection result in the first frame image can be identified using the target detection model. The target detection result can be understood as the detection frame in the first frame image. For example, when the target type is a pedestrian, if there is a pedestrian in the first frame image, then the detection frame can be the area framed by the detection frame containing the pedestrian. It can be understood that the target detection based on the first frame image in this embodiment is actually a static detection, which only detects the target in the current frame.

[0111] S30, determining a dynamic target based on the moving object detection result and the target detection result.

[0112] After obtaining the target detection result and the moving object detection result in the video according to the video, the vehicle can determine the dynamic target in the first frame image based on the moving object detection result and the target detection result.

[0113] Specifically, for example, the vehicle can fuse the moving object detection results and the target detection results. In this way, by fusing the moving object detection results and the target detection results, the final dynamic target can be obtained. By fusing the dynamic target and static target detection results, it can be effectively applied to target-dense places and can achieve high-precision dynamic target detection without high computing power.

[0114] Therefore, in the embodiment of the present application, the vehicle can determine the moving object detection result and the target detection result based on the video, and then determine the dynamic target based on the moving object detection result and the target detection result. It can be seen that the present application combines dynamic target detection and static target detection to determine the dynamic target in the video, so that the dynamic target can be accurately detected, and the interference of the complex parking environment on the dynamic target recognition can be avoided, so that the alarm can be timely issued when the dynamic target is detected, thereby ensuring the safety of vehicle driving.

[0115] In one embodiment, in the above S10, “obtaining a moving object detection result of a first frame image in a video” may include:

[0116] S101, obtaining the first moving area information based on a first frame image in a video and adjacent frame images of the first frame image through a first recognition model.

[0117] It should be noted that, in this embodiment, the moving object detection result may include first moving area information of the moving object in the first frame image, and the target detection result may include second moving area information of the target.

[0118] For example, the first moving area information may specifically be the center point position of the moving object in the first frame image, and the second moving area information may specifically be the detection frame information of the target in the adjacent frame image.

[0119] And, if Figure 3 As shown, the motion detection model in this embodiment may include a parallel moving object detection model (i.e., the first recognition model in this embodiment) and a target detection model (i.e., the second recognition model in this embodiment), wherein the moving object detection model outputs a moving object detection result based on at least two frames of images, and the target detection model outputs a target detection result based on the first frame of image. In this way, the motion detection model can output a dynamic target in the first frame of image based on the moving object detection result and the target detection result.

[0120] On this basis, the first recognition model in this embodiment is a moving object detection model, and the moving object detection model can obtain a moving object detection result in the first frame image based on at least two adjacent frames of images.

[0121] In one embodiment, in the above S101, “obtaining the first moving area information based on a first frame image and adjacent frame images of the first frame image in a video by using a first recognition model” may include:

[0122] S1011, respectively extracting feature maps of the first frame image and the adjacent frame images by using the first recognition model;

[0123] S1012, obtaining a target thermal map of the first frame image based on the feature map;

[0124] S1013: Determine first moving area information of the moving object in the first frame image according to the target heat map.

[0125] In this embodiment, the convolutional neural network in the moving object detection model can perform feature processing on at least two adjacent frames of images to obtain a feature map, and output a target heat map based on the feature map, and then determine the first moving area information of the moving object in the first frame of image based on the target heat map.

[0126] In a specific embodiment, in the above S1011, “using the first recognition model to extract feature maps of the first frame image and the adjacent frame images respectively” may include:

[0127] Step a: extract features from the first frame image and the adjacent frame images respectively through a convolutional neural network with shared weights in the first recognition model to obtain feature maps corresponding to the first frame image and the adjacent frame images.

[0128] In this embodiment, two adjacent frames of images include a first frame of image frame1 and an adjacent frame of image frame2, and two convolutional neural networks with shared weights in the moving object detection model are input simultaneously, including a first convolutional neural network and a second convolutional neural network, such as Figure 4 As shown, the first frame image frame1 is input into the first convolutional neural network, and the adjacent frame image frame2 is input into the second convolutional neural network, so as to extract features of the first frame image frame1 and the adjacent frame image frame2 respectively, and obtain the first feature image feature_map1 corresponding to the first frame image frame1, and the second feature image feature_map2 corresponding to the adjacent frame image frame2. Then, according to the first feature image feature_map1 and the second feature image feature_map2, the target heat map can be obtained:

[0129] predict_heatmap=feature_map2-feature_map1.

[0130] In a specific embodiment, in the above S1012, “obtaining a target heat map of the first frame image based on the feature map” may include:

[0131] Step b: performing bitwise difference between the feature map of the first frame image and the feature map of the adjacent frame image to obtain a target thermal map.

[0132] In this embodiment, the above feature_map1 and feature_map2 can be bitwise subtracted to obtain the target heat map predicted by the moving object detection model. For example, feature_map1 and feature_map2 are matrices, and the values ​​at the same positions of the two matrices are directly subtracted to obtain the target heat map.

[0133] In one embodiment, in the above S1013, “determining first moving area information of the moving object in the first frame image according to the target heat map” may include:

[0134] Step c, determining a target pixel point whose pixel value is a local maximum value among the pixels according to the pixel value of the pixel point in the target heat map and the neighborhood of the pixel point;

[0135] Step d: determining first moving area information of the moving object in the first frame image according to the target pixel point.

[0136] In this embodiment, the moving object detection model can output a target heat map predict_heatmap with the same size as the input image, and then the target heat map predict_heatmap can be processed according to the following steps:

[0137] (1) Truncate the original output predict_heatmap:

[0138]

[0139] Where (x, y) is the coordinate value of any point on the heatmap, and θ is the set threshold. In this embodiment, the truncation operation can remove low-value noise, which may represent unimportant feature responses or noise, so that the features displayed in the heatmap are more prominent and the visualization effect can be enhanced, making the color changes in the heatmap more obvious, which is convenient for observation and analysis;

[0140] (2) Calculate the center point of the moving object: Find the local maximum on the target heat map predict_heatmap, that is, traverse each coordinate point on the heat map. If the value of the point is greater than the value in the upper, lower, left, and right ε*ε neighborhood (preset parameters), then it means that the point is the target pixel point of the local maximum, that is, the center point of the moving object, where ε*ε represents a square area with a side length of ε.

[0141] Furthermore, the first moving area information of the moving object in the first frame image can be determined based on the target pixel point. For example, the position coordinates of the target pixel point can be set to the first moving area information of the moving object in the first frame image (i.e., the center point position of the moving object).

[0142] In one embodiment, the target detection result includes the second moving area information of the target; in the above S20, "obtaining the target detection result in the video" may include:

[0143] S201, performing target detection based on a first frame image in the video using a second recognition model, and outputting second moving area information of the target in the first frame image.

[0144] In combination with the above description, the second recognition model in the motion detection model in this embodiment can specifically be a target detection model. The target detection model can obtain the target detection result in the first frame image based on the first frame image. The target detection model in this embodiment can refer to existing ones such as YOLO, SSD, Faster R-CNN, etc., which will not be repeated here.

[0145] On this basis, the target detection model can perform target detection based on the first frame image in the video and output the second moving area information of the target in the first frame image.

[0146] In one embodiment, in the above S30, “determining a dynamic target based on the moving object detection result and the target detection result” may include:

[0147] S301, determining distance information between a moving object and a target according to the first moving area information and the second moving area information;

[0148] S302: Determine a dynamic target in the first frame image according to the distance information.

[0149] In combination with the above description, in this embodiment, the moving object detection result may include the first moving area information of the moving object in the first frame image, and the target detection result specifically includes the second moving area information of the target. In combination with the above description, the target may be a preset type of target, such as pedestrians, vehicles, etc.

[0150] In a specific embodiment, the first moving area information includes the center point position of the moving object in the first frame image, and the second moving area information includes the detection box information. For example, in combination with the above description, the detection box information can be the center point position of the detection box bbox, and the center point position of the moving object can be the point corresponding to the local maximum value in the heat map corresponding to the first frame image. The subsequent embodiments will be described in detail and will not be repeated here. The embodiments will be described later by taking the example that the first moving area information includes the center point position of the moving object in the first frame image and the second moving area information includes the center point position of the detection box bbox.

[0151] In this way, the vehicle can determine the distance information between the moving object and the target according to the first moving area information and the second moving area information, and determine the dynamic target in the first frame image according to the distance information.

[0152] Specifically, for example, the vehicle can determine the dynamic target in the first frame image based on the distance information between the center point position of the moving object in the first frame image and the center point position of the detection box bbox.

[0153] In one embodiment, the moving object detection result further includes a target heat map corresponding to the first frame image;

[0154] In the above S302, “determining the dynamic target of the first frame image according to the distance information” may include:

[0155] S3021: If the distance information is less than a preset distance threshold, determine the dynamic target in the first frame image according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information.

[0156] In this embodiment, after calculating the distance information between the moving object and the detection box bbox, the vehicle can determine whether the distance information is less than a preset distance threshold.

[0157] If it is determined that the distance information is less than the preset distance threshold, it means that the detection box bbox belongs to a moving object, then the pixel points in the detection box in the target heat map (that is, the pixel points corresponding to the second moving area information in this embodiment) are obtained, and the dynamic target in the first frame image is determined based on the pixel points and the detection box.

[0158] If it is determined that the distance information is greater than or equal to the preset distance threshold, it means that the detection box bbox does not belong to a moving object. Then it can be output that the detection box bbox does not belong to a moving object, end this judgment, and recycle the next new detection box bbox.

[0159] In a specific embodiment, in the above S3021, “determining the dynamic target in the first frame image according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information” may include:

[0160] Determine a first number of pixels corresponding to the second moving area information in the target heat map, and a second number of non-zero pixels in the pixels;

[0161] If the ratio between the first number and the second number is smaller than a preset threshold, it is determined that the target corresponding to the second moving area information is a dynamic target.

[0162] If the ratio between the first number and the second number is greater than the preset threshold, it is determined that the target corresponding to the second moving area information is not a dynamic target.

[0163] In this embodiment, each detection box in the image can be replaced by bbox 1 ,bbox 2 ,,bbox N Represents (N is the number of detection boxes in the image).

[0164] Then the center point coordinates of each moving object are replaced by (cx 1 ,cy 1 ),(cx 2 ,cy 2 ),(cx M ,cy M ) represents (M is the number of moving objects in the image).

[0165] Determine in turn whether each detection box bbox belongs to a moving object, including:

[0166] (1) Find the center point of the nearest moving object corresponding to the center point of the bbox, recorded as (cx, cy);

[0167] (2) If the distance d between the center point (x, y) and the center point (cx, cy) satisfies d<δ (δ is the above-mentioned preset distance threshold, which is not specifically limited), proceed to the next step; otherwise, output that the bbox does not belong to a moving object, end this judgment, and recycle the next bbox;

[0168] (3) Calculate the number of non-zero pixel points (i.e., the second number) and the bbox area (i.e., the bbox area can represent the first number of pixel points in the bbox) in the area corresponding to the bbox on the heatmap (i.e., the target heatmap in this embodiment). If the second number / bbox area is greater than the preset threshold β, the bbox belongs to a moving object, that is, it can be determined that the target framed by the detection box bbox is a dynamic target; otherwise, the bbox does not belong to a moving object. End this judgment and recycle the next bbox.

[0169] In this way, this embodiment can determine which detection frames belong to moving objects, determine the dynamic targets in the first frame image, and then determine whether the sentry mode should sound an alarm by calculating the distance between these dynamic targets and the vehicle camera.

[0170] It can be seen that considering that the current moving object detection is mainly based on traditional image processing algorithms, the general outline of the moving object is obtained by calculating the pixel difference between the previous and next frames, or between the current frame and the background template, and then performing operations such as binarization, expansion and corrosion. This type of method is very sensitive to pixel changes caused by changes in lighting and weather, and relies on a large number of preset parameters, so robustness is difficult to guarantee.

[0171] In the embodiment of the present application, based on the heatmap, the probability of each pixel in the current frame belonging to a moving object is calculated, and then the pixel area of ​​the moving object is obtained. Compared with the traditional method, this embodiment does not need to preset many parameters, and has stronger generalization for different scene environments. In addition, since the heatmap algorithm generally generates a heatmap for each key point, it cannot be directly applied to the mobile target detection scene in this embodiment. Therefore, in this embodiment, all moving objects in each image can be mapped to the same heatmap to be integrated with the subsequent target detection results, while saving computing power.

[0172] In one embodiment, the dynamic target detection method in the present application may further include:

[0173] S40, performing moving object detection based on the sample video by using the first recognition model to be trained to obtain a feature map of the image in the sample video;

[0174] S50, obtaining a predicted heat map of the sample video based on a feature map of adjacent frame images using a first recognition model to be trained; and training the first recognition model to be trained according to the predicted heat map to obtain a first recognition model.

[0175] In combination with the above description, in this embodiment, a first recognition model to be trained is pre-constructed, and the first recognition model to be trained is based on a sample video and outputs a predicted heat map.

[0176] The sample video may be a driving video collected by the vehicle in real time, or a driving video recorded during the historical driving process.

[0177] Then, the vehicle can train the first recognition model to be trained according to the predicted heat map to obtain the first recognition model, that is, the moving object detection model in this embodiment.

[0178] In a specific embodiment, in the above S50, “training the first recognition model to be trained according to the predicted heat map to obtain the first recognition model” may include:

[0179] S501, obtaining the loss of a first recognition model according to the predicted heat map and a preset real heat map;

[0180] S502: Train the first recognition model based on the loss to obtain a trained first recognition model.

[0181] In this embodiment, two adjacent frames of the sample video can be simultaneously input into two convolutional neural networks with shared weights in the moving object detection model to extract features from the two images respectively to obtain corresponding feature images. Then, the predicted heat map can be obtained based on the two feature images. Here, the generation method of the target heat map can be referred to, and no further description is given here.

[0182] Furthermore, if Figure 3 , the true heat map true_heatmap corresponding to the predicted heat map and the real labeled data can be passed into wingLoss (i.e., the loss in this embodiment) to train the network until the model converges, wherein wingLoss is intended to improve the deep neural network's ability to handle small and medium errors during training.

[0183] Furthermore, based on the above loss, the first recognition model can be trained to obtain a trained first recognition model.

[0184] In one embodiment, the step of generating a real heat map in the present application may include:

[0185] S60, for a video frame in the video, obtaining a circumscribed rectangular frame of a moving object in the video frame, and obtaining a center point position of the circumscribed rectangular frame and a pixel point position of a pixel point in the video frame;

[0186] S70, determining the thermal value of the pixel point according to the pixel point position and the center point position, and obtaining a real thermal map of the video frame.

[0187] In this embodiment, for a video sequence to be marked, all moving objects that appear in each frame are found. For each moving object, the minimum circumscribed rectangular frame of the moving object needs to be marked, and the priority of the rectangular frame is determined. The priority definition includes: if the rectangular frame does not overlap with any other frame, the priority level is 1; if there is an overlap, the rectangular frame closest to the camera has a priority level of 1, and the next layer of rectangular frames has a priority level of 2, and so on.

[0188] Specifically, for example, each external rectangular box in the video frame is replaced by bbox 1 ,bbox 2 ,bbox N Indicates that the coordinates of the center point of each circumscribed rectangular box (x 1 ,y 1 ),(x 2 ,y 2 ),(x N ,y N )(N is the number of rectangular frames in the video frame).

[0189] On this basis, the vehicle can obtain the center point position of the circumscribed rectangular box and the pixel position of the pixel in the video frame, and obtain the thermal value of the video frame based on the pixel point position and the center point position to obtain a real thermal map.

[0190] In one embodiment, in the above S70, “determining the thermal value of the pixel point according to the pixel point position and the center point position to obtain a real thermal map of the video frame” may include:

[0191] S701, determining a target circumscribed rectangular frame corresponding to the pixel point;

[0192] S702, determining the thermal value of the pixel point according to the position of the pixel point and the center point position of the target circumscribed rectangular frame, and obtaining a real thermal map of the video frame.

[0193] In this embodiment, for each pixel point position (x, y) in the video frame, find all the rectangular boxes to which the point belongs (i.e., the target bounding rectangular boxes in this embodiment), and record the center point positions of these target bounding rectangular boxes as (x 1 ,y 1 ),(x 2 ,y 2 ),(x T ,y T ), T is the number of rectangular boxes where the point (x, y) is located.

[0194] Furthermore, the vehicle can obtain the thermal value of the video frame according to the pixel position and the center point position of the target circumscribed rectangular frame to obtain a real thermal map.

[0195] In a specific embodiment, in the above S701, “determining a target circumscribed rectangular frame corresponding to the pixel point” may include:

[0196] S7011, determining the circumscribed rectangular frame to which the pixel point belongs;

[0197] S7012, if there is one circumscribed rectangular frame, the circumscribed rectangular frame is the target circumscribed rectangular frame;

[0198] S7013: If there are multiple circumscribed rectangular frames, select a target circumscribed rectangular frame from the circumscribed rectangular frames according to the priorities of the circumscribed rectangular frames.

[0199] In this embodiment, combined with the above-mentioned embodiment, each external rectangular frame in the video frame is sequentially represented by bbox 1 ,bbox 2 ,bbox N Indicates that the coordinates of the center point of each circumscribed rectangular box (x 1 ,y 1 ),(x 2 ,y 2 ),(x N ,y N )(N is the number of rectangular frames in the video frame).

[0200] For each pixel point position (x, y) in the video frame, find all the rectangular boxes to which the point belongs (i.e., the target circumscribed rectangular boxes in this embodiment), and record the center point positions of these target circumscribed rectangular boxes as (x 1 ,y 1 ),(x 2 ,y 2 ),(x T ,y T ), T is the number of rectangular boxes where the point (x, y) is located.

[0201] Then, for each pixel position (x, y) in the video frame, calculate its thermal value:

[0202]

[0203] Where u = argmin i∈T S i , σ is the configured variance coefficient; Si represents the priority of rectangular box i.

[0204] The above formula indicates that only when a pixel point falls within a rectangular box, the heatmap value of the point is calculated by the distance from the point to the center point of the rectangular box to which it belongs; otherwise, the heatmap value corresponding to the point is directly 0; if a pixel point belongs to multiple rectangular boxes at the same time, the center point coordinates of the rectangular box with the highest priority are taken to calculate the heatmap value of the pixel point.

[0205] On this basis, if there is only one bounding rectangular frame, the bounding rectangular frame is directly set as the target bounding rectangular frame.

[0206] If there are multiple target bounding rectangular frames, the priority of each target bounding rectangular frame is obtained according to the depth information of the target bounding rectangular frame (such as the distance from the camera), and the target bounding rectangular frame with the highest priority is selected from the bounding rectangular frames according to the priority.

[0207] For example, the priority level of the rectangular frame closest to the camera is 1, the priority level of the next rectangular frame is 2, and so on.

[0208] In one embodiment, in the above S702, “determining the thermal value of the pixel point according to the pixel point position and the center point position of the target circumscribed rectangular frame to obtain a real thermal map of the video frame” may include:

[0209] S7021, obtaining distance information between the pixel point position and the center point position of the target circumscribed rectangular frame;

[0210] S7022: Determine the thermal value of the pixel point according to the distance information to obtain a real thermal map of the video frame.

[0211] In this embodiment, the real heat map can be determined based on the distance information between the pixel point position and the center point position of the target circumscribed rectangular frame.

[0212] Specifically, for example, according to:

[0213] Calculate the heat value of the video frame to get the real heat map.

[0214] Wherein, referring to the above embodiment, u=argmin i∈T S i , σ is the configured variance coefficient; S i Indicates the priority of rectangle i.

[0215] In another embodiment, heatmap(x, y) may also be truncated to obtain the final true heatmap true_heatmap(x, y):

[0216]

[0217] Where θ is the set threshold, which can be obtained by combining manual experience calibration.

[0218] Therefore, in general, the embodiment of the present application integrates mobile object detection and target detection through a two-stage vehicle sentinel mode implementation method and a new mobile object detection method, through a mobile object detection model and a target detection model, to improve the accuracy of the dynamic target detection algorithm, and then promptly issue an alarm when suspicious targets appear around the vehicle.

[0219] Accordingly, the embodiment of the present application further provides an alarm control method, which can be applied to a vehicle and may include:

[0220] Step A, determining whether to alarm based on the relative position between the dynamic target and the alarm subject.

[0221] It should be noted that, in this embodiment, the alarm subject may be a vehicle or an alarm device installed on the vehicle, and there is no specific limitation on this.

[0222] In this embodiment, in combination with the above embodiments, the vehicle can collect the current driving video after turning on the sentry mode, and then determine the moving object detection result in the first frame image based on at least two adjacent images in the video, and can determine the target detection result in the first frame image, and determine the dynamic target in the first frame image based on the moving object detection result and the target detection result, refer to the above embodiment, which will not be repeated here.

[0223] For example, when the alarm subject is a vehicle, after determining that there is a dynamic target, the vehicle can determine whether to alarm based on the relative position between the dynamic target and the vehicle.

[0224] For example, if the relative position between the detected dynamic target and the vehicle is less than a preset safety distance, the vehicle can be controlled to sound an alarm, otherwise a reminder message that a dynamic target has been detected can be output.

[0225] In general, this embodiment can accurately detect dynamic targets, avoid the interference of complex driving environment on dynamic target recognition, and issue an alarm in time when a dynamic target is detected, thereby ensuring the safety of vehicle driving.

[0226] Accordingly, the present application also provides a dynamic target detection device, such as Figure 5 As shown, the device may include:

[0227] A first acquisition module (1001) is used to acquire a moving object detection result in a video;

[0228] A second acquisition module (1002), used to acquire the target detection result in the video;

[0229] A determination module (1003) is used to determine a dynamic target based on the moving object detection result and the target detection result.

[0230] Optionally, the moving object detection result includes first moving area information of the moving object in the first frame image in the video;

[0231] The first acquisition module (1001) is further used for:

[0232] The first moving area information is obtained through a first recognition model based on a first frame image in a video and adjacent frame images of the first frame image.

[0233] Optionally, the first acquisition module (1001) is further configured to:

[0234] Extracting feature maps of the first frame image and the adjacent frame images respectively through the first recognition model;

[0235] Based on the feature map, obtaining a target thermal map of the first frame image;

[0236] According to the target thermal map, first moving area information of the moving object in the first frame image is determined.

[0237] Optionally, the first acquisition module (1001) is further configured to:

[0238] Through the convolutional neural network with shared weights in the first recognition model, feature extraction is performed on the first frame image and the adjacent frame images respectively to obtain feature maps corresponding to the first frame image and the adjacent frame images.

[0239] Optionally, the first acquisition module (1001) is further configured to:

[0240] The feature map of the first frame image is subtracted from the feature map of the adjacent frame image to obtain a target thermal map.

[0241] Optionally, the first acquisition module (1001) is further configured to:

[0242] Determine a target pixel whose pixel value is a local maximum value among the pixels according to the pixel value of the pixel in the target heat map and the neighborhood of the pixel;

[0243] According to the target pixel point, first moving area information of the moving object in the first frame image is determined.

[0244] Optionally, the target detection result includes second moving area information of the target;

[0245] The second acquisition module (1002) is further configured to:

[0246] Target detection is performed based on the first frame image in the video using a second recognition model, and second moving area information of the target in the first frame image is output.

[0247] Optionally, the determining module (1003) is further configured to:

[0248] Determining distance information between the moving object and the target according to the first moving area information and the second moving area information;

[0249] A dynamic target in the first frame of image is determined according to the distance information.

[0250] Optionally, the moving object detection result further includes a target heat map corresponding to the first frame image;

[0251] The determining module (1003) is further used for:

[0252] If the distance information is less than a preset distance threshold, the dynamic target in the first frame image is determined according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information.

[0253] Optionally, the determining module (1003) is further configured to:

[0254] Determine a first number of pixels corresponding to the second moving area information in the target heat map, and a second number of non-zero pixels in the pixels;

[0255] If the ratio between the first number and the second number is smaller than a preset threshold, it is determined that the target corresponding to the second moving area information is a dynamic target.

[0256] Optionally, the determining module (1003) is further configured to:

[0257] If the ratio between the first number and the second number is greater than or equal to the preset threshold, it is determined that the target corresponding to the second moving area information is not a dynamic target.

[0258] Optionally, the first moving area information includes a center point position of a moving object in the first frame image, and the second moving area information includes detection frame information.

[0259] Optionally, the dynamic target detection device in the present application further includes:

[0260] A region output module is used to detect moving objects based on the sample video by using the first recognition model to be trained to obtain a feature map of the image in the sample video; and to obtain a predicted heat map of the sample video based on the feature map of adjacent frame images by using the first recognition model to be trained;

[0261] A training module is used to train the first recognition model to be trained according to the predicted heat map and the predicted position information to obtain a first recognition model.

[0262] Optionally, the training module is also used to:

[0263] Obtaining the loss of the first recognition model according to the predicted heat map and the preset real heat map;

[0264] The first recognition model is trained based on the loss to obtain a trained first recognition model.

[0265] Optionally, the dynamic target detection device in the present application further includes:

[0266] A position information acquisition module, for acquiring, for a video frame in the video, a bounding rectangular frame of a moving object in the video frame, and acquiring a center point position of the bounding rectangular frame and a pixel point position of a pixel point in the video frame;

[0267] The real thermal map acquisition module is used to determine the thermal value of the pixel point according to the pixel point position and the center point position to obtain the real thermal map of the video frame.

[0268] Optionally, the real heat map acquisition module is also used to:

[0269] Determine the target circumscribed rectangular frame corresponding to the pixel point;

[0270] According to the position of the pixel point and the center point position of the target circumscribed rectangular frame, the thermal value of the pixel point is determined to obtain a real thermal map of the video frame.

[0271] Optionally, the real heat map acquisition module is also used to:

[0272] Determine the circumscribed rectangular frame to which the pixel point belongs;

[0273] If there is one circumscribed rectangular frame, the circumscribed rectangular frame is the target circumscribed rectangular frame;

[0274] If there are multiple circumscribed rectangular frames, a target circumscribed rectangular frame is selected from the circumscribed rectangular frames according to the priorities of the circumscribed rectangular frames.

[0275] Optionally, the real heat map acquisition module is also used to:

[0276] Obtaining distance information between the pixel point position and the center point position of the target circumscribed rectangular frame;

[0277] According to the distance information, the thermal value of the pixel point is determined to obtain a real thermal map of the video frame.

[0278] The present application also provides an alarm control device, such as Figure 6 As shown, the device can be arranged on a vehicle, and includes:

[0279] The alarm module (1004) is used to determine whether to alarm according to the relative position between the dynamic target and the vehicle determined in the dynamic target detection device or the alarm control method.

[0280] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0281] Accordingly, the present application also provides an electronic device, such as Figure 7 As shown, Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. It will be understood by those skilled in the art that the vehicle structure shown in the figure does not constitute a limitation on the vehicle, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0282] The processor 1101 is the control center of the electronic device 1100, and uses various interfaces and lines to connect various parts of the entire electronic device 1100. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, the processor 1101 executes various functions of the electronic device 1100 and processes data, thereby monitoring the electronic device 1100 as a whole. The processor 1101 can be a processor (Central Processing Unit, CPU), a graphics processing unit (graphics processing unit, GPU), a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application.

[0283] In the embodiment of the present application, the processor 1101 in the electronic device 1100 will load instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions, such as:

[0284] Get the moving object detection results in the video;

[0285] Obtaining target detection results in the video;

[0286] Based on the moving object detection result and the target detection result, a dynamic target is determined.

[0287] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0288] Optional, such as Figure 7 As shown, the electronic device 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107, respectively. Those skilled in the art can understand that Figure 7 The vehicle structure shown in the figure does not constitute a limitation on the vehicle, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0289] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the vehicle, which can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light-emitting diode (OLED, Organic Light-Emitting Diode), etc. The touch panel can be used to collect the user's touch operation on or near it (such as the user uses any suitable object or accessory such as a finger, stylus, etc. on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel may include two parts: a touch display system and a touch controller. Among them, the touch display system detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch display system, converts it into the touch point coordinates, and then sends it to the processor 1101, and can receive the command sent by the processor 1101 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event, and then the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.

[0290] The RF circuit 1104 may be used to send and receive RF signals to establish wireless communication with network devices or other vehicles through wireless communication, and to send and receive signals with network devices or other vehicles.

[0291] The audio circuit 1105 can be used to provide an audio interface between the user and the vehicle through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and converted into audio data, and then the audio data is output to the processor 1101 for processing, and then sent to another vehicle through the radio frequency circuit 1104, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earplug jack to provide communication between an external headset and the vehicle.

[0292] The input unit 1106 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0293] The power supply 1107 is used to supply power to various components of the electronic device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management device, so that the power management device can manage charging, discharging, power consumption and other functions. The power supply 1107 can also include one or more DC or AC power supplies, recharging devices, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.

[0294] although Figure 7 Not shown, the electronic device 1100 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.

[0295] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0296] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0297] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of computer programs are stored. The computer program can be loaded by a processor to execute any dynamic target detection method provided in the embodiment of the present application. The computer program can execute the following steps of the dynamic target detection method:

[0298] Get the moving object detection results in the video;

[0299] Obtaining target detection results in the video;

[0300] Based on the moving object detection result and the target detection result, a dynamic target is determined.

[0301] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0302] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0303] Since the computer program stored in the computer-readable storage medium can execute any one of the dynamic target detection methods and alarm control methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the dynamic target detection methods provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0304] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0305] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable target detection device to generate a machine, so that the instructions executed by the processor of the computer or other programmable target detection device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0306] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable target detection device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0307] These computer program instructions may also be loaded onto a computer or other programmable target detection device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0308] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0309] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0310] Computer readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated communication signals and carrier waves.

[0311] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0312] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0313] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0314] The above are only preferred embodiments of the present application and do not constitute any form of limitation to the present application. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A dynamic target detection method, characterized in that: The method comprises: Get the moving object detection results in the video; Obtaining target detection results in the video; Based on the moving object detection result and the target detection result, a dynamic target is determined.

2. The dynamic target detection method according to claim 1, characterized in that: The moving object detection result includes first moving area information of the moving object in the first frame image in the video; The step of obtaining a moving object detection result in a video includes: The first moving area information is obtained through a first recognition model based on a first frame image in a video and adjacent frame images of the first frame image.

3. The dynamic target detection method according to claim 2, characterized in that: The obtaining the first moving area information based on a first frame image and adjacent frame images of the first frame image in a video by using a first recognition model includes: Extracting feature maps of the first frame image and the adjacent frame images respectively through the first recognition model; Based on the feature map, obtaining a target thermal map of the first frame image; According to the target thermal map, first moving area information of the moving object in the first frame image is determined.

4. The dynamic target detection method according to claim 3, characterized in that: The extracting feature maps of the first frame image and the adjacent frame images respectively through the first recognition model includes: Through the convolutional neural network with shared weights in the first recognition model, feature extraction is performed on the first frame image and the adjacent frame images respectively to obtain feature maps corresponding to the first frame image and the adjacent frame images.

5. The dynamic target detection method according to claim 3, characterized in that: The step of obtaining a target thermal map of the first frame image based on the feature map includes: The feature map of the first frame image is subtracted from the feature map of the adjacent frame image to obtain a target thermal map.

6. The dynamic target detection method according to claim 3, characterized in that: The determining, according to the target heat map, first moving area information of the moving object in the first frame image includes: Determine a target pixel whose pixel value is a local maximum value among the pixels according to the pixel value of the pixel in the target heat map and the neighborhood of the pixel; According to the target pixel point, first moving area information of the moving object in the first frame image is determined.

7. The dynamic target detection method according to claim 2, characterized in that: The target detection result includes second moving area information of the target; The obtaining of the target detection result in the video includes: Target detection is performed based on the first frame image in the video using a second recognition model, and second moving area information of the target in the first frame image is output.

8. The dynamic target detection method according to claim 7, characterized in that: The determining of the dynamic target based on the moving object detection result and the target detection result includes: Determining distance information between the moving object and the target according to the first moving area information and the second moving area information; A dynamic target in the first frame of image is determined according to the distance information.

9. The dynamic target detection method according to claim 8, characterized in that: The moving object detection result further includes a target heat map corresponding to the first frame image; The step of determining the dynamic target in the first frame image according to the distance information includes: If the distance information is less than a preset distance threshold, the dynamic target in the first frame image is determined according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information.

10. The dynamic target detection method according to claim 9, characterized in that: The determining of the dynamic target in the first frame image according to the pixel points corresponding to the second moving area information in the target heat map and the second moving area information comprises: Determine a first number of pixels corresponding to the second moving area information in the target heat map, and a second number of non-zero pixels in the pixels; If the ratio between the first number and the second number is smaller than a preset threshold, it is determined that the target corresponding to the second moving area information is a dynamic target.

11. The dynamic target detection method according to claim 10, characterized in that: The method further comprises: If the ratio between the first number and the second number is greater than or equal to the preset threshold, it is determined that the target corresponding to the second moving area information is not a dynamic target.

12. The dynamic target detection method according to claim 7, characterized in that: The first moving area information includes the center point position of the moving object in the first frame image, and the second moving area information includes detection frame information.

13. The dynamic target detection method according to any one of claims 1 to 12, characterized in that: The method further comprises: Performing moving object detection based on the sample video using the first recognition model to be trained, obtaining a feature map of the image in the sample video; Obtaining a predicted heat map of the sample video based on a feature map of adjacent frame images using a first recognition model to be trained; The first recognition model to be trained is trained according to the predicted heat map to obtain a first recognition model.

14. The dynamic target detection method according to claim 13, characterized in that: The step of training the first recognition model to be trained according to the predicted heat map to obtain the first recognition model includes: Obtaining the loss of the first recognition model according to the predicted heat map and the preset real heat map; The first recognition model is trained based on the loss to obtain a trained first recognition model.

15. The dynamic target detection method according to claim 14, characterized in that: The steps of generating the real heat map include: For a video frame in the video, obtain a circumscribed rectangular frame of a moving object in the video frame, and obtain a center point position of the circumscribed rectangular frame and a pixel point position in the video frame; The thermal value of the pixel point is determined according to the pixel point position and the center point position to obtain a real thermal map of the video frame.

16. The dynamic target detection method according to claim 15, characterized in that: The step of determining the thermal value of the pixel point according to the pixel point position and the center point position to obtain a real thermal map of the video frame includes: Determine the target circumscribed rectangular frame corresponding to the pixel point; According to the position of the pixel point and the center point position of the target circumscribed rectangular frame, the thermal value of the pixel point is determined to obtain a real thermal map of the video frame.

17. The dynamic target detection method according to claim 16, characterized in that: The determining of the target circumscribed rectangular frame corresponding to the pixel point includes: Determine the circumscribed rectangular frame to which the pixel point belongs; If there is one circumscribed rectangular frame, the circumscribed rectangular frame is the target circumscribed rectangular frame; If there are multiple circumscribed rectangular frames, a target circumscribed rectangular frame is selected from the circumscribed rectangular frames according to the priorities of the circumscribed rectangular frames.

18. The dynamic target detection method according to claim 16, characterized in that: The step of determining the thermal value of the pixel point according to the pixel point position and the center point position of the target circumscribed rectangular frame to obtain a real thermal map of the video frame includes: Obtaining distance information between the pixel point position and the center point position of the target circumscribed rectangular frame; The thermal value of the pixel is determined according to the distance information to obtain a real thermal map of the video frame.

19. An alarm control method, characterized in that: The method comprises: Whether to issue an alarm is determined based on the relative position between the dynamic target and the alarm subject determined according to any one of claims 1 to 18.

20. A dynamic target detection device, characterized in that: include: A first acquisition module (1001) is used to acquire a moving object detection result in a video; A second acquisition module (1002), used to acquire the target detection result in the video; A determination module (1003) is used to determine a dynamic target based on the moving object detection result and the target detection result.

21. An alarm control device, characterized in that: The device comprises: An alarm module (1004) is used to determine whether to alarm according to the relative position between the dynamic target and the vehicle determined in claim 20 or the alarm control method of claim 19.

22. An electronic device (1100), characterized in that: It comprises a processor (1101) and a memory (1102), wherein the memory (1102) stores a computer program, and when the computer program is executed by the processor (1101), the processor (1101) executes any method described in claim 1 to claim 19.

23. A vehicle, characterized in that: The vehicle is provided with at least one of the dynamic target detection device according to claim 20, the alarm control device according to claim 21 and the electronic device (1100) according to claim 22.

24. A computer-readable storage medium, characterized in that: It includes a computer program. When the computer program is run on an electronic device (1100), the computer program is used to enable the electronic device (1100) to execute any method described in claim 1 to claim 19.