Training Method, Device, Processing Device and Storage Medium of Object Detection Network

By introducing feature processing subnets and image detection subnets into the object detection network, combining scene motion reference values ​​and reference impact values, and adjusting network parameters, the problem of low accuracy in motion object detection in the prior art is solved, and more efficient and accurate object detection is achieved.

CN114596519BActive Publication Date: 2025-06-24ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210086588.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-06-24
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

In the prior art, the accuracy of detecting moving objects is low, especially when the background difference is large, making it difficult to accurately identify the target object.

Method used

A training method for an object detection network is provided, including a feature processing subnet and an image detection subnet. By fusing feature information, determining scene motion reference value and reference impact value, and adjusting network parameters to improve the accuracy of object detection.

Benefits of technology

The object detection network generated by this method can detect moving objects more accurately, improving detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596519B_ABST
    Figure CN114596519B_ABST
Patent Text Reader

Abstract

The present application relates to a training method, device, processing equipment, and storage medium for a target detection network. The target detection network at least includes a feature processing sub-network and an image detection sub-network. The method includes: using the feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; determining at least one reference influence value based on the fusion feature information, where the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; using the feature processing sub-network and the image detection sub-network to respectively perform target detection sub-operations on each video frame in the second video frame sequence to obtain the prediction loss information corresponding to each video frame in the second video frame sequence, and adjusting the network parameters of at least one network in the feature processing sub-network and the image detection sub-network based on the prediction loss information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular, to a training method, device, processing device, and storage medium for a target detection network. Background Art

[0002] The detection and recognition technology of targets is one of the important branches of machine vision, which refers to extracting targets from the background and analyzing them. Among them, the detection of moving targets is a branch of target detection. The task of moving target detection is to detect moving targets in a given video, which serves as a preprocessing part of intelligent video analysis and lays a foundation for subsequent target recognition, target tracking, and action recognition in the video.

[0003] In related technologies, a neural network model is usually used to detect moving targets. The neural network model generally uses two adjacent frames as samples to extract the foreground of the target. However, this method may not be able to accurately identify the target object from subsequent video frames due to the large background differences between the video frames of the video to be detected, thereby reducing the detection accuracy of moving targets.

[0004] Therefore, how to improve the accuracy of detecting moving targets is an urgent problem to be considered. Summary of the Invention

[0005] In this embodiment, a training method, device, processing device, and storage medium for a target detection network are provided to solve the problem of low accuracy in detecting moving targets in related technologies.

[0006] In a first aspect, a training method for a target detection network is provided in this embodiment. The target detection network includes at least a feature processing sub-network and an image detection sub-network. The method includes:

[0007] Using the feature processing sub-network, determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the video to be detected collected for a target scene;

[0008] Obtain a scene motion reference value based on the fusion feature information; the scene motion reference value represents the motion information of the target included in the first video frame sequence; determine at least one reference influence value based on the scene motion reference value; the reference influence value is used to represent the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a continuous plurality of video frames in the video to be detected, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frame;

[0009] Using the feature processing sub-network and the image detection sub-network, perform target detection sub-operations on each video frame in the second video frame sequence to obtain prediction loss information corresponding to each video frame in the second video frame sequence, and adjust network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the target detection sub-operation for the i-th video frame includes:

[0010] Using the feature processing sub-network, based on at least one determined reference influence value, process the historical memory information of the i-th video frame and the candidate memory information of the i-th video frame to obtain the memory information of the i-th video frame; the candidate memory information is determined based on video frames before the i-th video frame in the video to be detected; and

[0011] Using the image detection sub-network, based on the memory information of the i-th video frame, determine the target detection result of the i-th video frame, and determine the prediction loss information corresponding to the i-th video frame based on the target detection result and the target annotation result of the i-th video frame.

[0012] In the technical solution provided by the embodiments of the present application, the fusion feature information of the first video frame sequence can be determined through the feature processing sub-network. Then, a scene motion reference value can be obtained based on the fusion feature information, and at least one reference influence value can be determined based on the scene motion reference value. Secondly, the prediction loss information corresponding to each video frame in the second video frame sequence can be obtained by using the feature processing sub-network and the image detection sub-network. Finally, network parameters of at least one of the feature processing sub-network and the image detection sub-network can be adjusted based on the prediction loss information. Thus, a more accurate target detection network can be generated, making the performance of the target detection network after parameter adjustment better, thereby improving the accuracy of moving target detection.

[0013] In a second aspect, an embodiment of the present application further provides a target detection method, including:

[0014] Obtain a video to be detected;

[0015] Input the video to be detected into a target detection network, and use the target detection network to perform feature processing and image detection to obtain target detection results of each video frame in the video to be detected, where:

[0016] The target detection network is obtained by the training method of the target detection network described above.

[0017] The object detection method provided by the embodiment of the present application can apply the object detection network trained by the above method to the object detection method, thereby improving the efficiency and accuracy of the object detection method for detecting moving objects.

[0018] In a third aspect, a training device for an object detection network is provided in this embodiment. The device includes:

[0019] A feature information acquisition module, configured to use the feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the to-be-detected video collected for the target scene;

[0020] An influence value determination module, configured to obtain a scene motion reference value based on the fusion feature information; the scene motion reference value characterizes the motion information of the target included in the first video frame sequence; determine at least one reference influence value based on the scene motion reference value; the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a continuous plurality of video frames in the to-be-detected video, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frame;

[0021] A network parameter adjustment module, configured to use the feature processing sub-network and the image detection sub-network to perform object detection sub-operations on each video frame in the second video frame sequence respectively, obtain the prediction loss information corresponding to each video frame in the second video frame sequence, and adjust the network parameters of at least one network in the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the object detection sub-operation of the i-th video frame includes:

[0022] Using the feature processing sub-network, based on the determined at least one reference influence value, process the historical memory information of the i-th video frame and process the candidate memory information of the i-th video frame to obtain the memory information of the i-th video frame; the candidate memory information is determined based on the video frames before the i-th video frame in the to-be-detected video; and

[0023] Using the image detection sub-network, based on the memory information of the i-th video frame, determine the object detection result of the i-th video frame, and determine the prediction loss information corresponding to the i-th video frame based on the object detection result and the object annotation result of the i-th video frame.

[0024] Fourth aspect, in this embodiment, a processing device is provided, including a memory and a processor. The memory stores computer program instructions, and when the processor executes the computer program instructions, the steps of the above-mentioned method are implemented.

[0025] Fifth aspect, in this embodiment, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the steps of the above-mentioned method are implemented.

[0026] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0028] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0029] Figure 2 is a schematic flowchart of a method for training a target detection network provided by an embodiment of the present application;

[0030] Figure 3 is a schematic diagram of the training process of a target detection network provided by an embodiment of the present application;

[0031] Figure 4 is a schematic diagram of the working principle of a GRU network provided by an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of the input and output of a target detection network provided by an embodiment of the present application;

[0033] Figure 6 is a schematic diagram of the module structure of a training device for a target detection network provided by an embodiment of the present application;

[0034] Figure 7 is a schematic diagram of the module structure of a processing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To understand the purpose, technical solution, and advantages of the present application more clearly, the present application is described and illustrated below with reference to the drawings and embodiments.

[0036] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings understood by those of ordinary skill in the technical field to which this application pertains. In this application, words such as "a", "an", "one", "the", "these", etc. do not indicate a limitation in quantity and can be singular or plural. The terms "include", "comprise", "have" and any variants thereof used in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connect", "be connected", "couple" and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" used in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. used in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0037] In addition, for a better illustration of this application, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that this application can still be implemented without certain specific details. In some instances, devices, means, elements and circuits well-known to those skilled in the art are not described in detail to highlight the gist of this application.

[0038] In the actual process of moving target detection, it is often necessary to perform moving target detection on the image sequence collected at high altitude. High-altitude collection usually uses a high-altitude imaging device to collect images. The high-altitude imaging device may include, for example, an unmanned aerial vehicle, a high-altitude surveillance camera, etc. Since the high-altitude imaging device is far from the scene, the target objects in the image are small and the background is complex. Therefore, in the process of moving target detection, the target objects in the image cannot be accurately detected. Based on this, the target detection network trained by using the training method of the target detection network provided in the embodiments of this application can be used to perform target detection on the images collected by the high-altitude imaging device, thereby improving the accuracy of the moving target detection result.

[0039] Figure 1 It is a schematic diagram of the application scenario provided by the embodiments of this application. Figure 1The figure shows the scenario of applying the object detection network provided by the embodiments of the present application to the detection of moving objects. Specifically, the acquisition device 101 can be used to collect a video sequence for the area to be detected. The acquisition device 101 can be an electronic device with data acquisition and data transceiver capabilities. For example, the acquisition device 101 can include an electronic device capable of collecting image information and / or video information of the area to be detected, such as a camera, a lidar, etc. Among them, the camera can include a monocular camera, an infrared camera, a multiocular camera, a depth camera, etc. The lidar can include a single-line lidar, a multi-line lidar, etc., and the present application does not limit this. The acquisition device 101 can send the collected video sequence to the object detection device 103, and the object detection device 103 segments the foreground features from each video frame of the video sequence. Then, the object detection device 103 can determine the position and category of the object in each video frame according to the foreground features;

[0040] As an embodiment, the object detection device 103 and the above-mentioned acquisition device 101 can be the same device / equipment, or different devices / equipment; for example, the acquisition device can be an ordinary camera, an intelligent camera, etc., and the object detection device 103 can be an intelligent camera, a server, etc.

[0041] The following will describe in detail the training method of the object detection network according to the present application with reference to the accompanying drawings. Figure 2 It is a schematic flowchart of a method for an embodiment of the training method of the object detection network provided by the present application. Although the present application provides method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present application. When the method is actually used in the training process of the object detection network or when the device executes, it can be executed in the method order shown in the embodiments or drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, an embodiment of the training method of the object detection network provided by the present application is as follows Figure 2 shown, the object detection network at least includes a feature processing sub-network and an image detection sub-network, and the method may include:

[0042] Step 201: Using the feature processing sub-network, determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the to-be-detected video collected for the target scene.

[0043] Step 203: Obtain a scene motion reference value based on the fused feature information; the scene motion reference value characterizes the motion information of the target included in the first video frame sequence; determine at least one reference influence value based on the scene motion reference value; the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a plurality of consecutive video frames in the video to be detected, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frames.

[0044] Step 205: Use the feature processing sub-network and the image detection sub-network to respectively perform target detection sub-operations on each video frame in the second video frame sequence, obtain the prediction loss information corresponding to each video frame in the second video frame sequence, and adjust the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the target detection sub-operation of the i-th video frame includes:

[0045] Use the feature processing sub-network to process the historical memory information of the i-th video frame and the candidate memory information of the i-th video frame based on the determined at least one reference influence value to obtain the memory information of the i-th video frame; the candidate memory information is determined based on the video frames before the i-th video frame in the video to be detected; and

[0046] Use the image detection sub-network to determine the target detection result of the i-th video frame based on the memory information of the i-th video frame, and determine the prediction loss information corresponding to the i-th video frame based on the target detection result and the target annotation result of the i-th video frame.

[0047] In the embodiments of the present application, the video to be detected may include a plurality of consecutive video frames collected for a certain target scene. The plurality of consecutive video frames may include consecutive multiple frames of images collected at a preset frame rate. For example, consecutive video frames collected for a certain road section, park, or other scenes. The preset frame rate may be set to 10fps, 8fps, 1fps, etc., without limitation here. Each video frame may include at least one target object, and there are various target types for the target object. For example, it may include pedestrians, vehicles, animals, buildings, etc. The first video frame sequence and the second video frame sequence may be consecutive multiple video frames in the video to be detected. It should be noted that the collection time of the video frames in the first video frame sequence is earlier than the collection time of the video frames in the second video frame, and the first video frame sequence and the second video frame may be two adjacent video segments in the same video, or the first video frame sequence and the second video frame sequence may also be two non-adjacent video segments in the same video. In addition, the number of video frames in the first video frame sequence and the number of video frames in the second video frame sequence are not limited. For example, the number of video frames in the first video frame sequence may be more than the number of video frames in the second video frame sequence, the number of video frames in the first video frame sequence may be equal to the number of video frames in the second video frame sequence, or the number of video frames in the first video frame sequence may also be less than the number of video frames in the second video frame sequence.

[0048] In one example, the first video frame sequence includes video frames with a collection time between 8:30:00 - 8:30:03, and the second video frame sequence includes video frames with a collection time between 8:30:04 - 8:30:05; or, if the same video includes a total of 50 video frames from video frame 1 to video frame 50 in the order of collection time from early to late, then it can be but not limited to considering video frames 1 to 30 as the video frames in the first video frame sequence, and considering video frames 31 to 50 as the video frames in the second video frame sequence; it can also consider video frames 1 to 20 as the video frames in the first video frame sequence, and consider video frames 21 to 50 as the video frames in the second video frame sequence.

[0049] In the embodiments of the present application, before inputting the video to be detected into the target detection network 300 for training, in order to eliminate irrelevant information in the video to be detected, preprocessing can be performed on each video frame in the video to be detected. The preprocessing may include, for example but not limited to, one or more operations such as smoothing processing, denoising processing, and normalization processing, etc. Specifically, in one example, normalization processing can be performed on the R, G, and B channels of each video frame in the video to be detected, and uniformly scaled to a size of 96*640.

[0050] In the embodiments of the present application, as Figure 3As shown, the target detection network 300 may include a feature processing sub-network 301 and an image detection sub-network 303. Among them, the feature processing sub-network 301 may determine the fusion feature information of the first video frame sequence from the feature information of each video frame in the first video frame sequence. The feature processing sub-network 301 may include a (Fully Convolutional Network, FCN), a recurrent neural network (Recurrent Neural Networks, RNN) network, a long short-term memory (Long Short Term Memory, LSTM) network, a gated recurrent unit (Gated Recurrent Unit, GRU) network, a bidirectional RNN network, etc., which are not limited in this application. The image detection sub-network 303 may include a detection network layer and a localization network layer.

[0051] In the embodiment of the present application, each video frame in the first video frame may be input into the feature extraction sub-network in sequence first, and the feature information of each video frame in the first video frame sequence may be output to the feature processing sub-network 301 through the feature extraction sub-network. The feature processing sub-network 301 may process the feature information of each video frame to determine the fusion feature information of the first video frame sequence. The fusion feature information may include the accumulation of the feature information of each video frame in the first video frame sequence. In an embodiment of the present application, after determining the fusion feature information, a scene motion reference value may be determined based on the fusion feature information, where the scene motion reference value is used to characterize the motion information of the target included in the first video frame sequence.

[0052] In the embodiments of the present application, in actual applications, since the position of the target object in the video to be detected may not be stationary in each video frame and may move. This will cause the background of each video frame in the video to be detected to change continuously, and the background difference is relatively large. The scene motion reference value can be used to characterize the motion information of the target and can also be used to reflect the background change speed of each video frame in the first video frame sequence. For example, when the target object is a movable object such as a person, an animal, or a vehicle, its position in each video frame will change. Therefore, the motion information of the target object can be determined based on the change, and then the scene motion reference value can be determined. The motion information reflects the motion condition of the target object, that is, the position change condition of the target object in different video frames. Specifically, in an embodiment of the present application, when the number of moving target objects is one, the scene motion reference value can be determined based on the motion speed of the moving target object. When the number of moving target objects is multiple, it can be determined based on the average value of the speeds of multiple moving target objects, or the maximum value of the motion speeds of multiple moving target objects can be selected. The present application does not limit this here.

[0053] In an embodiment of the present application, since before the target detection network 300 is trained, each video frame in the first video frame sequence is labeled with a target object using a bounding box, the motion information of the target object in the first video frame sequence can be determined based on the position change of the bounding box of the target object in the first video frame sequence. In an embodiment of the present application, the position change of the bounding box of the target object in the first video frame sequence can be determined based on the intersection over union (IOU) of the corresponding bounding boxes in different video frames. The intersection over union can be used to represent the overlapping rate of two bounding boxes. Specifically, in an example, the bounding box of the first video frame and the corresponding bounding box of the last frame in the first video frame sequence can be selected. When it is determined that the intersection over union of the two bounding boxes is greater than the preset threshold, the target object can be determined as a moving target object. Otherwise, the target object can be determined as a stationary target object. It should be noted that in an embodiment of the present application, in order to improve the training speed of the target detection network 300, the bounding boxes of the stationary target objects can be removed and not used as training samples.

[0054] In an embodiment of the present application, the scene motion reference value may be a specific speed value, such as 30 km / h, 80 km / h, etc., or a speed level, such as fast, medium, slow, etc. In an embodiment of the present application, the speed level may be determined based on a preset reference speed threshold. When the scene motion reference value is greater than the reference speed threshold, the speed level may be determined as fast. Specifically, in one example, according to the preset reference speed threshold, the scene motion reference value of the first video frame sequence collected for the high-altitude parabolic scene may be determined as fast, and the scene motion reference value of the first video frame sequence collected for the panoramic viaduct scene may be determined as slow.

[0055] In an embodiment of the present application, after determining the scene motion reference value, at least one reference influence value may be determined. Since the scene motion reference values are different, the background feature information of the video frame sequence to be accumulated is also different. For example, in the case of a large scene motion reference value such as the high-altitude parabolic scene, since the target object moves too fast, the background difference between the front and back video frames in the collected video frame sequence is large, and the background changes quickly. Therefore, it may not be necessary to accumulate a large amount of background feature information in the target detection network 300. Thus, the reference influence value of the first video frame sequence may be selected and adjusted, so as to adjust the proportion of the information volume of the first video frame sequence flowing into the second video frame sequence.

[0056] As an embodiment, one reference influence value may be determined in the present application, or multiple reference influence values may be determined, that is, the number of the reference influence values may be one, two or more than two. In one example, the reference influence value may include one or both of a first influence value and a second influence value. The first influence value may be used to characterize the importance degree of the feature information of the first video frame sequence, and the second influence value may be used to characterize the importance degree of the feature information of the second video frame sequence.

[0057] As an embodiment, the sum value of the first influence value and the second influence value may be 1, or other preset values. When the sum value is 1, the ratio of the first influence value and the second influence value can be directly determined, so that the weight ratio of the feature information of the first video frame sequence and / or the second video frame sequence can be directly determined without further calculation. For example, the first influence value may be 0.75 and the second influence value may be 0.25. It should be noted that when there is only one reference influence value, the second influence value may be set as a constant.

[0058] Through the above embodiments, the at least one reference influence value can be determined based on the scene motion reference value, so that the trained target detection network has strong generalization ability, can output accurate target detection results for each video frame, and significantly improves the accuracy of target detection.

[0059] In the embodiments of the present application, after determining the at least one reference influence value, the feature processing sub-network 301 and the image detection sub-network 303 can be used to perform target detection sub-operations on each video frame in the second video frame sequence respectively, and determine the prediction loss information corresponding to each video frame in the second video frame sequence. Specifically, in one embodiment of the present application, the target detection sub-operation of the i-th video frame may include two steps. The first operation may include obtaining the memory information of the i-th video frame by using the feature processing sub-network 301. The second operation may include determining the target detection result of the i-th video frame by using the image detection sub-network 303 and determining the corresponding prediction loss information of the i-th video frame.

[0060] The following non-limitingly takes the GRU network as an example to illustrate the first operation performed by the feature processing sub-network 301 on the i-th video frame. Figure 4 It is a schematic diagram of the module structure of the GRU network at time t. The GRU model may include two gating units, a reset gating unit and an update gating unit. The GRU model may have two inputs. One is the feature of the input i-th video frame The other is the amount of information retained by the (i - 1)-th video frame The GRU model may also have two outputs. One is the output information yt of the input i-th video frame, and the other is the amount of information transmitted to the (i + 1)-th video frame

[0061]

[0062]

[0063]

[0064]

[0065] Among them, the r t may be the parameter of the reset gating unit, and the z t may be the parameter of the update gating unit. The is the candidate memory information, which can be used to represent the amount of information contained in the input i-th video frame. The is the memory information of the i-th video frame. Among them, the can be used to represent the amount of information retained by forgetting the (i - 1)-th video frame, and the is used to represent the amount of information contained in the i-th video frame. In an embodiment of the present application, the parameter r of the reset gating unit t can be used to adjust the contained in the proportion of the amount of information. The parameter z of the update gating unit t can be used to adjust the degree of forgetting and the degree of memory.

[0066] In an embodiment of the present application, based on the determined at least one reference influence value, the historical memory information of the i-th video frame and the candidate memory information of the i-th video frame can be processed to obtain the memory information of the i-th video frame. The processing may include performing weighted fusion processing on the historical memory information and the candidate memory information to obtain the memory information of the i-th video frame. In an embodiment of the present application, one of the first influence value and the second influence value can be set as a constant, and the constant can be 1 or other constant values. Then, the other influence value of the first influence value and the second influence value can be determined according to the scene motion reference value. Specifically, the second influence value can be set to the value 1, that is, the amount of information of the candidate memory information is not adjusted. Then, the first influence value can be determined according to the scene motion reference value. After determining the first influence value, the weight value of the historical memory information can be determined. For example, in one example, after determining the first influence value as m, the memory information of the i-th video frame can be In another embodiment of the present application, the first influence value and the second influence value can be determined respectively according to the scene motion reference value, and the first influence value and the second influence value can be independent of each other. Then, weighted fusion can be performed on the historical memory information and the candidate memory information according to the first influence value and the second influence value. For example, in one example, after determining the first influence value as b and the second influence value as q, the memory information of the i-th video frame can be

[0067] Of course, in some other embodiments of the present application, the sum value of the first influence value and the second influence value can also be set as a set value, and the set value can be 1 or other constant values. Therefore, in an embodiment of the present application, only by determining the first influence value, the second influence value can be determined according to the relationship that the sum value of the first influence value and the second influence value is the set value. Or, only by determining the second influence value, the first influence value can be determined according to the relationship that the sum value of the first influence value and the second influence value is the set value. For example, in an example, after determining that the first influence value is a, where 0 < a < 1, the second influence value can be determined to be 1 - a according to the relationship that the sum value of the first influence value and the second influence value is 1. The memory information of the i-th video frame can be

[0068] In the embodiments of the present application, as Figure 3 shown, after determining the memory information of the i-th video frame, the memory information can be sent to the image detection sub-network 303, and the image detection sub-network 303 determines the target detection result of the i-th video frame. Since the target object is labeled with a target box in the second video frame sequence, the corresponding prediction loss information of the i-th video frame can be determined based on the target detection result and the target annotation result of the i-th video frame. Specifically, in an embodiment of the present application, the prediction loss information can be determined based on a loss function, and the loss function can include any one or more of a cross-entropy loss function, a Smooth L1 loss function, and the like.

[0069] In the embodiments of the present application, after determining the prediction loss information, the network parameters of at least one of the feature processing sub-network 301 and the image detection sub-network 303 can be adjusted based on the prediction loss information. In an embodiment of the present application, the parameters of the feature processing sub-network 301 can be adjusted according to the prediction loss information. In an example, the parameters of the feature processing sub-network 301 can include the parameters of the reset gate unit, the parameters of the update gate unit, and the like. In another embodiment of the present application, the parameters of the image detection sub-network 303 can be adjusted according to the prediction loss information. Of course, in other embodiments of the present application, the parameters of the feature processing sub-network 301 and the parameters of the image detection sub-network 303 can also be adjusted simultaneously, and the present application does not limit this here.

[0070] In summary, in the technical solution provided by the embodiment of the present application, the fusion feature information of the first video frame sequence can be determined by the feature processing sub-network 301. Then, a scene motion reference value can be obtained based on the fusion feature information, and at least one reference influence value can be determined based on the scene motion reference value. Among them, the scene motion reference value can be used to characterize the motion information of the target included in the first video frame sequence. Secondly, the prediction loss information corresponding to each video frame in the second video frame sequence can be obtained by using the feature processing sub-network 301 and the image detection sub-network 303. Finally, the network parameters of at least one of the feature processing sub-network and the image detection sub-network can be adjusted based on the prediction loss information. Thus, a more accurate target detection network can be generated, so that the performance of the target detection network after parameter adjustment is better, and finally the accuracy of the target object determined from the target video frame is higher.

[0071] In practical applications, the prediction loss information corresponding to each video frame in the second video frame sequence can be determined respectively to determine the comprehensive prediction loss information of the second video frame sequence. Specifically, in an embodiment of the present application, the adjusting the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information includes:

[0072] Step 401: Obtain comprehensive prediction loss information based on the prediction loss information corresponding to each video frame in the second video frame sequence;

[0073] Step 403: Adjust the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the comprehensive prediction loss information.

[0074] In the embodiment of the present application, the sub-prediction loss information corresponding to each video frame in the second video frame sequence can be determined respectively. Then, the comprehensive prediction loss information can be determined by processing the sub-prediction loss information. In an embodiment of the present application, the sum value of the sub-prediction loss information can be statistically calculated, and the sum value can be used as the comprehensive prediction loss information. In another embodiment of the present application, the average value of the sub-prediction loss information can be obtained, and the comprehensive prediction loss information can be determined based on the average value and the number of frames in the second video frame sequence. Of course, in other embodiments of the present application, the sub-prediction loss information can also be weighted and summed to determine the comprehensive prediction loss information, which is not limited in this application. When determining the comprehensive prediction loss information, the network parameters of at least one of the feature processing sub-network 301 and the image detection sub-network 303 can be adjusted accordingly.

[0075] Further, in an embodiment of the present application, the reference influence value includes a first influence value set for the first video frame sequence;

[0076] Determining the at least one reference influence value based on the scene motion reference value includes:

[0077] Step 501: Determine a first preset reference value interval corresponding to the scene motion reference value;

[0078] Step 503: Determine the candidate reference influence value corresponding to the first preset reference value interval as the first influence value.

[0079] In an embodiment of the present application, the reference influence value may include a first influence value set for the first video frame sequence. The association relationship between the scene motion reference value and the first influence value may be a segmented corresponding relationship. For example, the scene motion reference value of the first video frame sequence may be divided into two or more preset reference value intervals. The preset reference value interval may be interval one where the scene motion reference value is greater than 50 km / h, and interval two where the scene motion reference value is not greater than 50 km / h. In an embodiment of the present application, in interval one and interval two, the corresponding relationship between the scene motion reference value and the first influence value may be different. For example, in interval one, the corresponding relationship between the scene motion reference value and the first influence value is y1(x). In interval two, the corresponding relationship between the scene motion reference value and the first influence value is y2(x). In an embodiment of the present application, after determining the scene motion reference value, the preset reference value interval corresponding to the scene motion reference value may be determined. Then, the corresponding relationship between the scene motion reference value and the first reference influence value in the preset reference value interval may be determined, and based on this, the first influence value may be determined. For example, after determining that the scene motion reference value is 30 km / h, it may be determined that the preset reference value interval is interval two, and then the corresponding relationship between the scene motion reference value and the first influence value in interval two may be determined as y2(x). Finally, the candidate first influence reference value corresponding to the scene motion reference value may be determined as the first influence value according to the corresponding relationship.

[0080] Further, in an embodiment of the present application, the reference influence value includes a second influence value set for the second video frame sequence;

[0081] Determining the at least one reference influence value based on the scene motion reference value includes:

[0082] Step 601: Determine a second preset reference value interval corresponding to the scene motion reference value;

[0083] Step 603: Determine the candidate reference influence value corresponding to the second preset reference value interval as the second influence value.

[0084] In the embodiments of the present application, the reference influence value may further include a second influence value set for the second video frame sequence. Specifically, the determination of the second influence value may refer to the method for determining the first influence value described above, which is not elaborated herein in the present application. It should be noted that the first preset reference value range and the second preset reference value range may be the same or different, and the present application does not limit this here.

[0085] Further, in an embodiment of the present application, the method is applied to the training process of the Nth round of the target detection network, and the reference influence value includes a first influence value set for the first video frame sequence;

[0086] Determining the at least one reference influence value based on the scene motion reference value includes:

[0087] Step 701: Determine the historical first influence value in the Nth round of training process of the target detection network;

[0088] Step 703: In response to the scene motion reference value being greater than the first reference value threshold, adjust the historical first influence value downwards to obtain the first influence value.

[0089] In the embodiments of the present application, during the Nth round of training of the target detection network 300, the first influence value may be determined based on the scene motion reference value. Specifically, in an embodiment of the present application, a reference value threshold may be preset, and the reference value threshold may be set by the user according to actual needs. In an embodiment of the present application, the historical first influence value in the Nth round of training process of the target detection network 300 may be determined. The historical first influence value may include the first influence value in the (N - 1)th round of training process of the target detection network 300. When the scene motion reference value is greater than the first reference threshold, the historical first influence value may be adjusted downwards to obtain the first influence value. Of course, when the scene motion reference value is not greater than the first reference threshold, the historical first influence value may be adjusted upwards to obtain the first influence value.

[0090] Further, in an embodiment of the present application, the method is applied to the training process of the Nth round of the target detection network, and the reference influence value includes a second influence value set for the second video frame sequence;

[0091] Determining the at least one reference influence value based on the scene motion reference value includes:

[0092] Step 801: Determine the historical second influence value in the Nth round of training process of the target detection network;

[0093] Step 803: In response to the scene motion reference value being greater than the second reference value threshold, increase the historical second influence value to obtain the second influence value.

[0094] In the embodiment of the present application, during the Nth round of training of the target detection network 300, the second influence value can be determined based on the scene motion reference value. Specifically, for the adjustment of the second influence value, reference can be made to the above method for adjusting the first influence value, which will not be elaborated herein. It should be noted that the first reference value threshold and the second reference value threshold can be set to the same threshold or different thresholds, which is not limited in the present application.

[0095] Further, in an embodiment of the present application, before using the feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence, the following steps are also included:

[0096] Step 901: Obtain the video to be detected;

[0097] Step 903: Determine the first video frame sequence based on the video to be detected;

[0098] Step 905: Perform convolution operation and downsampling processing on each video frame of the first video frame sequence to obtain the feature information of each video frame.

[0099] In the embodiment of the present application, there are multiple ways to obtain the video to be detected. In some embodiments, the video to be detected can be collected in real time by the collection device 101. In other embodiments, the video to be detected can also be obtained from a standard image library. The standard image library can include, for example, the PASCAL VOC dataset, the MS COCO dataset, etc. After obtaining the video to be detected, the first video frame sequence can be determined. Then, convolution operation and downsampling processing can be performed on each video frame of the first video frame sequence to obtain the feature information of each video frame. Specifically, in an embodiment of the present application, the feature information can be obtained by using a feature extraction network to extract features from each video frame. The feature information can include feature information such as the gray scale, edge, texture, color, and histogram of oriented gradients of the target object.

[0100] Further, in an embodiment of the present application, the method for using the image detection sub-network to determine the target detection result of the ith video frame based on the memory information of the ith video frame includes:

[0101] Step 1001: Perform foreground modeling on the ith video frame based on the memory information of the ith video frame to obtain a foreground image;

[0102] Step 1003: Perform image detection on the foreground image to obtain the object detection result of the i-th video frame.

[0103] In the embodiment of the present application, the image detection sub-network 303 can be used to perform foreground modeling on the i-th video frame to determine the foreground image of the i-th video frame. The foreground image may include an image of a moving target object. Then, image detection can be performed on the foreground image to determine the object detection result of the i-th video frame. The object detection result may include the position information and category information of the moving object. In another embodiment of the present application, foreground modeling may also be performed through the feature processing sub-network 301 to obtain the foreground image, and then the image detection sub-network may perform image detection on the foreground image to determine the object detection result of the i-th video frame. The present application does not limit this here.

[0104] After training the object detection network according to the above method, a high-precision object detection network can be obtained. During the detection process of moving objects, the object detection network can be used to perform object detection on the video frame sequence. Based on this, on the other hand, the present application also provides an object detection method, which may include:

[0105] Step 1101: Obtain a video to be detected;

[0106] Step 1103: Input the video to be detected into the object detection network, and use the object detection network to perform feature processing and image detection to obtain the object detection results of each video frame in the video to be detected, where: the object detection network is obtained by the training method of the object detection network described above.

[0107] In the embodiment of the present application, the object detection network can perform object detection on the video to be detected. The object detection network is trained according to the training method of the object detection network provided in the above embodiment. As Figure 5As shown, the target detection network 300 may include a cascaded structure consisting of a feature extraction sub-network 305, a feature processing sub-network 301, and an image detection sub-network 303. Among them, the feature extraction sub-network 305 may be composed of multiple convolutional layers and pooling layers, and is used to extract the feature information of each video frame in the image to be detected. The output end of the feature extraction sub-network 305 may be connected to the input end of the feature processing sub-network 301, so as to send the feature information of each video frame to the feature processing sub-network 301. The feature processing sub-network 301 may be used to process the feature information and determine the foreground feature information corresponding to the feature information, so as to obtain the foreground feature information of each video frame. The output end of the feature processing sub-network 301 may be connected to the input end of the image detection sub-network 303, so as to send the foreground feature information to the image detection sub-network 303. The image detection sub-network 303 may include a detection network layer and a positioning network layer. The image detection sub-network 303 may perform convolution operations and downsampling processing on the foreground feature information to determine the position information and category information of the moving targets in each video frame of the video to be detected.

[0108] In the embodiments of the present application, the video to be detected can be obtained in various ways. For example, it can be obtained by real-time collection by the above-mentioned collection device 101 for a certain area to be detected. After obtaining the video to be detected, the target detection network can be used to process the video to be detected to determine the target detection results of each video frame in the video to be detected. In an embodiment of the present application, the target detection network 300 can be used to perform feature processing and image detection on the video to be detected. Specifically, each video frame in the video to be detected can be input into the feature extraction sub-network 305 in sequence. The feature extraction sub-network 305 can extract features from each video frame to determine the feature information of each video frame. As Figure 5As shown, taking the i-th video frame in the video to be detected as an example, the feature processing sub-network 301 can be used to process the feature information of the i-th video frame to determine the memory information of each video frame. The processing may include processing the historical memory information of the i-th video frame and the candidate memory information of the i-th video frame to obtain the memory information of the i-th video frame. Specifically, at least one influence value can be used to perform weighted fusion processing on the historical memory information and the candidate memory information. The historical memory information may include the memory information of the (i - 1)-th video frame, and the candidate memory information may include the feature information of the i-th video frame. After determining the memory information of the i-th video frame, the image detection sub-network 303 can be used to process the memory information to determine the target detection result of the i-th video frame. The target detection result may include the position information and category information of the target object in the i-th video frame.

[0109] Through the above embodiments, the target detection network trained by the above embodiments can be applied to the target detection method, thereby improving the efficiency and accuracy of detecting targets by the target detection method.

[0110] On the other hand, the present application also provides a training device 600 for a target detection network, as Figure 6 shown, the device 600 includes:

[0111] A feature information acquisition module 601, configured to use the feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the video to be detected collected for the target scene;

[0112] An influence value determination module 603, configured to obtain a scene motion reference value based on the fusion feature information; the scene motion reference value characterizes the motion information of the target included in the first video frame sequence; determine at least one reference influence value based on the scene motion reference value; the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a continuous plurality of video frames in the video to be detected, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frame;

[0113] The network parameter adjustment module 605 is configured to perform object detection sub-operations on each video frame in the second video frame sequence by using the feature processing sub-network and the image detection sub-network, obtain prediction loss information corresponding to each video frame in the second video frame sequence, and adjust network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the object detection sub-operation of the i-th video frame includes:

[0114] Using the feature processing sub-network, based on at least one determined reference influence value, process the historical memory information of the i-th video frame and process the candidate memory information of the i-th video frame to obtain the memory information of the i-th video frame; the candidate memory information is determined based on video frames before the i-th video frame in the video to be detected; and

[0115] Using the image detection sub-network, based on the memory information of the i-th video frame, determine the object detection result of the i-th video frame, and based on the object detection result and the object annotation result of the i-th video frame, determine the corresponding prediction loss information of the i-th video frame.

[0116] Optionally, in an embodiment of the present application, the adjusting network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information includes:

[0117] Obtain comprehensive prediction loss information based on the prediction loss information corresponding to each video frame in the second video frame sequence;

[0118] Adjust network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the comprehensive prediction loss information.

[0119] Optionally, in an embodiment of the present application, the reference influence value includes a first influence value set for the first video frame sequence;

[0120] The determining the at least one reference influence value based on the scene motion reference value includes:

[0121] Determine a first preset reference value interval corresponding to the scene motion reference value;

[0122] Determine the candidate reference influence value corresponding to the first preset reference value interval as the first influence value.

[0123] Optionally, in an embodiment of the present application, the reference influence value includes a second influence value set for the second video frame sequence;

[0124] Determining the at least one reference influence value based on the scene motion reference value includes:

[0125] Determining a second preset reference value interval corresponding to the scene motion reference value;

[0126] Determining the candidate reference influence value corresponding to the second preset reference value interval as the second influence value.

[0127] Optionally, in an embodiment of the present application, the method is applied to the training process of the Nth round of the target detection network, and the reference influence value includes a first influence value set for the first video frame sequence;

[0128] Determining the at least one reference influence value based on the scene motion reference value includes:

[0129] Determining the historical first influence value during the Nth round of training of the target detection network;

[0130] In response to the scene motion reference value being greater than the first reference value threshold, reducing the historical first influence value to obtain the first influence value.

[0131] Optionally, in an embodiment of the present application, determining the at least one reference influence value based on the scene motion reference value further includes:

[0132] In response to the scene motion reference value not being greater than the first reference value threshold, increasing the historical first influence value to obtain the first influence value.

[0133] Optionally, in an embodiment of the present application, the method is applied to the training process of the Nth round of the target detection network, and the reference influence value includes a second influence value set for the second video frame sequence;

[0134] Determining the at least one reference influence value based on the scene motion reference value includes:

[0135] Determining the historical second influence value during the Nth round of training of the target detection network;

[0136] In response to the scene motion reference value being greater than the second reference value threshold, increasing the historical second influence value to obtain the second influence value.

[0137] Optionally, in an embodiment of the present application, determining the at least one reference influence value based on the scene motion reference value further includes:

[0138] In response to the scene motion reference value not being greater than the second reference value threshold, reducing the historical second influence value to obtain the second influence value.

[0139] Optionally, in an embodiment of the present application, the determining the object detection result of the i-th video frame based on the memory information of the i-th video frame by using the image detection sub-network includes:

[0140] Performing foreground modeling on the i-th video frame based on the memory information of the i-th video frame to obtain a foreground image;

[0141] Performing image detection on the foreground image to obtain the object detection result of the i-th video frame.

[0142] Optionally, in an embodiment of the present application, before determining the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence by using the feature processing sub-network, it further includes:

[0143] Obtaining a video to be detected;

[0144] Determining a first video frame sequence based on the video to be detected;

[0145] Performing a convolution operation and downsampling processing on each video frame of the first video frame sequence to obtain the feature information of each video frame.

[0146] On the other hand, the present application further provides a processing device, including a memory and a processor, where computer program instructions are stored in the memory, and the processor is configured to run the computer program instructions to execute the training method of the object detection network described in each of the above embodiments.

[0147] Among them, the processing device may be a physical device or a cluster of physical devices, or a virtualized cloud device, such as at least one cloud computing device in a cloud computing cluster. For ease of understanding, the present application takes the processing device as an independent physical device to illustrate the structure of the processing device.

[0148] As Figure 7 shown, the processing device 700 includes: a processor and a memory for storing computer program instructions of the processor; among them, the processor is configured to implement the above device when executing the computer program instructions. The processing device 700 includes a memory 701, a processor 703, a bus 705, and a communication interface 707. The memory 701, the processor 703, and the communication interface 707 communicate with each other through the bus 705. The bus 705 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7It is only represented by a thick line, but it does not mean that there is only one bus or one type of bus. The communication interface 707 is used for external communication.

[0149] Among them, the processor 703 can be a central processing unit (CPU). The memory 701 can include volatile memory, such as random access memory (RAM). The memory 701 can also include non-volatile memory, such as read-only memory (ROM), flash memory, HDD or SSD, etc.

[0150] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0151] On the other hand, this application also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the steps of the training method of the above-mentioned target detection network are realized.

[0152] A computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, such as punch cards or raised structures in grooves storing instructions thereon, and any suitable combination of the above.

[0153] The computer program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer program instructions from the network and forwards the computer program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0154] The computer program instructions for performing the operations of the present application can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Small talk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., by using an Internet service provider to connect via the Internet). In some embodiments, by using the status information of the computer program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer program instructions to implement various aspects of the present application.

[0155] Aspects of the present application are described herein with reference to the flowchart and / or block diagram of methods, apparatuses according to embodiments of the present application. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions.

[0156] These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, result in an apparatus that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0157] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses, systems, and methods according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.

[0159] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A training method for an object detection network, characterized in that The target detection network at least includes a feature processing sub-network and an image detection sub-network, and the method includes: Using the feature processing sub-network, determining the fused feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the to-be-detected video collected for the target scene; Obtaining a scene motion reference value based on the fused feature information; the scene motion reference value characterizes the motion information of the targets included in the first video frame sequence; determining at least one reference influence value based on the scene motion reference value; the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a continuous plurality of video frames in the to-be-detected video, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frame; Using the feature processing sub-network and the image detection sub-network, respectively performing a target detection sub-operation on each video frame in the second video frame sequence, obtaining the prediction loss information corresponding to each video frame in the second video frame sequence, and adjusting the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the target detection sub-operation of the i-th video frame includes: Using the feature processing sub-network, based on the determined at least one reference influence value, performing a weighted fusion process on the historical memory information of the i-th video frame and the candidate memory information of the i-th video frame, obtaining the memory information of the i-th video frame; the candidate memory information is determined based on the video frames before the i-th video frame in the to-be-detected video; the historical memory information is determined based on the memory information of the (i - 1)-th video frame in the to-be-detected video; and Using the image detection sub-network, performing foreground modeling on the i-th video frame based on the memory information of the i-th video frame, obtaining a foreground image, performing image detection on the foreground image, determining the target detection result of the i-th video frame, and determining the prediction loss information corresponding to the i-th video frame based on the target detection result and the target annotation result of the i-th video frame.

2. The training method of the object detection network according to claim 1, characterized in that The adjusting the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the prediction loss information includes: Obtaining comprehensive prediction loss information based on the prediction loss information corresponding to each video frame in the second video frame sequence; Adjusting the network parameters of at least one of the feature processing sub-network and the image detection sub-network based on the comprehensive prediction loss information.

3. The training method of the object detection network according to claim 1, wherein, The reference influence value includes a first influence value set for the first video frame sequence; The determining the at least one reference influence value based on the scene motion reference value includes: Determining a first preset reference value interval corresponding to the scene motion reference value; Determining the candidate reference influence value corresponding to the first preset reference value interval as the first influence value.

4. The training method of the object detection network according to claim 1, characterized in that The reference influence value includes a second influence value set for the second video frame sequence; Determining the at least one reference influence value based on the scene motion reference value includes: Determining a second preset reference value interval corresponding to the scene motion reference value; Determining the candidate reference influence value corresponding to the second preset reference value interval as the second influence value.

5. The training method of the object detection network according to claim 1, characterized in that The method is applied in the training process of the Nth round of the target detection network, and the reference influence value includes a first influence value set for the first video frame sequence; Determining the at least one reference influence value based on the scene motion reference value includes: Determining the historical first influence value in the Nth round of training process of the target detection network; In response to the scene motion reference value being greater than the first reference value threshold, reducing the historical first influence value to obtain the first influence value.

6. The training method of the object detection network according to claim 5, wherein Determining the at least one reference influence value based on the scene motion reference value further includes: In response to the scene motion reference value not being greater than the first reference value threshold, increasing the historical first influence value to obtain the first influence value.

7. The training method of the object detection network according to claim 1, wherein The method is applied in the training process of the Nth round of the target detection network, and the reference influence value includes a second influence value set for the second video frame sequence; Determining the at least one reference influence value based on the scene motion reference value includes: Determining the historical second influence value in the Nth round of training process of the target detection network; In response to the scene motion reference value being greater than the second reference value threshold, increasing the historical second influence value to obtain the second influence value.

8. The training method of the object detection network according to claim 7, wherein Determining the at least one reference influence value based on the scene motion reference value further includes: In response to the scene motion reference value not being greater than the second reference value threshold, reducing the historical second influence value to obtain the second influence value.

9. The training method of the object detection network according to any one of claims 1-8, characterized in that Using the image detection sub-network to determine the target detection result of the ith video frame based on the memory information of the ith video frame includes: Performing foreground modeling on the ith video frame based on the memory information of the ith video frame to obtain a foreground image; Performing image detection on the foreground image to obtain the target detection result of the ith video frame.

10. The training method of the object detection network according to any one of claims 1-8, characterized in that, Before using the feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence, it further includes: Obtaining a video to be detected; Determining a first video frame sequence based on the video to be detected; Performing a convolution operation and downsampling process on each video frame of the first video frame sequence to obtain the feature information of each video frame.

11. A target detection method, characterized in that, Includes: Obtaining a video to be detected; Inputting the video to be detected into the target detection network, and using the target detection network to perform feature processing and image detection to obtain the target detection results of each video frame in the video to be detected, where: The target detection network is obtained by the training method of the target detection network described in claims 1-10.

12. A training device for an object detection network, characterized in that, The device includes: A feature information acquisition module, configured to use a feature processing sub-network to determine the fusion feature information of the first video frame sequence based on the feature information of each video frame in the first video frame sequence; the first video frame sequence includes a continuous plurality of video frames in the to-be-detected video collected for the target scene; An influence value determination module, configured to obtain a scene motion reference value based on the fusion feature information; the scene motion reference value characterizes the motion information of the target included in the first video frame sequence; determine at least one reference influence value based on the scene motion reference value; the reference influence value is used to characterize the importance degree of the feature information of the first video frame sequence and / or the second video frame sequence; the second video frame sequence includes a continuous plurality of video frames in the to-be-detected video, and the acquisition time of the video frames in the first video frame sequence is earlier than the acquisition time of the video frames in the second video frame sequence; A network parameter adjustment module, configured to use the feature processing sub-network and the image detection sub-network to perform a target detection sub-operation on each video frame in the second video frame sequence respectively, obtain the prediction loss information corresponding to each video frame in the second video frame sequence, and adjust the network parameters of at least one network in the feature processing sub-network and the image detection sub-network based on the prediction loss information; wherein, the target detection sub-operation of the i-th video frame includes: Using the feature processing sub-network, based on the determined at least one reference influence value, processing the historical memory information of the i-th video frame and performing weighted fusion processing on the candidate memory information of the i-th video frame to obtain the memory information of the i-th video frame; the candidate memory information is determined based on the video frames before the i-th video frame in the to-be-detected video; the historical memory information is determined based on the memory information of the (i - 1)-th video frame in the to-be-detected video; and Using the image detection sub-network, performing foreground modeling on the i-th video frame based on the memory information of the i-th video frame to obtain a foreground image, performing image detection on the foreground image to determine the target detection result of the i-th video frame, and determining the prediction loss information corresponding to the i-th video frame based on the target detection result and the target annotation result of the i-th video frame.

13. A processing device, comprising a memory and a processor, the memory storing computer program instructions, characterized in that, When the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Detecting objects and determining confidence scores

    CN111133447A

  • Real-time event abstracting method based on consistency monitoring

    CN111639176A