Target detection method and device and electronic equipment

By determining the ontology information and its correlation degree of the pending subject target in the target detection, the problem of low accuracy of target detection is solved, false detection and missed detection are reduced, and the accuracy of detection is improved.

CN120198652APending Publication Date: 2025-06-24INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495109.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The accuracy of target detection is low, and there are problems of false detection and missed detection.

Method used

When a pending subject target is detected, its ontology information is determined, and whether the pending subject target is a subject target is further determined based on the first degree of correlation with the associated target subset and the second degree of correlation with the non-associated target subset.

Benefits of technology

The error detection and missed detection situations occur when target detection is performed based on the ontological characteristics of the to-determined subject targets are reduced, and the accuracy of target detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198652A_ABST
    Figure CN120198652A_ABST
Patent Text Reader

Abstract

The invention provides a target detection method, a target detection device and electronic equipment, and relates to the technical field of computers. Further determining whether the to-be-determined subject target is a subject target according to the ontology information of the to-be-determined subject target, a first association degree between the to-be-determined subject target and the associated target subset, and a second association degree between the to-be-determined subject target and the non-associated target subset; according to the invention, false detection and missing detection conditions generated when target detection is carried out only according to the ontology characteristics of the to-be-determined main target can be reduced, and the accuracy of target detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to an object detection method, apparatus, and electronic device. Background Art

[0002] Object detection is one of the core problems in the field of computer vision. Its task is to find the objects of interest in an image or video and determine the categories and locations of the objects. In related technologies, only the ontological features of the objects of interest are used to detect whether there are objects of interest in the image or video, resulting in problems of false detection and missed detection of objects, and the accuracy of object detection is relatively low. Summary of the Invention

[0003] The present disclosure provides an object detection method, apparatus, and electronic device. Its main purpose is to solve the problem of relatively low accuracy of object detection.

[0004] According to a first aspect of the present disclosure, there is provided an object detection method, including:

[0005] When it is detected that there is a to-be-detected main object in the to-be-detected video data, determining the ontological information of the to-be-detected main object in the to-be-detected video data;

[0006] When the ontological information meets the ontological logic requirements, determining a non-main object set in the to-be-detected video data except the to-be-detected main object, where the non-main object set includes an associated object subset and a non-associated object subset;

[0007] Determining a first association degree between the to-be-detected main object and the associated object subset and a second association degree between the to-be-detected main object and the non-associated object subset;

[0008] When the first association degree meets the first association degree requirement and the second association degree meets the second association degree requirement, determining the to-be-detected main object as the main object.

[0009] According to a second aspect of the present disclosure, there is provided an object detection apparatus, including:

[0010] An information determination unit, configured to determine the ontological information of the to-be-detected main object in the to-be-detected video data when it is detected that there is a to-be-detected main object in the to-be-detected video data;

[0011] A set determination unit, configured to determine a non-main object set in the to-be-detected video data except the to-be-detected main object when the ontological information meets the ontological logic requirements, where the non-main object set includes an associated object subset and a non-associated object subset;

[0012] An association detection unit, configured to determine a first association degree between a to-be-determined subject target and an associated target subset, and a second association degree between the to-be-determined subject target and a non-associated target subset;

[0013] A condition determination unit, configured to determine the to-be-determined subject target as a subject target when the first association degree meets the first association degree requirement and the second association degree meets the second association degree requirement.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the foregoing first aspect.

[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the method described in the foregoing first aspect when executed by a processor.

[0020] Through the present disclosure, after detecting the to-be-determined subject target in the to-be-detected video data, further determining whether the to-be-determined subject target is a subject target according to the ontology information of the to-be-determined subject target, the first association degree between the to-be-determined subject target and the associated target subset, and the second association degree between the to-be-determined subject target and the non-associated target subset, it is possible to reduce the false detection and missed detection situations generated when only performing target detection according to the ontology features of the to-be-determined subject target, and improve the accuracy of target detection.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0023] Figure 1 It is a schematic flowchart of a target detection method provided by an embodiment of the present disclosure;

[0024] Figure 2 A flowchart of another target detection method provided by an embodiment of the present disclosure;

[0025] Figure 3 A flowchart of another target detection method provided by an embodiment of the present disclosure;

[0026] Figure 4 A schematic diagram of the structure of a target detection device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] The target detection method, device, and electronic device according to the embodiments of the present disclosure are described below with reference to the accompanying drawings.

[0029] Figure 1 A flowchart of a target detection method provided by an embodiment of the present disclosure.

[0030] like Figure 1 As shown, the method comprises the following steps:

[0031] Step 101 : when it is detected that there is a subject target to be determined in the video data to be detected, the entity information of the subject target to be determined in the video data to be detected is determined.

[0032] The video data to be detected refers to the video data that needs to be detected. Video data refers to a continuous image sequence, which is actually composed of a group of continuous images. For the image itself, there is no structural information except the order in which it appears.

[0033] The video data to be detected can be obtained by placing an imaging device, such as a camera or a video camera, at a target location. For example, when the type of the subject target to be determined is a vehicle, the video data to be detected can be obtained by a camera installed at a gas station.

[0034] The pending subject target refers to an object detected in the video data to be detected and needs to be further confirmed whether it is a subject target. The pending subject target may be, for example, a target detected in the video data to be detected according to features corresponding to the subject target.

[0035] Among them, the main target refers to the target of interest that needs to be found as indicated by the target detection task. The ontology information refers to the information related to the ontology of the to-be-determined main target.

[0036] Step 102: When the ontology information meets the ontology logic requirements, determine the set of non-main targets in the video data to be detected except for the to-be-determined main target.

[0037] Among them, the ontology logic requirements refer to the requirements related to the ontology behavior logic of the main target. For example, when the type of the main target is a vehicle, the ontology logic requirements corresponding to the vehicle include requirements such as the vehicle being on the road and the vehicle not being in the sky. Therefore, when the to-be-determined main target is detected, the accuracy of target detection can be improved by further confirming the to-be-determined main target according to the ontology logic requirements.

[0038] Among them, the set of non-main targets refers to a set composed of other non-main targets detected in the video data to be detected except for the to-be-determined main target. The set of non-main targets includes an associated target subset and a non-associated target subset. The associated targets in the associated target subset refer to the targets that have an association relationship with the main target. The associated targets in the non-associated target subset refer to the targets that have no association relationship with the main target. For example, when the type of the main target is a vehicle, there is an association relationship between the vehicle and street lights, roadblocks, traffic lights, and roads, and there is no association relationship between the vehicle and coffee tables, sofas, and refrigerators.

[0039] Step 103: Determine the first association degree between the to-be-determined main target and the associated target subset and the second association degree between the to-be-determined main target and the non-associated target subset.

[0040] Among them, the first association degree refers to the degree of association between the to-be-determined main target and the associated target subset in the video data to be detected.

[0041] Among them, the second association degree refers to the degree of association between the to-be-determined main target and the non-associated target subset in the video data to be detected.

[0042] Step 104: When the first association degree meets the first association degree requirement and the second association degree meets the second association degree requirement, determine the to-be-determined main target as the main target.

[0043] Among them, if the first correlation degree meets the first correlation degree requirement and the second correlation degree meets the second correlation degree requirement, it indicates that there is a strong correlation between the to-be-detected subject target and the associated target subset in the to-be-detected video data, and there is no strong correlation between the to-be-detected subject target and the non-associated target subset. In this case, it can be determined that the to-be-detected subject target is the subject target, thereby determining that there is a subject target in the to-be-detected video data. Therefore, it is possible to reduce the situation where the to-be-detected subject target is determined to be the subject target because there is a strong correlation between the to-be-detected subject target and the associated target subset, but there is also a strong correlation between the to-be-detected subject target and the non-associated target subset. For example, determining the vehicle displayed on the living room TV screen in the to-be-detected video data as the subject target, resulting in false detection, and the accuracy of target detection can be improved.

[0044] In summary, for the method provided by the embodiments of the present disclosure, after detecting the to-be-detected subject target in the to-be-detected video data, further determining whether the to-be-detected subject target is the subject target according to the ontology information of the to-be-detected subject target, the first correlation degree between the to-be-detected subject target and the associated target subset, and the second correlation degree between the to-be-detected subject target and the non-associated target subset can reduce the false detection and missed detection situations generated when performing target detection only based on the ontology characteristics of the to-be-detected subject target, and can improve the accuracy of target detection.

[0045] It should be noted that there may be multiple steps in the embodiments of the present disclosure. For the convenience of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0046] Furthermore, in a possible implementation manner of this embodiment, Figure 2 is a schematic flowchart of another target detection method provided by the embodiments of the present disclosure. As Figure 2 shown, this method includes the following steps:

[0047] Step 201, obtain the to-be-detected video data, and perform a preliminary detection of the subject target on the to-be-detected video data.

[0048] According to some embodiments, the preliminary detection of the subject target for the video data to be detected can be performed by constructing the image features corresponding to the subject target and combining with a sliding window; among them, the image features corresponding to the subject target can be constructed by means such as Histogram of Oriented Gradient (HOG), Scale Invariant Feature Transform (SIFT), etc. The preliminary detection of the subject target for the video data to be detected can also be performed by combining a classifier with a sliding window; among them, the classifier includes but is not limited to Support Vector Machine (SVM), etc. It is also possible to adopt a model based on Convolutional Neural Networks (CNN) to realize the joint optimization of feature automatic extraction and target detection through end-to-end training, so as to realize the preliminary detection of the subject target for the video data to be detected.

[0049] In some embodiments, when performing the preliminary detection of the subject target for the video data to be detected based on the target detection network, the following steps can be adopted:

[0050] Step 2011, access the video stream and obtain the video data to be detected from the video stream.

[0051] Among them, pictures can be taken out from the video data to be detected, and the taken-out pictures are preprocessed, so as to scale the taken-out pictures to a preset size. During the process of scaling the picture size, the aspect ratio of the picture can be kept unchanged for scaling, and the maximum side is scaled to the preset size, and the small side is filled with grayscale.

[0052] Among them, the preset size does not specifically refer to a certain fixed value. The preset size can be adjusted according to the actual application scenario. For example, the preset size can be 640*640.

[0053] Step 2012, perform continuous video frame processing on the obtained video data to be detected, and form a picture sequence according to the time sequence for the first number of video frames obtained after the continuous video frame processing.

[0054] According to some embodiments, the first number is not greater than the carrying threshold, and the carrying threshold refers to the maximum number of video frames that the target detection network can parse in real time. The carrying threshold can be, for example, the maximum number of video frames that can be parsed by the hardware.

[0055] In some embodiments, the first number is greater than the minimum number threshold. The minimum number threshold does not specifically refer to a certain fixed threshold. The minimum number threshold can be adjusted according to the actual application scenario. For example, the minimum number threshold can be 2.

[0056] Taking a scenario as an example, when the main target is a moving target, the first quantity threshold can be determined according to the following formula:

[0057]

[0058] Where F is the number of video frames selected per second, i.e., the first quantity, max_frequent is the carrying threshold, obj_v is the maximum moving speed of the main target, obj_s is the minimum size of the main target, and obj_num is the number of main targets.

[0059] That is to say, F needs to satisfy the above 5 conditions, F is greater than or equal to 3, F is less than or equal to max_frequent, and F is proportional to obj_v, obj_s, and obj_num.

[0060] Taking a scenario as an example, the video frames obtained by processing and parsing consecutive video frames can be first composed into an initial picture sequence according to the time sequence; then, the initial picture sequence is input into the target detection network to test the first quantity, and when the number of initial video frames is less than the first quantity, new video data is continuously obtained to reorganize the video frames to obtain the initial picture sequence until the number of initial video frames is not less than the first quantity; and when the number of initial video frames is greater than the carrying threshold, the video frames are reorganized according to the carrying threshold to obtain the picture sequence.

[0061] It should be noted that by using the maximum moving speed and the carrying threshold for two-way limitation, the number of video frames input into the target detection network is dynamically changed. For example, if the moving speed of the to-be-determined main target is slow, the minimum quantity threshold can be used; if the target moving speed is relatively fast, it can be set as the carrying threshold. Therefore, the detection efficiency and detection effect of the subsequent target detection network for target detection can be improved.

[0062] Step 2013, input the picture sequence into the target detection network to determine whether there is a to-be-determined main target in the to-be-detected video data.

[0063] According to some embodiments, the scores of multiple targets in the to-be-detected video data can be obtained through the target detection network; in the case where the score is greater than the set threshold, it is determined that there is a to-be-determined main target in the to-be-detected video data; in the case where the score is not greater than the score threshold, it is determined that there is no to-be-determined main target in the to-be-detected video data.

[0064] In some embodiments, the target box moving speed, target box position, and target box category of multiple targets in the to-be-detected video data can also be obtained through the target detection network; according to the coordinates corresponding to the target box position, the target box size and aspect ratio are calculated; according to the position of the to-be-determined main target in the picture sequence, the detection target box picture is obtained.

[0065] According to some embodiments, in the object detection network, first, video frames in a picture sequence can be sequentially grouped into picture groups in accordance with the time sequence, starting from the first frame and in accordance with a second quantity; then, the obtained picture groups can be sequentially input into a first-layer long short-term memory (LSTM) recurrent neural network and a dropout network, a second-layer LSTM and dropout network, and a feature extraction convolutional neural network (CNN) to obtain the extracted feature maps; thereafter, the extracted feature maps can be input into a region proposal network (RPN) to obtain the object recommendation regions of the picture; then, the obtained feature maps and object recommendation regions can be simultaneously input into a region of interest align (ROI Align) network to obtain the feature maps of the object size; then, the feature maps of the object size can be input into a head layer to obtain the object detection regions; finally, the object detection regions can be input into a CNN layer to respectively obtain the moving speed of the object bounding box, the position of the object bounding box, the category of the object bounding box, and the score.

[0066] It should be noted that when the object detection network inputs video frames, processing them as picture groups with a continuous second quantity of video frames as a group can not only achieve the input of different quantities of video frames but also avoid the problem of limited model input caused by not knowing the quantity of video frames.

[0067] Among them, the second quantity does not specifically refer to a fixed value. The second quantity can be adjusted according to the actual application scenario. For example, the second quantity can be 3. In this case, pictures with serial numbers 0, 1, and 2 in the picture sequence can be grouped into a picture group, pictures with serial numbers 1, 2, and 3 can be grouped into a picture group, and so on.

[0068] In some embodiments, in the ROI Align network, the object recommendation regions containing bounding boxes (bboxes) can be equally divided according to the object size required by the output to obtain the feature maps of the object size.

[0069] Among them, the object size does not specifically refer to a fixed size. The object size can be adjusted according to the actual application scenario. The object size can be, for example, 2x2.

[0070] Among them, since when performing equal division, the vertices do not fall on real pixel points, in this case, 4 fixed blue points can be taken in each block. For each blue point, the values of the 4 real pixel points closest to it are weighted and / or bilinearly interpolated to obtain the values of these 4 blue points, and the maximum value among the values of these four blue points is taken as the output value of this block.

[0071] Step 202, when it is detected that there is a to-be-detected subject target in the to-be-detected video data, determine the body information of the to-be-detected subject target in the to-be-detected video data.

[0072] According to some embodiments, the body information includes static information and dynamic information.

[0073] Among them, the static information refers to the information recorded for the to-be-detected subject target at a certain point in time, which will not be continuously updated over time. The static information includes, but is not limited to, the position, size, etc. of the to-be-detected subject target on each video frame in the to-be-detected video data.

[0074] Among them, the dynamic information refers to the information recorded when the to-be-detected subject target is moving, which can be continuously updated over time. For example, the dynamic information includes, but is not limited to, the speed and direction of the to-be-detected subject target in the to-be-detected video data.

[0075] Step 203, when the time sequence of the video frames indicating the existence of the to-be-detected subject target in the static information meets the time sequence requirement, and the dynamic information meets the body dynamic logic requirement, determine the non-subject target set in the to-be-detected video data except the to-be-detected subject target.

[0076] According to some embodiments, the time sequence requirement refers to the requirement adopted when determining whether the static information meets the subject logic requirement. For example, when the type of the to-be-detected subject target is a vehicle and it is not occluded in the to-be-detected video data, if the time sequence corresponding to the to-be-detected subject target is discontinuous instead of the complete continuous time sequence corresponding to the to-be-detected video data, it indicates that there is a misjudgment in the detection of the subject target, that is, the to-be-detected subject target is not a subject target, because a vehicle cannot appear discontinuously in the real scene.

[0077] In some embodiments, the body dynamic logic requirement refers to the requirement adopted when determining whether the dynamic information meets the subject logic requirement. For example, when the type of the to-be-detected subject target is a vehicle, if the speed corresponding to the to-be-detected subject target is within the preset speed range, it can be considered that the dynamic information meets the body dynamic logic requirement. If the speed corresponding to the to-be-detected subject target is not within the preset speed range, it indicates that there is a misjudgment in the detection of the subject target. Among them, the preset speed range can be determined according to the actual application scenario.

[0078] Taking a scenario as an example, when the type of the to-be-detected subject target is a vehicle and it has been determined that the time sequence corresponding to the to-be-detected subject target is the complete continuous time sequence corresponding to the to-be-detected video data, the speed corresponding to the to-be-detected subject target can be determined according to the pixel position of the first picture, the pixel position of the last picture, and the sampling frequency of the to-be-detected subject target in the to-be-detected video data.

[0079] In some embodiments, the speed corresponding to the subject target to be determined can be calculated according to the following formula:

[0080]

[0081] where v is the speed corresponding to the subject target to be determined, H is the vertical height from the camera to the ground, f is the camera focal length, (x1, y1) is the pixel position of the subject target to be determined in the first picture of the video data to be detected, and (x2, y2) is the pixel position of the subject target to be determined in the last picture of the video data to be detected.

[0082] Among them, by inputting the time interval of the video frames and the pixel positions of the subject target to be determined in the video frames, the speed of the subject target to be determined can be simply calculated.

[0083] It should be noted that by combining the static information and dynamic information of the subject target to be determined to determine whether the ontology information meets the ontology logic requirements, the detection accuracy and accuracy of target detection can be improved.

[0084] Step 204, determine the first spatial association degree and / or the first semantic association degree between the subject target to be determined and the associated target subset to obtain the first association degree; determine the second spatial association degree and / or the second semantic association degree between the subject target to be determined and the non-associated target subset to obtain the second association degree.

[0085] According to some embodiments, the spatial association degree is used to indicate the spatial relationship between the subject target to be determined in the video data to be detected and the non-subject target set. The first spatial association degree can be obtained by determining the distances and angles between the subject target to be determined and multiple associated targets in the associated target subset; the second spatial association degree can be obtained by determining the distances and angles between the subject target to be determined and multiple non-associated targets in the non-associated target subset.

[0086] In some embodiments, the distance and angle between the subject target to be determined and any non-subject target in the non-subject target set can be calculated according to the following formula:

[0087]

[0088] where D is the distance between the subject target to be determined and any non-subject target in the non-subject target set, is the angle between the subject target to be determined and any non-subject target in the non-subject target set, (X a , Y a ) is the coordinate of the subject target to be determined in the world coordinate system, and (X b , Y b ) is the coordinate of any non-subject target in the non-subject target set in the world coordinate system.

[0089] Among them, the coordinates of any non-subject target in the world coordinate system in the set of to-be-determined subject targets or non-subject targets can be determined according to the following formula:

[0090]

[0091] Among them, θ is the camera elevation angle, (x0, y0) is the pixel coordinates of the camera optical center, and (x, y) is the coordinates of any non-subject target in the set of to-be-determined subject targets or non-subject targets in the camera coordinate system.

[0092] According to some embodiments, the semantic correlation degree is used to indicate the semantic relationship between the to-be-determined subject target and the set of non-subject targets in the video data to be detected. The first co-occurrence probability between the to-be-determined subject target and multiple associated targets in the associated target subset can be obtained, and the first semantic correlation degree can be determined according to the first co-occurrence probability; the second co-occurrence probability between the to-be-determined subject target and multiple non-associated targets in the non-associated target subset can be obtained, and the second semantic correlation degree can be determined according to the second co-occurrence probability.

[0093] Among them, the first semantic correlation degree is the sum of the first co-occurrence probabilities corresponding to multiple associated targets in the associated target subset. The first co-occurrence probability refers to the probability that the to-be-determined subject target and a certain associated target in the associated target subset appear together in the picture. This first co-occurrence probability can be obtained, for example, by acquiring a set of picture samples related to the subject target and analyzing this set of picture samples.

[0094] Among them, the second semantic correlation degree is the sum of the second co-occurrence probabilities corresponding to multiple non-associated targets in the non-associated target subset. The second co-occurrence probability refers to the probability that the to-be-determined subject target and a certain non-associated target in the non-associated target subset appear together in the picture. This second co-occurrence probability can also be obtained, for example, by acquiring a set of picture samples related to the subject target and analyzing this set of picture samples.

[0095] In some embodiments, the first co-occurrence probability and the second co-occurrence probability can be determined according to the following formula:

[0096]

[0097] Among them, R sem (i1) is the first co-occurrence probability, R sem (i2) is the second co-occurrence probability, P1 is the score of the to-be-determined subject target calculated by the target detection network, is the first co-occurrence probability between the subject target and i1, is the second co-occurrence probability between the subject target and i2.

[0098] It should be noted that by determining whether the to-be-determined main target is a main target according to the spatial correlation degree and / or semantic correlation degree between the to-be-determined main target and the non-main target, the detection accuracy and accuracy of target detection can be improved.

[0099] Step 205: When the first correlation degree meets the first correlation degree requirement and the second correlation degree meets the second correlation degree requirement, determine that the to-be-determined main target is a main target.

[0100] Step 206: When the first correlation degree does not meet the first correlation degree requirement and / or the second correlation degree does not meet the second correlation degree requirement, determine that the to-be-determined main target is not a main target.

[0101] Among them, when at least one of the following situations exists, it can be considered that the first correlation degree does not meet the first correlation degree requirement:

[0102] The distance between the to-be-determined main target and multiple associated targets in the associated target subset is greater than the first distance threshold;

[0103] The angle between the to-be-determined main target and multiple associated targets in the associated target subset is outside the first angle range;

[0104] The first semantic correlation degree is less than the first semantic correlation degree threshold.

[0105] Among them, when at least one of the following situations exists, it can be considered that the second correlation degree does not meet the second correlation degree requirement:

[0106] The distance between the to-be-determined main target and multiple non-associated targets in the non-associated target subset is less than the second distance threshold;

[0107] The angle between the to-be-determined main target and multiple non-associated targets in the non-associated target subset is outside the second angle range;

[0108] The second semantic correlation degree is greater than the second semantic correlation degree threshold.

[0109] Furthermore, in a possible implementation manner of this embodiment, Figure 3 is a schematic flowchart of another target detection method provided by the embodiments of the present disclosure. It is for connecting a gas station camera to a hardware system deployed for the target detection method, and the type of the main target is a vehicle. As Figure 3As shown in the figure, first, video data of the gas station camera is collected; then, continuous video frames are processed to obtain a sequence of pictures; then, the number of video frames that can be parsed in real time by the hardware, that is, the number of the sequence of pictures, is determined. If the number of the determined sequence of pictures does not reach the first number, the sequence of pictures is input into the target detection network to determine the number of video frames that can be detected in real time by the hardware, that is, the first number, and the video data of the gas station camera is collected again and the first number is updated; if the number of the sequence of pictures reaches the first number, the first number is directly updated; then, it is determined whether the number of the sequence of pictures is not greater than the maximum number of video frames that can be parsed by the hardware. If so, the preprocessed sequence of pictures is directly input into the target detection network. If not, the sequence of pictures is reorganized with the maximum number of video frames that can be parsed by the hardware and the reorganized sequence of pictures is input into the target detection network; then, the target detection network calculates the size, position, category, and score of the target box according to the input sequence of pictures; then, it is determined whether there is a pending main target according to the score. If not, the detection of the pending main target is performed again according to the target detection network. If so, it is determined whether the time series corresponding to the pending main target is continuous; if the time series is not continuous, the detection of the pending main target is performed according to the target detection network; if the time series is continuous, the speed of the pending main target is determined according to the time series; if the speed is not within the preset speed range, the detection of the pending main target is performed according to the target detection network; if the speed is within the preset speed range, it is determined whether the pending main target satisfies the spatial relationship with other non-main targets; if the spatial relationship is not satisfied, the detection of the pending main target is performed according to the target detection network; if the spatial relationship is satisfied, it is determined whether the pending main target satisfies its own semantic relationship; if the semantic relationship is not satisfied, the detection of the pending main target is performed according to the target detection network; if it is satisfied, the pending main target is determined to be the main target, and the target detection ends.

[0110] It should be noted that, without changing the target detection network and when a pending main target is detected in the video data to be detected according to the target detection network, by introducing long-chain thinking into the target detection network and further determining whether the pending main target is the main target according to the ontology information of the pending main target, the first association degree between the pending main target and the associated target subset, and the second association degree between the pending main target and the non-associated target subset, the detection accuracy and accuracy of the target detection can be improved; moreover, the method provided by the embodiments of the present disclosure abandons the target tracking network, which can avoid the errors caused by target tracking, reduce the calculation amount of the target detection at the same time, and improve the detection accuracy of the target detection.

[0111] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. However, in many cases, the former is a better implementation method.

[0112] According to an embodiment of the present disclosure, the present disclosure also provides an object detection device.

[0113] Exemplarily, Figure 4 FIG. is a schematic structural diagram of an object detection device provided by an embodiment of the present disclosure. The object detection device 400 includes: an information determination module 401, a set determination module 402, an association detection module 403, and a condition determination module 404; wherein,

[0114] The information determination module 401 is configured to determine the body information of the to-be-determined main target in the to-be-detected video data when it is detected that there is a to-be-determined main target in the to-be-detected video data.

[0115] The set determination module 402 is configured to determine a non-main target set in the to-be-detected video data except for the to-be-determined main target when the body information meets the body logic requirements, where the non-main target set includes an associated target subset and a non-associated target subset.

[0116] The association detection module 403 is configured to determine a first association degree between the to-be-determined main target and the associated target subset and a second association degree between the to-be-determined main target and the non-associated target subset.

[0117] The condition determination module 404 is configured to determine that the to-be-determined main target is the main target when the first association degree meets the first association degree requirement and the second association degree meets the second association degree requirement.

[0118] Further, the body information includes static information and dynamic information. After determining the body information of the to-be-determined main target in the to-be-detected video data, the condition determination module 404 is further configured to:

[0119] If the timing of the video frame indicating the existence of the to-be-determined main target in the to-be-detected video data in the static information meets the timing requirement and the dynamic information meets the body dynamic logic requirement, then the body information meets the body logic requirement.

[0120] Further, the type of the to-be-determined main target is a vehicle, and the dynamic information includes speed. The condition determination module 404 is further configured to:

[0121] If the speed is within a preset speed range, then the dynamic information meets the body dynamic logic requirement.

[0122] Further, when the association detection module 403 is configured to determine the first association degree between the to-be-determined main target and the associated target subset and the second association degree between the to-be-determined main target and the non-associated target subset, specifically:

[0123] Determine the first spatial correlation degree and / or the first semantic correlation degree between the to-be-determined subject target and the associated target subset to obtain the first correlation degree;

[0124] Determine the second spatial correlation degree and / or the second semantic correlation degree between the to-be-determined subject target and the non-associated target subset to obtain the second correlation degree.

[0125] Further, when the association detection module 403 is used to determine the first spatial correlation degree between the to-be-determined subject target and the associated target subset and determine the second spatial correlation degree between the to-be-determined subject target and the non-associated target subset, it specifically is used for:

[0126] Determine the distances and angles between the to-be-determined subject target and multiple associated targets in the associated target subset to obtain the first spatial correlation degree;

[0127] Determine the distances and angles between the to-be-determined subject target and multiple non-associated targets in the non-associated target subset to obtain the second spatial correlation degree.

[0128] Further, after determining the first correlation degree between the to-be-determined subject target and the associated target subset and the second correlation degree between the to-be-determined subject target and the non-associated target subset, the condition determination module 404 is further used for:

[0129] If the distance between the to-be-determined subject target and multiple associated targets in the associated target subset is greater than the first distance threshold, the first correlation degree does not meet the first correlation degree requirement;

[0130] If the angle between the to-be-determined subject target and multiple associated targets in the associated target subset is outside the first angle range, the first correlation degree does not meet the first correlation degree requirement;

[0131] If the distance between the to-be-determined subject target and multiple non-associated targets in the non-associated target subset is less than the second distance threshold, the second correlation degree does not meet the second correlation degree requirement;

[0132] If the angle between the to-be-determined subject target and multiple non-associated targets in the non-associated target subset is outside the second angle range, the second correlation degree does not meet the second correlation degree requirement.

[0133] Further, when the association detection module 403 is used to determine the first semantic correlation degree between the to-be-determined subject target and the associated target subset and determine the second semantic correlation degree between the to-be-determined subject target and the non-associated target subset, it specifically is used for:

[0134] Obtain the first coexistence probability between the to-be-determined subject target and multiple associated targets in the associated target subset, and determine the first semantic correlation degree according to the first coexistence probability, where the first semantic correlation degree is the sum of the first coexistence probabilities corresponding to multiple associated targets in the associated target subset;

[0135] Obtain the second coexistence probability between the target of the subject to be determined and multiple non-associated targets in the non-associated target subset, and determine the second semantic association degree according to the second coexistence probability, where the second semantic association degree is the sum of the second coexistence probabilities corresponding to multiple non-associated targets in the non-associated target subset.

[0136] Further, after determining the first association degree between the target of the subject to be determined and the associated target subset and the second association degree between the target of the subject to be determined and the non-associated target subset, the condition determination module 404 is further configured to:

[0137] If the first semantic association degree is less than the first semantic association degree threshold, the first association degree does not meet the first association degree requirement;

[0138] If the second semantic association degree is greater than the second semantic association degree threshold, the second association degree does not meet the second association degree requirement.

[0139] It should be noted that for the description of the features in the embodiments corresponding to the target detection device, reference can be made to the relevant descriptions in the embodiments corresponding to the target detection method, which will not be elaborated here one by one.

[0140] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned target detection method embodiments.

[0141] An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned target detection method embodiments when running.

[0142] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.

[0143] An embodiment of the present disclosure further provides a computer program product. The above-mentioned computer program product includes a computer program, and the steps in any of the above-mentioned target detection method embodiments are implemented when the computer program is executed by a processor.

[0144] An embodiment of the present disclosure further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in any of the above-mentioned target detection method embodiments are implemented when the computer program is executed by a processor.

[0145] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0146] The above has introduced in detail a target detection method provided by the present disclosure. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present disclosure, several improvements and modifications can be made to the present disclosure, and these improvements and modifications also fall within the protection scope of the claims of the present disclosure.

Claims

1. A target detection method, characterized in that: include: In the case where it is detected that there is a subject target to be determined in the video data to be detected, determining the entity information of the subject target to be determined in the video data to be detected; In the case where the ontology information meets the ontology logic requirements, determining a set of non-subject targets other than the subject target to be determined in the video data to be detected, wherein the non-subject target set includes an associated target subset and an unassociated target subset; Determine a first degree of association between the pending subject target and the associated target subset and a second degree of association between the pending subject target and the non-associated target subset; When the first degree of association satisfies a first degree of association requirement and the second degree of association satisfies a second degree of association requirement, the undetermined subject target is determined as the subject target.

2. The method according to claim 1, characterized in that The entity information includes static information and dynamic information. After determining the entity information of the subject target to be determined in the video data to be detected, the method further includes: If the static information indicates that the timing of the video frame of the to-be-determined subject target in the to-be-detected video data satisfies the timing requirement, and the dynamic information satisfies the ontology dynamic logic requirement, then the ontology information satisfies the ontology logic requirement.

3. The method according to claim 2, characterized in that The type of the subject target to be determined is a vehicle, the dynamic information includes speed, and the method further includes: If the speed is within the preset speed range, the dynamic information meets the dynamic logic requirements of the entity.

4. The method according to claim 1, characterized in that: The determining of a first degree of association between the pending subject target and the associated target subset and a second degree of association between the pending subject target and the non-associated target subset includes: Determine a first spatial correlation and / or a first semantic correlation between the undetermined main target and the associated target subset to obtain a first correlation; Determine a second spatial correlation and / or a second semantic correlation between the undetermined main target and the non-associated target subset to obtain a second correlation.

5. The method according to claim 4, characterized in that Determining a first spatial correlation between the pending main target and the associated target subset, and determining a second spatial correlation between the pending main target and the non-associated target subset, comprises: Determine the distance and angle between the undetermined main target and a plurality of associated targets in the associated target subset to obtain a first spatial association degree; The distance and angle between the undetermined main target and a plurality of non-associated targets in the non-associated target subset are determined to obtain a second spatial correlation degree.

6. The method according to claim 5, characterized in that After determining the first degree of association between the pending subject target and the associated target subset and the second degree of association between the pending subject target and the non-associated target subset, the method further includes: If the distance between the undetermined main target and multiple associated targets in the associated target subset is greater than a first distance threshold, then the first degree of association does not meet the first degree of association requirement; If the angle between the undetermined main target and multiple associated targets in the associated target subset is outside the first angle range, the first degree of association does not meet the first degree of association requirement; If the distance between the undetermined main target and a plurality of non-associated targets in the non-associated target subset is less than a second distance threshold, then the second degree of association does not meet the second degree of association requirement; If the angle between the undetermined main target and a plurality of non-associated targets in the non-associated target subset is outside the second angle range, the second degree of association does not meet the second degree of association requirement.

7. The method according to claim 4, characterized in that Determining a first semantic association between the pending subject target and the associated target subset, and determining a second semantic association between the pending subject target and the non-associated target subset, comprises: Obtaining a first coexistence probability between the undetermined main target and a plurality of associated targets in the associated target subset, and determining a first semantic association degree according to the first coexistence probability, wherein the first semantic association degree is the sum of the first coexistence probabilities corresponding to the plurality of associated targets in the associated target subset; Obtain a second coexistence probability between the undetermined main target and multiple non-associated targets in the non-associated target subset, and determine a second semantic association based on the second coexistence probability, wherein the second semantic association is the sum of the second coexistence probabilities corresponding to multiple non-associated targets in the non-associated target subset.

8. The method according to claim 7, characterized in that After determining the first degree of association between the pending subject target and the associated target subset and the second degree of association between the pending subject target and the non-associated target subset, the method further includes: If the first semantic relevance is less than a first semantic relevance threshold, then the first relevance does not meet the first relevance requirement; If the second semantic relevance is greater than a second semantic relevance threshold, the second relevance does not meet a second relevance requirement.

9. A target detection device, characterized in that: include: An information determination module, for determining the entity information of the to-be-determined subject target in the to-be-determined video data when detecting that there is a to-be-determined subject target in the to-be-determined video data; A set determination module, used for determining a set of non-subject targets other than the subject target to be determined in the video data to be detected, when the subject information satisfies the subject logic requirements, wherein the non-subject target set includes an associated target subset and an unassociated target subset; An association detection module, used to determine a first degree of association between the pending main target and the associated target subset and a second degree of association between the pending main target and the non-associated target subset; A condition determination module is used to determine that the pending main target is a main target when the first correlation degree meets a first correlation degree requirement and the second correlation degree meets a second correlation degree requirement.

10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.