Target detection method and device, electronic equipment and storage medium
By utilizing the features of the target detection box and the 3D coordinate matching algorithm in multi-view multi-target tracking, the problem of insufficient accuracy in pedestrian detection is solved, and the accuracy of multi-view multi-target tracking is improved.
Patent Information
- Application Number
- CN202210759276.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2042-06-29
AI Technical Summary
In multi-view, multi-target tracking, the accuracy of pedestrian detection technology is poor, resulting in low accuracy of multi-view, multi-target tracking.
By extracting target detection boxes and features from multiple different perspectives of the current frame image, and combining the three-dimensional coordinates of the target detection boxes in the world coordinate system with the features and coordinates of the identity identifiers in the database, a weighted and similarity matching algorithm is used to determine the identity information of the target detection boxes.
It improves the accuracy of multi-view, multi-target tracking and ensures that the matching results of target detection boxes and identity identifiers are relatively accurate.
Smart Images

Figure CN115131705B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of pedestrian detection, and particularly relates to a target detection method and device, electronic equipment and a storage medium. BACKGROUND
[0002] Multi-view multi-target tracking is an important problem in the field of computer vision and intelligent video monitoring, which refers to detecting, identifying and tracking different human bodies appearing in a scene according to different view information shot by multiple cameras, so as to obtain accurate and complete human body trajectories. Multi-view multi-target tracking has important applications in many fields, including human behavior analysis, crowd counting and people flow analysis, human trajectory analysis, and detection of abnormal behavior. In the related art, the effect of pedestrian detection technology is general, and the accuracy needs to be improved. In the related art, the accuracy of multi-view multi-target tracking is poor. SUMMARY
[0003] The present disclosure provides a target detection method, device, electronic equipment and storage medium to solve the defects in the related art.
[0004] According to a first aspect of an embodiment of the present disclosure, a target detection method is provided, comprising:
[0005] extracting at least one target detection box in a current frame image of multiple different views and a first target feature of each target detection box;
[0006] determining a first three-dimensional coordinate of each target detection box in a world coordinate system according to the position of each target detection box in the current image frame;
[0007] determining identity information of at least one target detection box according to the first target feature and the first three-dimensional coordinate of each target detection box, and the second target feature and the second three-dimensional coordinate of each identity identifier in the database.
[0008] In one embodiment, the identity information of at least one target detection box is determined according to the first target feature and the first three-dimensional coordinate of each target detection box, and the second target feature and the second three-dimensional coordinate of each identity identifier in the database, comprising:
[0009] determining a first feature similarity between the first target feature of each target detection box and the second target feature of each identity identifier;
[0010] determining a first relative distance between the first three-dimensional coordinate of each target detection box and the second three-dimensional coordinate of each identity identifier;
[0011] According to the determined first feature similarity and the first relative distance, identity information of at least one of the target bounding boxes is determined.
[0012] In one embodiment, the identity information of at least one of the target bounding boxes is determined according to the determined first feature similarity and the first relative distance, comprising:
[0013] According to the first feature weight, the first distance weight, and the determined first feature similarity and the first relative distance, a matching degree of each matching combination is determined, wherein the matching combination comprises one target bounding box and one identity label.
[0014] According to the matching degree of each matching combination, the matching combination that successfully matches is determined.
[0015] The identity label included in each successfully matched matching combination is determined as the identity information of the target bounding box included in the matching combination.
[0016] In one embodiment, the matching degree of each matching combination is determined according to the first feature weight, the first distance weight, and the determined first feature similarity and the first relative distance, comprising:
[0017] According to the first feature weight, the first distance weight, the time coefficient of the identity label in each matching combination, and the determined first feature similarity and the first relative distance, the matching degree of each matching combination is determined, wherein the time coefficient is determined by the time when the identity label is last tracked and the time corresponding to the current frame image.
[0018] In one embodiment, further comprising:
[0019] The confidence of each target bounding box is determined.
[0020] The matching degree of each matching combination is determined according to the first feature weight, the first distance weight, and the determined first feature similarity and the first relative distance, comprising:
[0021] According to the first feature weight, the first distance weight, the confidence of the target bounding box in each matching combination, and the determined first feature similarity and the first relative distance, the matching degree of each matching combination is determined.
[0022] In one embodiment, further comprising:
[0023] In the plurality of target bounding boxes without identity information, according to the first target feature and the first three-dimensional coordinates of each target bounding box, at least one identity with the first target feature and the first three-dimensional coordinates is newly created.
[0024] In one embodiment, further comprising:
[0025] According to the first target feature and the first three-dimensional coordinates of each newly created identity, and the second target feature and the second three-dimensional coordinates of each identity in the database, the identity similarity between each newly created identity and each identity in the database is determined.
[0026] In the case that the identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, the identity information of the target bounding box corresponding to the newly created identity is determined as the corresponding identity in the database.
[0027] In the case that the identity similarity between any newly created identity and each identity in the database is less than the preset similarity threshold, the newly created identity is added to the database.
[0028] In one embodiment, further comprising:
[0029] In the database, according to the time when each identity is tracked last time and the time corresponding to the current frame image, the second feature weight and the second distance weight of each identity are determined.
[0030] According to the first target feature and the first three-dimensional coordinates of each newly created identity, and the second target feature and the second three-dimensional coordinates of each identity in the database, the identity similarity between each newly created identity and each identity in the database is determined.
[0031] According to the first target feature and the first three-dimensional coordinates of each newly created identity, and the second target feature, the second three-dimensional coordinates, the second feature weight and the second distance weight of each identity in the database, the identity similarity between each newly created identity and each identity in the database is determined.
[0032] In one embodiment, after the identity information of at least one target bounding box is determined, further comprising:
[0033] According to the first target feature and the first three-dimensional coordinates of each target bounding box with successfully determined identity information, the second target feature and the second three-dimensional coordinates of the corresponding identity in the database are updated; and / or,
[0034] in a case where an identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, further comprising:
[0035] updating, according to the first target feature and the first three-dimensional coordinate of the newly created identity, a second target feature and a second three-dimensional coordinate of the identity in the database.
[0036] In an embodiment, after the identity information of the at least one target detection frame is determined, further comprising:
[0037] updating, in the database, a state of the identity corresponding to each target detection frame whose identity information is successfully determined to be tracked, and updating a state of other identities to be lost;
[0038] after the identity information of the target detection frame corresponding to the newly created identity is determined to be the identity in the database, further comprising:
[0039] updating a state of the identity in the database to be tracked.
[0040] In an embodiment, the newly creating at least one identity having a first target feature and a first three-dimensional coordinate according to the first target feature and the first three-dimensional coordinate of each target detection frame comprises:
[0041] determining a second feature similarity between the first target feature of each target detection frame and the first target feature of each other target detection frame;
[0042] determining a second relative distance between the first three-dimensional coordinate of each target detection frame and the first three-dimensional coordinate of each other target detection frame;
[0043] determining a matching degree of each target detection frame combination according to a first feature weight, a first distance weight, and the second feature similarity and the second relative distance between two target detection frames in each target detection frame combination, wherein the target detection frame combination includes two different target detection frames;
[0044] newly creating at least one identity according to the matching degree of each target detection frame combination.
[0045] According to a second aspect of the embodiments of the present disclosure, a target detection device is provided, comprising:
[0046] an extraction module configured to extract at least one target detection frame in a current frame image of multiple different perspectives and a first target feature of each target detection frame;
[0047] a coordinate module, configured to determine a first three-dimensional coordinate of each of the target bounding boxes in a world coordinate system according to a position of each of the target bounding boxes in the current image frame;
[0048] a detection module, configured to determine identity information of at least one of the target bounding boxes according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes, and the second target feature and the second three-dimensional coordinate of each of the identity labels.
[0049] In an embodiment, the detection module is specifically configured to:
[0050] determine a first feature similarity between the first target feature of each of the target bounding boxes and the second target feature of each of the identity labels;
[0051] determine a first relative distance between the first three-dimensional coordinate of each of the target bounding boxes and the second three-dimensional coordinate of each of the identity labels;
[0052] determine the identity information of at least one of the target bounding boxes according to the determined first feature similarity and the first relative distance.
[0053] In an embodiment, when the detection module is configured to determine the identity information of at least one of the target bounding boxes according to the determined first feature similarity and the first relative distance, the detection module is specifically configured to:
[0054] determine a matching degree of each matching combination according to a first feature weight, a first distance weight, and the determined first feature similarity and the first relative distance, wherein the matching combination comprises one of the target bounding boxes and one of the identity labels;
[0055] determine a successfully matched matching combination according to the matching degree of each of the matching combinations;
[0056] determine the identity label included in each of the successfully matched matching combinations as the identity information of the target bounding box included in the matching combination.
[0057] In an embodiment, when the detection module is configured to determine the matching degree of each matching combination according to a first feature weight, a first distance weight, and the determined first feature similarity and the first relative distance, the detection module is specifically configured to:
[0058] determine the matching degree of each of the matching combinations according to a first feature weight, a first distance weight, a time coefficient of the identity label in each of the matching combinations, and the determined first feature similarity and the first relative distance, wherein the time coefficient is determined by a time when the identity label is last tracked and a time corresponding to the current frame image.
[0059] In an embodiment, further comprising a confidence module configured to:
[0060] determine a confidence of each of the target bounding boxes;
[0061] The detection module is configured to determine the matching degree of each matching combination according to the first feature weight, the first distance weight, and the determined first feature similarity and first relative distance, and specifically configured to:
[0062] determine the matching degree of each matching combination according to the first feature weight, the first distance weight, the confidence of the target bounding box in each matching combination, and the determined first feature similarity and first relative distance.
[0063] In an embodiment, further comprising a new module configured to:
[0064] In the plurality of target bounding boxes without determined identity information, create at least one identity label having the first target feature and the first three-dimensional coordinates according to the first target feature and the first three-dimensional coordinates of each of the target bounding boxes.
[0065] In an embodiment, further comprising a merging module configured to:
[0066] determine an identity similarity between each of the newly created identity labels and each of the identity labels in the database according to the first target feature and the first three-dimensional coordinates of each of the newly created identity labels, and the second target feature and the second three-dimensional coordinates of each of the identity labels in the database.
[0067] In a case where the identity similarity between any of the newly created identity labels and the identity labels in the database is greater than or equal to a preset similarity threshold, determine that the identity information of the target bounding box corresponding to the newly created identity label is the corresponding identity label in the database.
[0068] In a case where the identity similarity between any of the newly created identity labels and each of the identity labels in the database is less than the preset similarity threshold, add the newly created identity label to the database.
[0069] In an embodiment, further comprising a weight module configured to:
[0070] In the database, determine a second feature weight and a second distance weight of each of the identity labels according to a time when each of the identity labels was last tracked and a time corresponding to the current frame image.
[0071] The merging module is configured to determine the identity similarity between each newly created identity and each identity in the database according to the first target feature and the first three-dimensional coordinate of each newly created identity and the second target feature and the second three-dimensional coordinate of each identity in the database, and specifically configured to:
[0072] The identity similarity between each newly created identity and each identity in the database is determined according to the first target feature and the first three-dimensional coordinate of each newly created identity and the second target feature, the second three-dimensional coordinate, the second feature weight and the second distance weight of each identity in the database.
[0073] In one embodiment, the updating module is further configured to:
[0074] After the identity information of at least one target detection frame is determined, the second target feature and the second three-dimensional coordinate of the corresponding identity in the database are updated according to the first target feature and the first three-dimensional coordinate of each target detection frame whose identity information is successfully determined; and / or,
[0075] In the case where the identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, the second target feature and the second three-dimensional coordinate of the identity in the database are updated according to the first target feature and the first three-dimensional coordinate of the newly created identity.
[0076] In one embodiment, the state module is further configured to:
[0077] After the identity information of at least one target detection frame is determined, the state of the identity corresponding to each target detection frame whose identity information is successfully determined is updated to be tracked, and the state of other identities is updated to be lost in the database.
[0078] After it is determined that the identity information of the target detection frame corresponding to the newly created identity is the corresponding identity in the database, the state of the corresponding identity in the database is updated to be tracked.
[0079] In one embodiment, the newly created module is specifically configured to:
[0080] The first target feature of each target detection frame and the second feature similarity between the first target feature of each target detection frame and the first target feature of other target detection frames are determined.
[0081] The first three-dimensional coordinate of each target detection frame and the second relative distance between the first three-dimensional coordinate of each target detection frame and the first three-dimensional coordinate of other target detection frames are determined.
[0082] determine a matching degree of each of the target bounding box combinations according to the first feature weight, the first distance weight, and a second feature similarity and a second relative distance between two of the target bounding boxes in each of the target bounding box combinations, wherein the target bounding box combinations include two different target bounding boxes;
[0083] create at least one identity according to the matching degrees of each of the target bounding box combinations.
[0084] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, the device comprising a memory and a processor, the memory being configured to store computer instructions executable on the processor, and the processor being configured to implement the method of the first aspect when executing the computer instructions.
[0085] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, the medium storing a computer program, the program being executable by a processor to implement the method of the first aspect.
[0086] According to the above embodiments, by extracting at least one target bounding box in a current frame image of multiple different perspectives and a first target feature of each of the target bounding boxes, a first three-dimensional coordinate of each of the target bounding boxes in a world coordinate system can be determined according to the position of each of the target bounding boxes in the current image frame, and finally the identity information of at least one of the target bounding boxes can be determined according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes, and the second target feature and the second three-dimensional coordinate of each of the identities in the database. Since the target bounding box and the identity are matched from two dimensions of target features and three-dimensional coordinates, the matching result is more accurate, the determined identity information is more accurate, and thus the accuracy of multi-perspective multi-target tracking can be improved.
[0087] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0088] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0089] Figure 1 is a flowchart of a target detection method according to an embodiment of the present disclosure;
[0090] Figure 2 is a flowchart of a target detection method according to another embodiment of the present disclosure;
[0091] Figure 3is a structural schematic diagram of a target detection device according to an embodiment of the present disclosure.
[0092] Figure 4 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0093] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to designate the same elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not meant to represent all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0094] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0095] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used only to distinguish one from another. For example, a first information can be termed a second information, and, similarly, a second information can also be termed a first information, without departing from the scope of the present disclosure. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to a determination".
[0096] Multi-view multi-target tracking is very complex, and human body detection, recognition and tracking can be affected by many factors. Pedestrians are easily blocked by other people or objects, and human body detection in crowded scenes is prone to missed detection or false detection. Meanwhile, the human body frame blocked by objects such as shelves is also incomplete, resulting in inaccurate human body position positioning. The appearance and features of the same pedestrian under different views will have great differences, resulting in the human bodies of the same pedestrian under different views being unable to be clustered or matched together. In addition, the appearance and features of different pedestrians under different views can also be relatively similar, which can cause human body clustering or matching errors under different views.
[0097] Therefore, in a first aspect, at least one embodiment of the present disclosure provides a target detection method. Please refer to FIG. 1, which shows the flow of the method, including steps S101 to S104. Figure 1 , which shows the flow of the method, including steps S101 to S104.
[0098] The method can be used for target detection of a video or an image sequence, that is, the identity information of each target in each image frame of the video or the identity information of each target in each image of the image sequence is identified. For example, the target is a pedestrian, and the method is a pedestrian detection method, that is, the identity information of each pedestrian in each image frame of the video or the identity information of each pedestrian in each image of the image sequence is identified. It can be understood that the method is repeatedly executed for each image in the video or the image sequence. The specific steps of the method are described below by taking the processing process of an image frame as an example. The same method is repeatedly executed for other image frames, so that the pedestrian detection of the entire video or image sequence is completed.
[0099] In addition, the method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA) handheld device, a computing device, a vehicle-mounted device, a wearable device, or the like. The method can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the method can be executed by a server, which can be a local server or a cloud server.
[0100] In step S101, at least one target detection box in a plurality of current frame images of different perspectives and a first target feature of each target detection box are extracted.
[0101] The plurality of current frame images of different perspectives can be current frame images collected by a plurality of different cameras. The plurality of different cameras are respectively installed according to respective perspectives, and there is an overlap between the fields of view of the plurality of different cameras, that is, the plurality of different cameras can collect images of the same scene from different perspectives. The plurality of different cameras can synchronously record a video or continuously collect images, that is, each camera synchronously updates a current frame image.
[0102] The detection model and the feature extraction model can be pre-trained or configured, so that the detection model can be used to detect the target bounding box in the current frame image, and then the feature extraction model can be used to extract features of each target bounding box to obtain target features of each target bounding box. Alternatively, a neural network model capable of performing the above detection task and the above feature extraction task can be pre-configured. After the image is input into the model, the model can output the target bounding box of the image and the target features of the target bounding box. Taking pedestrians as an example, a human body detection model and a pedestrian ReID model can be used to detect and analyze the features of the human body in the multi-view current frame image, to obtain the human body detection box, the confidence of the human body detection box, and the human body features, respectively.
[0103] In one possible embodiment, the plurality of detected target bounding boxes can also be filtered before the target features of the target bounding boxes are extracted. For example, the plurality of target bounding boxes can be filtered by the aspect ratio of the target bounding box and the distance from the bottom edge of the target bounding box to the bottom edge of the image, that is, the target bounding boxes with an aspect ratio exceeding a preset aspect ratio range and / or a distance from the bottom edge to the bottom edge of the image exceeding a distance threshold are deleted. Thus, the quality of the target bounding box can be improved, and low-quality or incomplete target bounding boxes can be avoided to interfere with the subsequent target detection results.
[0104] In addition, the confidence of the target bounding box can also be obtained at the same time as the target bounding box is detected. Further, the confidence of each target bounding box can also be normalized.
[0105] In step S102, a first three-dimensional coordinate of each target bounding box in a world coordinate system is determined according to a position of each target bounding box in the current image frame.
[0106] In the above embodiment, the position of the target bounding box can be represented by a point at a specific position in the target bounding box, that is, the three-dimensional coordinate of the point at the specific position in the three-dimensional coordinate system is determined to represent the first three-dimensional coordinate of the target bounding box in the world coordinate system. For example, the point at the specific position can be the center point of the bottom edge of the target bounding box.
[0107] For example, the center point of the bottom edge of the target bounding box in the current frame image can be calculated, and then the pixel coordinate can be converted into an actual three-dimensional coordinate in the world coordinate system according to the camera calibration parameters of the camera used to capture the current frame image, so as to serve as the first three-dimensional coordinate of the target bounding box in the world coordinate system.
[0108] In step S103, identity information of at least one of the target bounding boxes is determined according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes, and the second target feature and the second three-dimensional coordinate of each of the identity labels in the database. The identity information is represented by an identity label.
[0109] This step can be performed when the current frame image is a non-first frame image. When the current frame image is a first frame image, at least one identity label with the second target feature and the second three-dimensional coordinate can be newly created according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes, and the newly created identity label is saved in the database.
[0110] In one possible embodiment, this step can be performed in the manner shown in FIG. 10, including sub-steps S1031 to S1032. Figure 2
[0111] In sub-step S1031, a first feature similarity between the first target feature of each of the target bounding boxes and the second target feature of each of the identity labels is determined.
[0112] For example, the cosine distance can be used to calculate the first feature similarity between the first target feature and the second target feature. A first feature similarity matrix A1_feat can be constructed to represent the first feature similarity between each of the first target features and each of the second target features. Each row of A1_feat represents a target bounding box, and each column represents an identity label. Therefore, an element in A1_feat is the first feature similarity a1_feat between the target bounding box represented by the row of the element and the identity label represented by the column of the element.
[0113] In sub-step S1032, a first relative distance between the first three-dimensional coordinate of each of the target bounding boxes and the second three-dimensional coordinate of each of the identity labels is determined.
[0114] For example, a first relative distance matrix can be constructed to represent the first relative distance between each of the first three-dimensional coordinates and each of the second three-dimensional coordinates. Each row of the first relative distance matrix represents a target bounding box, and each column represents an identity label. Therefore, an element in the first relative distance matrix is the first relative distance between the target bounding box represented by the row of the element and the identity label represented by the column of the element.
[0115] In sub-step S1033, identity information of at least one of the target bounding boxes is determined according to the determined first feature similarity and the first relative distance.
[0116] The identity information of the target bounding box determined in this sub-step is the result of target detection. Since the target bounding box is extracted from the current frame image from multiple different perspectives, different target bounding boxes can be determined as the same identity information.
[0117] Optionally, the following manner (i.e., the following first step, second step and third step) is used to perform this sub-step:
[0118] In the first step, a matching degree of each matching combination is determined according to the first feature weight, the first distance weight, and the first feature similarity and the first relative distance between each of the target bounding boxes and each of the identity labels, wherein the matching combination includes one of the target bounding boxes and one of the identity labels.
[0119] For example, the first relative distance d1 can be converted into a first correlation degree a1_d according to the following formula: a1_d = exp(-lambda_m*d1), where lambda_m is a first conversion coefficient. Further, each element in the first relative distance matrix can be converted into a corresponding first correlation degree, thereby obtaining a first correlation degree matrix A1_d. Each row in A1_d represents a target bounding box, and each column represents an identity label. Therefore, a certain element in A1_d is the first correlation degree between the target bounding box represented by the row where the element is located and the identity label represented by the column where the element is located. Then, the first correlation degree is used to calculate the matching degree a1 of the matching combination: a1 = a1_d*(1-alpha_f)*a1_feat*alpha_f, where 1-alpha_f is the first distance weight, and alpha_f is the first feature weight. It can be understood that if A1_feat and A_d have been constructed, the matrix can be calculated directly according to the above formula, thereby obtaining a first matching degree matrix A1. Each row in A1 represents a target bounding box, and each column represents an identity label. Therefore, a certain element in A1 is the matching degree a1 of the matching combination composed of the target bounding box represented by the row where the element is located and the identity label represented by the column where the element is located.
[0120] In a possible embodiment, the matching degree of each matching combination can be calculated based on the first feature weight, the first distance weight, the first feature similarity and the first relative distance determined between each of the target detection boxes and each of the identity labels, and further combined with a time coefficient of each identity label. That is, the matching degree of each matching combination is determined according to the first feature weight, the first distance weight, the time coefficient of the identity label in each matching combination, and the first feature similarity and the first relative distance determined between each of the target detection boxes and each of the identity labels, wherein the time coefficient is determined by the time when the identity label is last tracked and the time corresponding to the current frame image. For example, the matching degree a1 can be calculated according to the following formula: a1 = a1_d*(1-alpha_f)*a1_feat*alpha_f*t_factor, wherein t_factor = exp(-lambda_t*(t-t_valid)), lambda_t is a second conversion coefficient, t is the time corresponding to the current frame image, and t_valid is the time when the identity label is last tracked. In this embodiment, the time when the identity label is last tracked is added to increase the influence on the matching degree, thereby further improving the accuracy of the matching degree calculation.
[0121] In a possible embodiment, the matching degree of each matching combination can be calculated based on the first feature weight, the first distance weight, the first feature similarity and the first relative distance determined between each of the target detection boxes and each of the identity labels, and further combined with a time coefficient of each identity label. That is, the matching degree of each matching combination is determined according to the first feature weight, the first distance weight, the time coefficient of the identity label in each matching combination, and the first feature similarity and the first relative distance determined between each of the target detection boxes and each of the identity labels, wherein the time coefficient is determined by the time when the identity label is last tracked and the time corresponding to the current frame image. For example, the matching degree a1 can be calculated according to the following formula: a1 = a1_d*(1-alpha_f)*a1_feat*alpha_f*t_factor, wherein t_factor = exp(-lambda_t*(t-t_valid)), lambda_t is a second conversion coefficient, t is the time corresponding to the current frame image, and t_valid is the time when the identity label is last tracked. In this embodiment, the time when the identity label is last tracked is added to increase the influence on the matching degree, thereby further improving the accuracy of the matching degree calculation.
[0122] It can be understood that the two embodiments described above can be combined to calculate the matching degree of each matching combination, that is, the matching degree of each matching combination is calculated according to the following formula: a1: a1 = a1_d*(1-alpha_f)*a1_feat*alpha_f*t_factor*norm_score.
[0123] Secondly, according to the matching degree of each matching combination, the successful matching matching combination is determined. For example, the target detection frame and the identity in the matching combination with a matching degree higher than a preset matching degree threshold can be determined as successful matching. Alternatively, a bipartite graph matching method can be used to determine whether each matching combination is successfully matched.
[0124] Thirdly, the identity in each successful matching matching combination is determined as the identity information of the target detection frame in the matching combination.
[0125] It can be understood that after step S103 is completed, the second target feature and the second three-dimensional coordinate of the identity in the database can also be updated according to the first target feature and the first three-dimensional coordinate of each target detection frame whose identity information is successfully determined. Thus, the second target feature of the identity can be further enriched, and the second three-dimensional coordinate of the identity can be kept as the latest three-dimensional coordinate. Further, the trajectory of each identity can be determined according to the historical update records of the second three-dimensional coordinate of each identity in the database. When updating the second target feature, the first target feature of the target detection frame whose identity information is successfully determined and the second target feature of the corresponding identity in the database can be weighted and summed according to a preset weight relationship, and the obtained result can replace the second target feature of the corresponding identity in the database. When updating the second three-dimensional coordinate of the identity, the first three-dimensional coordinate of the target detection frame can be used to replace the second three-dimensional coordinate of the identity.
[0126] It can be understood that after step S103 is completed, the state of the identity corresponding to each target detection frame whose identity information is successfully determined can be updated to tracked in the database, and the state of other identities can be updated to lost. Thus, the states of the identities in the database can be updated in real time, and the trajectory of each identity can be easily determined by a user. In addition, the identities in the database that have been in the lost state for a duration longer than a preset time threshold can be deleted, so as to avoid that invalid identities occupy the database memory for a long time, and to reduce the calculation when step S103 is performed.
[0127] According to the above embodiment, by extracting at least one target detection box in a plurality of current frame images of different perspectives and a first target feature of each target detection box, the first three-dimensional coordinates of each target detection box in the world coordinate system can be determined according to the position of each target detection box in the current image frame, and finally the identity information of at least one target detection box can be determined according to the first target feature and the first three-dimensional coordinates of each target detection box, and the second target feature and the second three-dimensional coordinates of each identity in the database. Since the target detection box and the identity are matched from two dimensions of target features and three-dimensional coordinates, the matching result is more accurate, the determined identity information is more accurate, and the accuracy of multi-perspective multi-target tracking can be improved.
[0128] In some embodiments of the present disclosure, after step S103 is completed, at least one identity having the first target feature and the first three-dimensional coordinates can be newly created in the plurality of target detection boxes whose identity information is not determined according to the first target feature and the first three-dimensional coordinates of each target detection box.
[0129] In one possible embodiment, the newly created identity can be added to the database, thereby increasing the number of identities in the database and improving the richness of the database. In this way, in the case of multi-perspective multi-target tracking, the identity of part of the targets in each frame of image is determined, and then a new identity is generated from the target whose identity is not determined, thereby gradually enriching the database. In other words, in the case of multi-perspective multi-target tracking, the targets in each frame of image are tracked, the successfully tracked targets are determined to have an identity, and the unsuccessfully tracked targets are used to generate a new identity.
[0130] In another possible embodiment, the identity similarity between each newly created identity and each identity in the database can be determined according to the first target feature and the first three-dimensional coordinates of each newly created identity, and the second target feature and the second three-dimensional coordinates of each identity in the database. Then, in the case where the identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, the identity information of the target detection box corresponding to the newly created identity is determined to be the corresponding identity in the database. In the case where the identity similarity between any newly created identity and each identity in the database is less than the preset similarity threshold, the newly created identity is added to the database.
[0131] Wherein, before calculating the identity similarity of the identity identifiers, the second feature weight and the second distance weight of each of the identity identifiers can be determined in the database according to the time when each of the identity identifiers was tracked last time and the time corresponding to the current frame image. Then, when calculating the identity similarity of two identity identifiers, the identity similarity of each of the newly created identity identifiers and each of the identity identifiers in the database can be determined according to the first target feature and the first three-dimensional coordinate of each of the newly created identity identifiers, and the second target feature, the second three-dimensional coordinate, the second feature weight and the second distance weight of each of the identity identifiers in the database. For example, the identity similarity of two identity identifiers can be calculated by the following formula s :
[0132] a s = exp(-lambda_merge_1*d s )*a s _feat, if abs(t-t_valid) <= t_thr
[0133] a s = exp(-lambda_merge_2*d s )*a s _feat, if abs(t-t_valid) > t_thr
[0134] Wherein, lambda_merge_1 is a third conversion coefficient, lambda_merge_2 is a fourth conversion coefficient, lambda_merge_1 > lambda_merge_2, d s is the relative distance between the three-dimensional coordinates of two identity identifiers, t is the time corresponding to the current frame image, t_valid is the time when the identity identifier was tracked last time, and t_thr is a pre-set time threshold.
[0135] Through the second distance weight and the second feature weight, the relative distance and the feature similarity of two identity information can be weighted equally when the identity information is tracked all the time, and the weight of the relative distance of two identity information can be reduced when the identity information is lost for a period of time. In the formula of the above example, the second distance weight and the second feature weight are adjusted by switching lambda_merge_1 and lambda_merge_2.
[0136] In the embodiment, after the new identity is created, whether the new identity and the identity in the database are the same identity is determined by identity similarity. If they are the same identity, the two identities are merged. If they are not the same identity, the new identity is added to the database. Thus, the identity in the database can be prevented from being duplicated, and the accuracy is improved. Moreover, the target detection frame whose identity is not successfully determined in step S103 can successfully determine the identity in the embodiment, and thus the accuracy of target detection is further improved.
[0137] It can be understood that, in a case where the identity similarity between any new identity and the identity in the database is greater than or equal to a preset similarity threshold, the second target feature and the second three-dimensional coordinate of the corresponding identity in the database can also be updated according to the first target feature and the first three-dimensional coordinate of the new identity. Thus, the second target feature of the identity can be further enriched, and the second three-dimensional coordinate of the identity is kept as the latest three-dimensional coordinate. The manner of updating the second target feature and the second three-dimensional coordinate has been described in detail in the foregoing, and thus will not be repeated here.
[0138] It can be understood that, after the identity information of the target detection frame corresponding to the new identity is determined as the corresponding identity in the database, the state of the corresponding identity in the database can also be updated to be tracked. Thus, the state of each identity in the database is updated in real time, and the trajectory of the target of each identity is convenient for a user to determine.
[0139] In some embodiments of the present disclosure, whether the first target feature and the first three-dimensional coordinate of the target detection frame in the first frame of image are used or the first target feature and the first three-dimensional coordinate of the target detection frame whose identity information is not determined are used to create the identity, the identity can be created in the following manner:
[0140] Firstly, a second feature similarity between the first target feature of each of the target bounding boxes and the first target feature of each of the other target bounding boxes is determined. That is, the second feature similarity between the first target features of each of the target bounding boxes is calculated. Exemplarily, the second feature similarity between the first target features can be calculated using a cosine distance. A second feature similarity matrix A2 feat can be constructed to represent the second feature similarity between each two of the first target features, each row of A2 feat represents a target bounding box (all target bounding boxes have a represented row), each column also represents a target bounding box (all target bounding boxes have a represented column), thus an element in A2 feat is the second feature similarity a2 feat between the target bounding box represented by the row of the element and the target bounding box represented by the column of the element.
[0141] Next, a second relative distance between the first three-dimensional coordinates of each of the target bounding boxes and the first three-dimensional coordinates of each of the other target bounding boxes is determined. That is, the second relative distance between the first three-dimensional coordinates of each of the target bounding boxes is calculated. Exemplarily, a second relative distance matrix can be constructed to represent the second relative distance between each two of the target bounding boxes, each row of the second relative distance matrix represents a target bounding box (all target bounding boxes have a represented row), each column also represents a target bounding box (all target bounding boxes have a represented column), thus an element in the second relative distance matrix is the second relative distance between the target bounding box represented by the row of the element and the target bounding box represented by the column of the element.
[0142] Next, according to the first feature weight, the first distance weight, and the second feature similarity and the second relative distance between each two of the target bounding boxes in each of the target bounding box combinations, a matching degree of each of the target bounding box combinations is determined, wherein the target bounding box combination includes two different target bounding boxes.
[0143] Exemplarily, the second relative distance d2 can be converted into a second correlation degree a2_d according to the following formula: a2_d = exp(-lambda_m*d2), where lambda_m is a first conversion coefficient. Each element in the second relative distance matrix can be further converted into a corresponding second correlation degree, so as to obtain a second correlation degree matrix A2_d. Each row in A2_d represents a target detection frame (all target detection frames have a represented row), and each column also represents a target detection frame (all target detection frames have a represented column). Therefore, a certain element in A2_d is the second correlation degree between the target detection frame represented by the row where the element is located and the target detection frame represented by the column where the element is located. Then, the second correlation degree is used to calculate the matching degree a2 of the target detection frame combination: a2 = a2_d*(1-alpha_f)*a2_feat*alpha_f, where 1-alpha_f is a first distance weight, and alpha_f is a first feature weight. It can be understood that, if A2_feat and A2_d have been constructed, the first matching degree matrix A2 can be directly calculated according to the above formula, so that each row in A2 represents a target detection frame (all target detection frames have a represented row), and each column also represents a target detection frame (all target detection frames have a represented column). Therefore, a certain element in A2 is the matching degree a2 of the target detection frame combination composed of the target detection frame represented by the row where the element is located and the target detection frame represented by the column where the element is located.
[0144] In addition, the confidence of at least one target detection frame in the target detection frame combination can also be added in the formula for calculating the matching degree a2, so that the formula is updated as: a2 = a2_d*(1-alpha_f)*a2_feat*alpha_f*norm_score, where norm_score can be the confidence of the target detection frame corresponding to the row, or the confidence of the target detection frame corresponding to the column, or the average of the confidence of the two target detection frames. By adding the confidence, the influence of the confidence of the target detection frame on the matching degree is increased, so as to further improve the accuracy of the matching degree calculation.
[0145] Finally, at least one identity is newly created according to the matching degree of each target detection frame combination.
[0146] Exemplarily, all the target detection boxes can be clustered according to the matching degrees of each target detection box combination and a clustering prior principle (for example, only one target detection box in the image of each view can be included in each cluster), for example, two target detection boxes in a target detection box combination with a matching degree higher than a preset matching degree threshold are determined as target detection boxes of the same target, and then the same target detection boxes in different target detection box combinations are combined, so that a plurality of target detection box cluster sets can be obtained, each cluster set corresponds to a target, and finally the target detection boxes in each cluster set are screened by using the clustering prior principle.
[0147] Further, at least one identity can be newly created according to the clustering result, for example, each cluster set is determined as an identity, the first target features of all the target detection boxes in each cluster set are weighted and summed to determine the first target feature of the corresponding identity, and the first three-dimensional coordinates of all the target detection boxes in each cluster set are averaged to determine the first three-dimensional coordinates of the corresponding identity.
[0148] Please refer to the accompanying drawings Figure 2 which exemplarily show the complete flow of the target detection method. As can be seen from the accompanying drawings Figure 2 , when the current frame image is the first frame image, a plurality of identities can be newly created according to the plurality of target detection boxes, the first target features and the first three-dimensional coordinates of each target detection box, and a database is constructed. As can be seen from the accompanying drawings Figure 2 , when the current frame image is a non-first frame image, the plurality of target detection boxes can be matched with the identities in the database according to the first target features and the first three-dimensional coordinates of each target detection box, the target detection boxes successfully matched with the identities in the database are determined as identity information, the first target features of the target detection boxes are used to update the second target features of the matched identities, and the first three-dimensional coordinates of the target detection boxes are used to update the second three-dimensional coordinates of the matched identities; a plurality of identities are newly created by using the target detection boxes that are not successfully matched with the identities in the database, and the identity similarities between the newly created identities and the identities in the database are determined, if the identity similarity is not less than a similarity threshold, the newly created identities are used to update the corresponding identities in the database, if the identity similarity is less than the similarity threshold, the newly created identities are added to the database; in addition, it can be judged whether the duration of the identity in the database remaining in a lost state exceeds a preset duration threshold, if yes, the corresponding identity is deleted, and thus the database of the current frame is updated to the database of the next frame.
[0149] According to a second aspect of the embodiments of the present disclosure, a target detection device is provided, please refer to the accompanying drawings Figure 3 , the device comprises:
[0150] The extraction module 301 is configured to extract at least one target detection frame in a plurality of different view angles of a current frame image and a first target feature of each target detection frame.
[0151] The coordinate module 302 is configured to determine a first three-dimensional coordinate of each target detection frame in a world coordinate system according to a position of each target detection frame in the current image frame.
[0152] The detection module 303 is configured to determine identity information of at least one target detection frame according to the first target feature and the first three-dimensional coordinate of each target detection frame and a second target feature and a second three-dimensional coordinate of each identity in a database.
[0153] In some embodiments of the present disclosure, the detection module is specifically configured to:
[0154] determine a first feature similarity between the first target feature of each target detection frame and the second target feature of each identity;
[0155] determine a first relative distance between the first three-dimensional coordinate of each target detection frame and the second three-dimensional coordinate of each identity;
[0156] determine the identity information of at least one target detection frame according to the determined first feature similarity and the first relative distance.
[0157] In some embodiments of the present disclosure, when the detection module is configured to determine the identity information of at least one target detection frame according to the determined first feature similarity and the first relative distance, the detection module is specifically configured to:
[0158] determine a matching degree of each matching combination according to a first feature weight, a first distance weight, and the determined first feature similarity and the first relative distance, wherein the matching combination includes one target detection frame and one identity;
[0159] determine a successfully matched matching combination according to the matching degree of each matching combination;
[0160] determine the identity information of the target detection frame included in the matching combination as the identity included in each successfully matched matching combination.
[0161] In some embodiments of the present disclosure, when the detection module is configured to determine the matching degree of each matching combination according to the first feature weight, the first distance weight, and the determined first feature similarity and the first relative distance, the detection module is specifically configured to:
[0162] determine a matching degree of each of the matching combinations according to the first feature weight, the first distance weight, a time coefficient of the identity in each of the matching combinations, and the determined first feature similarity and the first relative distance, wherein the time coefficient is determined according to a time when the identity is last tracked and a time corresponding to the current frame image.
[0163] In some embodiments of the present disclosure, further comprising a confidence module configured to:
[0164] determine a confidence of each of the target bounding boxes;
[0165] When the detection module is configured to determine a matching degree of each of the matching combinations according to the first feature weight, the first distance weight, and the determined first feature similarity and the first relative distance, the detection module is specifically configured to:
[0166] determine a matching degree of each of the matching combinations according to the first feature weight, the first distance weight, a confidence of the target bounding box in each of the matching combinations, and the determined first feature similarity and the first relative distance.
[0167] In some embodiments of the present disclosure, further comprising a new module configured to:
[0168] In the plurality of target bounding boxes without determined identity information, create at least one identity having a first target feature and a first three-dimensional coordinate according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes.
[0169] In some embodiments of the present disclosure, further comprising a merging module configured to:
[0170] determine an identity similarity between each of the newly created identities and each of the identities in the database according to the first target feature and the first three-dimensional coordinate of each of the newly created identities and the second target feature and the second three-dimensional coordinate of each of the identities in the database;
[0171] In a case where the identity similarity between any of the newly created identities and the identities in the database is greater than or equal to a preset similarity threshold, determine that the identity information of the target bounding box corresponding to the newly created identity is the corresponding identity in the database;
[0172] In a case where the identity similarity between any of the newly created identities and each of the identities in the database is less than the preset similarity threshold, add the newly created identity to the database.
[0173] In some embodiments of the present disclosure, further comprising a weight module configured to:
[0174] In the database, a second feature weight and a second distance weight of each identity are determined according to a time when each identity is last tracked and a time corresponding to the current frame image;
[0175] The merging module is configured to determine an identity similarity between each newly created identity and each identity in the database according to the first target feature and the first three-dimensional coordinate of each newly created identity and the second target feature and the second three-dimensional coordinate of each identity in the database, and specifically configured to:
[0176] The merging module is configured to determine an identity similarity between each newly created identity and each identity in the database according to the first target feature and the first three-dimensional coordinate of each newly created identity and the second target feature, the second three-dimensional coordinate, the second feature weight and the second distance weight of each identity in the database.
[0177] In some embodiments of the present disclosure, the updating module is further configured to:
[0178] After the identity information of at least one target detection frame is determined, the second target feature and the second three-dimensional coordinate of the corresponding identity in the database are updated according to the first target feature and the first three-dimensional coordinate of each target detection frame whose identity information is successfully determined; and / or,
[0179] In a case where the identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, the second target feature and the second three-dimensional coordinate of the identity in the database are updated according to the first target feature and the first three-dimensional coordinate of the newly created identity.
[0180] In some embodiments of the present disclosure, the state module is further configured to:
[0181] After the identity information of at least one target detection frame is determined, the state of the corresponding identity of each target detection frame whose identity information is successfully determined is updated to tracked in the database, and the state of other identities is updated to lost.
[0182] After the identity information of the target detection frame corresponding to the newly created identity is determined to be the corresponding identity in the database, the state of the corresponding identity in the database is updated to tracked.
[0183] In some embodiments of the present disclosure, the newly creating module is specifically configured to:
[0184] determine a second feature similarity between the first target feature of each of the target bounding boxes and a first target feature of each of the other target bounding boxes;
[0185] determine a second relative distance between the first three-dimensional coordinate of each of the target bounding boxes and a first three-dimensional coordinate of each of the other target bounding boxes;
[0186] determine a matching degree of each of the target bounding box combinations according to the first feature weight, the first distance weight, and the second feature similarity and the second relative distance between the two target bounding boxes in each target bounding box combination, wherein the target bounding box combination includes two different target bounding boxes;
[0187] create at least one identity according to the matching degree of each of the target bounding box combinations.
[0188] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method of the third aspect, and thus will not be described in detail here.
[0189] In a third aspect, the disclosure provides an apparatus, referring to the accompanying drawings Figure 4 which shows the structure of the apparatus, the apparatus comprising a memory and a processor, the memory being configured to store computer instructions executable on the processor, and the processor being configured to detect a target based on the method of any one of the first aspect when executing the computer instructions.
[0190] In a fourth aspect, the disclosure provides a computer readable storage medium having stored thereon a computer program, the program being executable by a processor to implement the method of any one of the first aspect.
[0191] In the disclosure, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood or implied to indicate or suggest relative importance. The term "multiple" refers to two or more, unless otherwise explicitly limited.
[0192] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. The disclosure is intended to cover any variations, uses, or adaptations of the disclosure following the general principles thereof and including such departures from the present disclosure as come within known or customary practice in the art to which the disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the disclosure are indicated by the following claims.
[0193] It should be understood that the present disclosure is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A target detection method characterized by, The method comprises: extracting at least one target bounding box in a current frame image of multiple different perspectives and a first target feature of each target bounding box; the multiple different perspectives of the current frame image are images collected by multiple different cameras from different perspectives for the same scene, and there is an overlap between the fields of view of the multiple different cameras; determining a first three-dimensional coordinate of each target bounding box in a world coordinate system according to the position of each target bounding box in the current frame image; determining the identity information of at least one target bounding box according to the first target feature and the first three-dimensional coordinate of each target bounding box, and the second target feature and the second three-dimensional coordinate of each identity in the database, comprising: determining the first feature similarity between the first target feature of each target bounding box and the second target feature of each identity; determining the first relative distance between the first three-dimensional coordinate of each target bounding box and the second three-dimensional coordinate of each identity; determining the matching degree of each matching combination according to the first feature weight, the first distance weight, the time coefficient of the identity in each matching combination, and the determined first feature similarity and first relative distance, wherein the time coefficient is determined by the time when the identity is last tracked and the time corresponding to the current frame image, and the matching combination comprises one target bounding box and one identity; determining the successfully matched matching combination according to the matching degree of each matching combination; determining the identity included in each successfully matched matching combination as the identity information of the target bounding box included in the matching combination.
2. The object detection method of claim 1, wherein, Further comprising: determining the confidence of each target bounding box; determining the matching degree of each matching combination according to the first feature weight, the first distance weight, and the determined first feature similarity and first relative distance, comprising: determining the matching degree of each matching combination according to the first feature weight, the first distance weight, the confidence of the target bounding box in each matching combination, and the determined first feature similarity and first relative distance.
3. The object detection method of claim 1, wherein, Further comprising: in the multiple target bounding boxes without determined identity information, creating at least one identity with the first target feature and the first three-dimensional coordinate according to the first target feature and the first three-dimensional coordinate of each target bounding box.
4. The object detection method of claim 3, wherein, Further comprising: determining the identity similarity between each newly created identity and each identity in the database according to the first target feature and the first three-dimensional coordinate of each newly created identity, and the second target feature and the second three-dimensional coordinate of each identity in the database; in the case where the identity similarity between any newly created identity and the identity in the database is greater than or equal to a preset similarity threshold, determining the identity information of the target bounding box corresponding to the newly created identity as the corresponding identity in the database. In a case that the identity similarity between any new identity and each identity in the database is less than the preset similarity threshold, the new identity is added to the database.
5. The object detection method of claim 4, wherein, Further comprising: In the database, a second feature weight and a second distance weight of each identity are determined according to a time when each identity is last tracked and a time corresponding to the current frame image; The identity similarity between each new identity and each identity in the database is determined according to the first target feature and the first three-dimensional coordinate of each new identity, and the second target feature and the second three-dimensional coordinate of each identity in the database, comprising: The identity similarity between each new identity and each identity in the database is determined according to the first target feature and the first three-dimensional coordinate of each new identity, and the second target feature, the second three-dimensional coordinate, the second feature weight and the second distance weight of each identity in the database.
6. The object detection method of claim 4, wherein, After the identity information of at least one target detection frame is determined, further comprising: updating the second target feature and the second three-dimensional coordinate of the corresponding identity in the database according to the first target feature and the first three-dimensional coordinate of each target detection frame whose identity information is successfully determined; And / or, In a case that the identity similarity between any new identity and the identities in the database is greater than or equal to the preset similarity threshold, further comprising: updating the second target feature and the second three-dimensional coordinate of the corresponding identity in the database according to the first target feature and the first three-dimensional coordinate of the new identity.
7. The object detection method of claim 3, wherein, The new identity with the first target feature and the first three-dimensional coordinate is created according to the first target feature and the first three-dimensional coordinate of each target detection frame, comprising: Determining the first target feature of each target detection frame and the second feature similarity between the first target feature of each target detection frame and the first target feature of other target detection frames; Determining the first three-dimensional coordinate of each target detection frame and the second relative distance between the first three-dimensional coordinate of each target detection frame and the first three-dimensional coordinate of other target detection frames; Determining the matching degree of each target detection frame combination according to the first feature weight, the first distance weight, and the second feature similarity and the second relative distance between two target detection frames in each target detection frame combination, wherein the target detection frame combination includes two different target detection frames; Creating at least one identity according to the matching degree of each target detection frame combination.
8. A target detection apparatus characterized by comprising: Comprising: The extraction module is configured to extract at least one target detection frame in a current frame image of multiple different perspectives and a first target feature of each target detection frame; The current frame image of multiple different perspectives is an image collected by multiple different cameras from different perspectives for the same scene, and there is an overlap between the fields of view of the multiple different cameras; The coordinate module is configured to determine a first three-dimensional coordinate of each target detection frame in a world coordinate system according to a position of each target detection frame in the current frame image; The coordinate module is configured to determine a first three-dimensional coordinate of each target detection frame in a world coordinate system according to a position of each target detection frame in the current frame image; The detection module is configured to determine identity information of at least one of the target bounding boxes according to the first target feature and the first three-dimensional coordinate of each of the target bounding boxes, and the second target feature and the second three-dimensional coordinate of each of the identity identifiers, including: determining a first feature similarity between the first target feature of each of the target bounding boxes and the second target feature of each of the identity identifiers; determining a first relative distance between the first three-dimensional coordinate of each of the target bounding boxes and the second three-dimensional coordinate of each of the identity identifiers; determining a matching degree of each of the matching combinations according to a first feature weight, a first distance weight, a time coefficient of the identity identifier in each of the matching combinations, and the determined first feature similarity and the first relative distance, wherein the time coefficient is determined by a time when the identity identifier is last tracked and a time corresponding to the current frame image, and each of the matching combinations includes one of the target bounding boxes and one of the identity identifiers; determining successfully matched matching combinations according to the matching degree of each of the matching combinations; and determining the identity identifier included in each of the successfully matched matching combinations as the identity information of the target bounding box included in the matching combination.
9. An electronic device, comprising: The device comprises a memory for storing computer instructions executable on a processor, and a processor for implementing the method of any one of claims 1 to 7 when executing the computer instructions.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program, when executed by a processor, implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-view target track generation method and device and electronic equipment
CN112686178A