Target recognition method, electronic device and computer-readable storage medium

Through the collaborative acquisition and filtering of target information by multiple cameras, and the attribute recognition combined with historical scheduling and attribute recognition conditions, the problem of low target recognition accuracy is solved and a more efficient target recognition effect is achieved.

CN119559411BActive Publication Date: 2025-05-09ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510132771.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-09
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The prior art has poor target recognition effect and defects in target recognition in long distances and relying on specific rules, resulting in low accuracy of target recognition.

Method used

By obtaining the target set to be detected and its positioning information in the images to be detected collected by multiple cameras, the candidate target set is filtered based on historical scheduling information and attribute recognition conditions, and the attribute recognition is performed, and the target fusion result is finally obtained through matching association.

Benefits of technology

It realizes the rapid acquisition of more structured information for targets with fewer computing resources, improving the accuracy of attribute recognition and overall accuracy of target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559411B_ABST
    Figure CN119559411B_ABST
Patent Text Reader

Abstract

The present application discloses a target recognition method, an electronic device and a computer-readable storage medium, the method comprising: obtaining a set of targets to be detected from a plurality of images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected; wherein, a plurality of cameras respectively collect corresponding images to be detected; based on historical scheduling information corresponding to the targets to be detected and pre-configured attribute recognition conditions, a candidate target set is screened from the set of targets to be detected, and attribute recognition is performed on the candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets; wherein, the historical scheduling information represents the scheduling interval for scheduling the targets to be detected for attribute recognition, and the attribute recognition conditions are at least related to the positioning information; based on the attribute recognition results and positioning information corresponding to the candidate targets, the candidate targets in the plurality of cameras are matched and associated to obtain target fusion results of the candidate targets. The above scheme can improve the accuracy of target recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a target recognition method, an electronic device and a computer-readable storage medium. Background Art

[0002] As the requirements for image processing accuracy continue to increase, the collection and analysis of target information in the image is a key technology. However, the current technical solutions have the following two problems. On the one hand, as the distance between the target and the camera increases, the effect of attribute recognition becomes worse and worse. On the other hand, attribute recognition depends on the triggering of specific rules, such as configuring detection lines. Otherwise, the target attributes cannot be obtained in time, and it is difficult to ensure the optimal position of attribute recognition, resulting in low accuracy of target recognition. Therefore, how to improve the accuracy of target recognition has become an urgent problem to be solved. Summary of the invention

[0003] The main technical problem solved by the present application is to provide a target recognition method, an electronic device and a computer-readable storage medium, which can improve the accuracy of target recognition.

[0004] To solve the above technical problems, the first aspect of the present application provides a target recognition method, comprising: obtaining a set of targets to be detected from multiple images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected; wherein, multiple cameras respectively collect corresponding images to be detected; based on historical scheduling information corresponding to the targets to be detected and pre-configured attribute recognition conditions, a set of candidate targets is screened from the set of targets to be detected, and attribute recognition is performed on the candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets; wherein, the historical scheduling information represents a scheduling interval for scheduling the targets to be detected for attribute recognition, and the attribute recognition condition is at least related to the positioning information; based on the attribute recognition results and positioning information corresponding to the candidate targets, the candidate targets in the multiple cameras are matched and associated to obtain target fusion results of the candidate targets.

[0005] To solve the above technical problem, the second aspect of the present application provides an electronic device, comprising a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method described in the first aspect.

[0006] In order to solve the above technical problem, the third aspect of the present application provides a computer-readable storage medium, which stores program instructions that can be executed by a processor, and the program instructions are used to implement the method described in the first aspect above.

[0007] The above scheme obtains a set of targets to be detected in a corresponding plurality of images to be detected respectively acquired by a plurality of cameras, and positioning information corresponding to the targets to be detected in the set of targets to be detected, screens the set of targets to be detected based on historical scheduling information corresponding to the targets to be detected and pre-configured attribute recognition conditions, obtains a set of candidate targets, performs attribute recognition on the candidate targets in the set of candidate targets, obtains attribute recognition results corresponding to the candidate targets, matches and associates the candidate targets in a plurality of cameras based on the attribute recognition results corresponding to the candidate targets and the positioning information corresponding to the candidate targets, outputs target fusion results of the candidate targets, reasonably allocates computing resources by analyzing the attribute recognition conditions and the historical scheduling information, achieves faster acquisition of structured information of more targets while using fewer computing resources, and obtains better attribute recognition effect, thereby improving the accuracy of target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0009] Figure 1 It is a flowchart of an implementation method of the target recognition method of the present application;

[0010] Figure 2 It is a flowchart of another implementation method of the target recognition method of the present application;

[0011] Figure 3 It is a schematic diagram of the target to be detected in this application being blocked;

[0012] Figure 4 It is a schematic diagram of the vehicle model and color metric distance matrix;

[0013] Figure 5 This is a schematic diagram of some multi-camera arrangements and sensing ranges supported by this application;

[0014] Figure 6 It is a structural schematic diagram of an embodiment of the electronic device of the present application;

[0015] Figure 7 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments, and different implementation methods can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0017] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.

[0018] See also Figure 1 , Figure 1 : is a flow chart of an implementation method of the target recognition method of the present application. The target recognition method includes:

[0019] S101: Acquire a set of targets to be detected in a plurality of images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected; wherein a plurality of cameras respectively capture corresponding images to be detected.

[0020] Specifically, a set of targets to be detected in a corresponding plurality of images to be detected respectively acquired by a plurality of cameras, and positioning information corresponding to the targets to be detected in the set of targets to be detected are obtained.

[0021] In one application, after multiple cameras respectively capture corresponding images to be detected, a target detection algorithm is used to obtain a set of targets to be detected in the multiple images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected.

[0022] In another application, after multiple cameras respectively capture corresponding images to be detected, a feature matching algorithm is used to obtain a set of targets to be detected in the multiple images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected.

[0023] Optionally, the positioning information corresponding to the target to be detected may be a 3D positioning result, such as coordinates and size, or may be 2D information of the target on the image, such as a rectangular bounding box.

[0024] Optionally, after obtaining the positioning information corresponding to the target to be detected, these targets can be continuously tracked based on a multi-target tracking algorithm, and continuous and unique tracking identifiers of the targets that match each camera are obtained.

[0025] It can be understood that the target set to be detected may include multiple different targets to be detected. When the image to be detected does not include the target to be detected, the target set to be detected may also be an empty set.

[0026] S102: Based on the historical scheduling information corresponding to the target to be detected and the pre-configured attribute recognition condition, a candidate target set is obtained from the target set to be detected, and the attributes of the candidate targets in the candidate target set are recognized to obtain the attribute recognition results corresponding to the candidate targets; wherein the historical scheduling information represents the scheduling interval for scheduling the target to be detected for attribute recognition, and the attribute recognition condition is at least related to the positioning information.

[0027] Specifically, based on the historical scheduling information corresponding to the target to be detected and the pre-configured attribute recognition conditions, the target set to be detected is screened to obtain a candidate target set, and the attributes of the candidate targets in the candidate target set are recognized to obtain the attribute recognition results corresponding to the candidate targets.

[0028] It should be noted that the historical scheduling information represents the scheduling interval for scheduling the target to be detected for attribute recognition, which can be described by the fact that the target to be detected has N frames without attribute recognition, and the attribute recognition condition is at least related to the positioning information corresponding to the target to be detected.

[0029] In one application method, the target set to be detected is first screened based on historical scheduling information, and then the screened target set to be detected is further screened based on pre-configured attribute recognition conditions to obtain a candidate target set, and attribute recognition is performed on the candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets.

[0030] In another application, the target set to be detected is screened based on historical scheduling information and weights corresponding to preset attribute recognition conditions to obtain a candidate target set, and attribute recognition is performed on the candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets.

[0031] In some application scenarios, the attribute recognition conditions include multiple attribute values, such as the pixel size of the target, the distance of the target, the occlusion rate of the target, the image quality of the target, etc.

[0032] S103: Based on the attribute recognition results and positioning information corresponding to the candidate targets, the candidate targets in multiple cameras are matched and associated to obtain target fusion results of the candidate targets.

[0033] Specifically, based on the attribute recognition results corresponding to the candidate targets and the positioning information corresponding to the candidate targets, the candidate targets in multiple cameras are matched and associated, and the target fusion results of the candidate targets are output.

[0034] In one application, the Hungarian algorithm is used to match and associate candidate targets in multiple cameras according to attribute recognition results and positioning information corresponding to the candidate targets, thereby obtaining target fusion results of the candidate targets.

[0035] In another application mode, a pre-trained convolutional neural network or other deep learning model is used to match and associate candidate targets in multiple cameras based on the attribute recognition results and positioning information corresponding to the candidate targets to obtain target fusion results of the candidate targets.

[0036] The above scheme obtains a set of targets to be detected in a corresponding plurality of images to be detected respectively acquired by a plurality of cameras, and positioning information corresponding to the targets to be detected in the set of targets to be detected, screens the set of targets to be detected based on historical scheduling information corresponding to the targets to be detected and pre-configured attribute recognition conditions, obtains a set of candidate targets, performs attribute recognition on the candidate targets in the set of candidate targets, obtains attribute recognition results corresponding to the candidate targets, matches and associates the candidate targets in a plurality of cameras based on the attribute recognition results corresponding to the candidate targets and the positioning information corresponding to the candidate targets, outputs target fusion results of the candidate targets, reasonably allocates computing resources by analyzing the attribute recognition conditions and the historical scheduling information, achieves faster acquisition of structured information of more targets while using fewer computing resources, and obtains better attribute recognition effect, thereby improving the accuracy of target recognition.

[0037] See also Figure 2 , Figure 2 : is a flow chart of another embodiment of the target recognition method of the present application. The target recognition method includes:

[0038] S201: Acquire a set of targets to be detected in a plurality of images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected.

[0039] Specifically, a plurality of cameras respectively capture corresponding images to be detected, obtain a set of targets to be detected in the plurality of images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected.

[0040] S202: Screening the target set to be detected based on historical scheduling information to obtain an initial target set.

[0041] Specifically, all the targets to be detected in the target set to be detected are screened based on the historical scheduling information to obtain an initial target set.

[0042] In one application scenario, historical scheduling information is described by N frames without attribute recognition. For example, when the target to be detected has not been identified for 1 frame, it means that the target to be detected has been identified in the previous frame. Therefore, the scheduling priority of the target to be detected will be ranked last. When the target to be detected has not been identified for 20 frames, it means that the target to be detected has not been identified for a long time. Therefore, the scheduling priority of the target to be detected will be ranked first, thereby performing a preliminary screening of all targets to be detected to obtain an initial target set.

[0043] In a specific application scenario, for the sake of convenience, the historical scheduling information is converted into discrete categories or intervals, and the mapping formula is as follows:

[0044]

[0045] (1)

[0046] Among them, mapping result 1 means "done in the previous frame", mapping result 2 means "done recently", mapping result 3 means "not done for a long time", and mapping result 4 means "never done". The higher the value, the higher the priority of this scheduling. and is the conversion threshold, which can be freely set according to the actual situation. If 5 is selected and there are 3 frames in which the target to be detected has not been attributed, it means that the target to be detected has been detected within 5 frames, which corresponds to the mapping result 2 "recently done".

[0047] S203: Screening the initial target set based on the attribute recognition condition to obtain a candidate target set.

[0048] Specifically, for targets with the same mapping results in the initial target set, they are screened again according to the attribute recognition condition to obtain a candidate target set.

[0049] In one application scenario, the attribute recognition condition includes multiple attribute values, and each attribute value is set with a corresponding weight. Some attribute values ​​are related to the positioning information of the target to be detected, and some attribute values ​​are related to the image area quality of the target to be detected in the image to be detected. Step S203 specifically includes: based on the weights corresponding to the multiple attribute values ​​included in the attribute recognition condition, the initial target set is screened to obtain a candidate target set.

[0050] Specifically, in order to adapt to different scenarios and functions, a weight can be set for each attribute value separately, and based on the multiple attribute values ​​and their corresponding weights, the initial target set is screened to obtain a candidate target set.

[0051] In a specific application scenario, the multiple attribute values ​​include target pixel size, target distance, target occlusion rate, target image quality score, etc.

[0052] It should be noted that the target pixel size refers to the height, width or area of ​​the target 2D detection frame. If there is no 2D detection frame but there is a target 3D detection frame, the target pixel size refers to the pixel height, width or area of ​​the circumscribed rectangle formed by the projection of the target 3D detection frame onto the image plane. The target distance can be the pixel distance from the bottom edge of the target 2D detection frame to the bottom edge of the image as a reference for the target distance. If there is no 2D detection frame but there is a target 3D detection frame, the target distance can be the distance from the target 3D center to the camera center or the projection of the camera center on the bottom surface as a reference for the target distance. The target occlusion rate refers to the degree to which the target to be detected is occluded by other foreground targets. Please refer to Figure 3 , Figure 3 This is a schematic diagram of the target to be detected in this application being blocked. The calculation formula of the target blocking rate is as follows:

[0053] (2)

[0054] in, represents the area of ​​the overlapped part, Indicates the area of ​​the obscured target.

[0055] The target image quality score mainly includes the target image clarity, lighting conditions, etc. The image quality can be scored manually first, and then a convolutional neural network is used to train a target quality score prediction model, and then the image quality can be scored using the trained target quality score prediction model.

[0056] Furthermore, taking the target occlusion rate as an example, if the target to be detected is severely occluded by other targets, the corresponding weight of this scheduling is lower, and if the target to be detected is not occluded by other targets, the corresponding weight of this scheduling is higher. For the sake of convenience, taking the attribute value of the target occlusion rate as an example, the target occlusion rate is converted into discrete categories or intervals, and the mapping formula is as follows:

[0057]

[0058] (3)

[0059] Among them, mapping result 3 means "not blocked", mapping result 2 means "slightly blocked", and mapping result 1 means "severely blocked". The higher the value, the higher the weight of this scheduling. is the conversion threshold, which can be freely set according to the actual situation. When the value is 0.5, targets with an occlusion rate greater than 0.5 are considered “severely occluded”.

[0060] It can be understood that other attribute values ​​can be mapped in a similar manner, so as to obtain a candidate target set consisting of K targets with the highest scheduling priority after weighted summation and sorting according to multiple attribute values ​​and their corresponding weights. By screening the targets to be detected through each attribute value and its corresponding weight included in the historical scheduling information and attribute identification conditions, the rationality of computing resource allocation can be further improved.

[0061] S204: Perform attribute recognition on the candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets.

[0062] Specifically, corresponding attribute recognition is performed on K candidate targets in the candidate target set to obtain attribute recognition results corresponding to the candidate targets.

[0063] Optionally, taking a vehicle target as an example, its attributes mainly include license plate recognition, vehicle model recognition, vehicle color recognition, etc. In other application scenarios, corresponding attribute recognition can also be performed on other types of targets, and this application does not impose specific restrictions on this.

[0064] In an application scenario, step S204 specifically includes: in response to the number of times that attribute recognition is performed on the candidate target meeting a preset threshold condition, obtaining multiple candidate recognition results; and obtaining an attribute recognition result based on the multiple candidate recognition results.

[0065] Specifically, when the number of attribute identifications for a candidate target meets a preset threshold condition, multiple candidate identification results are obtained, and based on the multiple candidate identification results, the attribute identification results corresponding to the candidate target are obtained. Through multiple identifications, the errors and uncertainties that may occur in a single identification can be effectively reduced, thereby improving the accuracy of the overall attribute recognition.

[0066] In a specific application scenario, taking the vehicle color recognition result as an example, the vehicle color candidate recognition results of the vehicle target in the past 5 frames are {black, black, black, dark blue, dark blue}, so the vehicle color recognition result of the vehicle target is black. Similarly, the vehicle model and license plate recognition can also obtain the corresponding recognition results in a similar way.

[0067] In one embodiment, the target to be detected and the candidate target have corresponding identifiers, and after step S204, the method further includes: adding the attribute recognition result to the attribute cache information of the candidate target based on the identifier of the candidate target.

[0068] Specifically, after the attribute identification of the candidate target is completed, the attribute identification result will be cached in the attribute cache information of the candidate target according to the identifier corresponding to the candidate target. The caching mechanism can significantly reduce the need for repeated complex calculations, thereby further improving the rationality of computing resource allocation.

[0069] In one implementation scenario, the target identification method also includes: in response to the current target triggering a preset rule, determining whether the current target is a candidate target, and if so, obtaining an attribute identification result from the attribute cache information; if not, performing attribute identification on the current target, and in response to the number of times the attribute identification is performed on the current target satisfying a preset threshold condition, obtaining multiple reference identification results of the current target, and based on the multiple reference identification results, obtaining an attribute identification result of the current target and adding the attribute identification result of the current target to the corresponding attribute cache information.

[0070] Specifically, for the current target that triggers a pre-set specific rule, for example, when the current target violates traffic rules, it will first be determined whether the target is a candidate target in the candidate target set. If so, the attribute recognition result is directly obtained from the attribute cache information of the corresponding candidate target. If not, the attribute recognition is performed on the current target. When the number of attribute recognitions for the current target meets the preset threshold condition, multiple reference recognition results are obtained, and based on the multiple reference recognition results, the attribute recognition result corresponding to the current target is obtained, and the attribute recognition result is cached in the corresponding attribute cache information, so that the attribute recognition of the current target and the supplement of the attribute information can be completed, avoiding the missing of attributes when the target information needs to be collected.

[0071] Optionally, for the current target that triggers a pre-set specific rule, if the target is also a candidate target in the candidate target set, but lacks one of the attribute identifications, for example, the vehicle target lacks license plate recognition, the license plate recognition operation will also be performed in addition, and the license plate recognition result will be updated to the attribute cache information corresponding to the vehicle target.

[0072] Optionally, if the attribute recognition result in the attribute cache information corresponding to the candidate target has not been updated for a long time, for example, attribute recognition needs to be redone after many frames have not been performed, the latest and best attribute recognition result is updated to the attribute cache information corresponding to the candidate target.

[0073] S205: Determine a reference distance between candidate targets based on the attribute recognition result and the positioning information.

[0074] Specifically, a reference distance between candidate targets is determined based on the attribute recognition result and the positioning information.

[0075] It should be noted that in order to ensure that the target information in the multi-camera fusion stage does not deviate greatly due to time differences, in some implementation scenarios, PTP (Precision Time Protocol) or NTP (Network Time Protocol) will be used to synchronize the frames from each camera at the same time. And because the positioning information output by each camera may be different in their respective coordinate systems, it must be converted to a unified coordinate system based on the calibrated camera extrinsics before multi-camera target fusion.

[0076] In an application scenario, the positioning information corresponds to Euclidean distance and heading angle difference, and the attribute recognition result corresponds to multiple attribute distances. Step S205 specifically includes: determining a reference distance between candidate targets based on the Euclidean distance, heading angle difference and multiple attribute distances.

[0077] Specifically, the reference distance between the candidate targets is calculated based on the Euclidean distance corresponding to the positioning information, the heading angle difference, and the multiple attribute distances corresponding to the attribute recognition result.

[0078] In a specific application scenario, taking a vehicle target as an example, the specific calculation process of the reference distance is as follows:

[0079] (4)

[0080] in, is the reference distance, is the Euclidean distance, is the heading angle difference, , , They are the distances of attributes such as car model, car color, and license plate, or the distances after similarity conversion. , , , , is the weight of the distance in each dimension. For the car model and car color, the metric distance matrix is ​​designed to obtain the distance after similarity conversion. Figure 4 , Figure 4 This is a schematic diagram of the vehicle model and color metric distance matrix. The value is the distance size, and the license plate is calculated based on the edit distance (Levenshtein distance) as a metric.

[0081] S206: Match and associate the candidate targets in multiple cameras based on the reference distance to obtain target fusion results of the candidate targets.

[0082] Specifically, the same candidate targets in multiple cameras are matched and associated based on the reference distance to obtain target fusion results of the candidate targets.

[0083] In one application scenario, see Figure 5 , Figure 5 This is a schematic diagram of some multi-camera arrangements and perception ranges supported by this application. Among them, targets in non-common view areas are generated by the camera that observes the only source, targets in common view areas are determined based on the size of the area inversely projected onto the image coordinate system of each camera, and are generated by the camera with the largest area and marked as the main camera of the common view area. Targets from multiple cameras in the common view area are matched and associated based on the calculated reference distance and the Hungarian algorithm, and when a target enters the observation area of ​​one camera from the observation area of ​​another camera, it will switch according to the associated target information, and the main camera is responsible for outputting the target in the common view area.

[0084] It should be noted that, taking the scene where the target is a vehicle passing through an intersection as an example, Figure 5 For the middle scene (a), the target is located in the field of view of the telephoto camera and the short-focus camera at the same time, that is, the target is located in the common viewing area of ​​the telephoto camera and the short-focus camera. Since the projection area of ​​the common viewing area is larger on the telephoto camera, the telephoto camera is used as the main camera in the common viewing area, and the target in the common viewing area is generated by it. When outputting the target fusion result, the result recognized by the main camera is also mainly used.

[0085] Furthermore, for Figure 5 For scene (b), the target is in the common view area of ​​camera 1 and camera 2. At this time, it can be determined based on the size of the area inversely projected onto the image coordinate system of each camera according to the common view area. The camera with the largest area is responsible for generating it and marked as the main camera of the common view area. For example, the area of ​​the target projected onto the image coordinate system of camera 2 is larger than the area projected onto the image coordinate system of camera 1. Therefore, camera 2 can be used as the main camera of the common view area. When the area of ​​the target projected onto the image coordinate system of camera 1 is larger than the area projected onto the image coordinate system of camera 2, camera 1 can be used as the main camera of the common view area. This application does not impose any restrictions on this.

[0086] Furthermore, for Figure 5 For scene (c), the target moves from the field of view of camera 2 to the field of view of camera 1. At this time, when the target is only in the field of view of camera 2, camera 2 is responsible for outputting the target. When the target moves to the common viewing area of ​​camera 1 and camera 2, since the area of ​​the target projected on the image coordinate system of camera 2 is larger than the area projected on the image coordinate system of camera 1 at the beginning, camera 2 is the main camera at this time. When camera 2 cannot capture the entire target and the target is in the field of view of camera 1, camera 1 is switched to be the main camera, and camera 1 is responsible for outputting the target. When the target is only in the field of view of camera 1, camera 1 continues to be responsible for outputting the target.

[0087] The above solution, through multi-camera collaboration, can obtain accurate trajectory information of the target in a larger range and high-definition images of multiple locations, so as to achieve more accurate behavior analysis and output reliable attribute recognition results in a timely manner.

[0088] See also Figure 6 , Figure 6 6 is a structural diagram of an embodiment of an electronic device of the present application, wherein the electronic device 60 comprises a memory 600 and a processor 602 coupled to each other, wherein the memory 600 stores program data (not shown), and the processor 602 calls the program data to implement the method in any of the above embodiments, and the description of the relevant contents can be found in the detailed description of the above method embodiments, which will not be repeated here. Specifically, the electronic device 60 includes but is not limited to: a desktop computer, a laptop computer, a tablet computer, a server, etc., which are not limited here. In addition, the processor 602 can also be referred to as a CPU (Center Processing Unit). The processor 602 may be an integrated circuit chip with signal processing capabilities. The processor 602 may also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. In addition, the processor 602 may be implemented by an integrated circuit chip.

[0089] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 70 stores program data 700. When the program data 700 is executed by a processor, the method in any of the above embodiments is implemented. For descriptions of related contents, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0090] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present implementation scheme.

[0091] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., and other media that can store program codes.

[0093] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A target recognition method, characterized in that: include: Acquire a set of targets to be detected in a plurality of images to be detected, and positioning information corresponding to the targets to be detected in the set of targets to be detected; wherein a plurality of cameras respectively acquire corresponding images to be detected; Based on the historical scheduling information corresponding to the target to be detected and the pre-configured attribute recognition conditions, a candidate target set is obtained by screening from the target to be detected set, and the attributes of the candidate targets in the candidate target set are recognized to obtain the attribute recognition results corresponding to the candidate targets; wherein the historical scheduling information represents the scheduling interval for scheduling the target to be detected for attribute recognition, and the attribute recognition conditions are at least related to the positioning information; wherein, based on the historical scheduling information corresponding to the target to be detected and the pre-configured attribute recognition conditions, a candidate target set is obtained by screening from the target to be detected set, and the attributes of the candidate targets in the candidate target set are recognized to obtain the attribute recognition results corresponding to the candidate targets, including: screening the target to be detected set based on the historical scheduling information to obtain an initial target set; screening the initial target set based on the attribute recognition conditions to obtain the candidate target set; recognizing the attributes of the candidate targets in the candidate target set to obtain the attribute recognition results corresponding to the candidate targets, and the historical scheduling information is described by the fact that the target to be detected has N frames without attribute recognition; Based on the attribute recognition results and positioning information corresponding to the candidate targets, the candidate targets in the multiple cameras are matched and associated to obtain target fusion results of the candidate targets.

2. The method according to claim 1, characterized in that: The attribute recognition condition includes a plurality of attribute values, and each of the attribute values ​​is correspondingly provided with a weight, some of the attribute values ​​are related to the positioning information of the target to be detected, and some of the attribute values ​​are related to the image area quality of the target to be detected in the image to be detected; The screening of the initial target set based on the attribute recognition condition to obtain the candidate target set includes: Based on the weights corresponding to the multiple attribute values ​​included in the attribute recognition condition, the initial target set is screened to obtain the candidate target set.

3. The method according to claim 1, characterized in that The performing attribute recognition on the candidate target in the candidate target set to obtain the attribute recognition result corresponding to the candidate target includes: In response to the number of times the attribute recognition of the candidate target is performed meeting a preset threshold condition, obtaining a plurality of candidate recognition results; Based on the plurality of candidate recognition results, the attribute recognition result is obtained.

4. The method according to claim 1, characterized in that: The target to be detected and the candidate target have corresponding identifiers; After performing attribute recognition on the candidate targets in the candidate target set and obtaining the attribute recognition results corresponding to the candidate targets, the method further includes: The attribute recognition result is added to the attribute cache information of the candidate target based on the identifier of the candidate target.

5. The method according to claim 4, characterized in that Also includes: In response to the current target triggering a preset rule, determining whether the current target is the candidate target, and if so, obtaining the attribute recognition result from the attribute cache information; If not, perform attribute identification on the current target, and in response to the number of times the attribute identification is performed on the current target satisfying a preset threshold condition, obtain multiple reference identification results of the current target, and based on the multiple reference identification results, obtain the attribute identification result of the current target and add the attribute identification result of the current target to the corresponding attribute cache information.

6. The method according to claim 1, characterized in that The matching and associating the candidate targets in the multiple cameras based on the attribute recognition results and positioning information corresponding to the candidate targets to obtain the target fusion results of the candidate targets includes: Determining a reference distance between the candidate targets based on the attribute recognition result and the positioning information; The candidate targets in the multiple cameras are matched and associated based on the reference distance to obtain target fusion results of the candidate targets.

7. The method according to claim 6, characterized in that The positioning information corresponds to the Euclidean distance and the heading angle difference, and the attribute recognition result corresponds to a plurality of attribute distances, and the attribute distances correspond to the distances of the plurality of attributes of the target to be detected or the distances after similarity conversion; The determining, based on the attribute recognition result and the positioning information, a reference distance between the candidate targets, comprises: The reference distance between the candidate targets is determined based on the Euclidean distance, the heading angle difference and the multiple attribute distances.

8. An electronic device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target detection method and device for driving scene, equipment and storage medium

    CN114898314A

  • Target attribute identification method, electronic equipment and computer readable storage medium

    CN119313994A